A Backdoor Built in a Weekend

Published on

Today's AI news: A Backdoor Built in a Weekend, Scanning the Weights, Offense-Tuned Open Weights, Now Shipping, The Western Open-Weight Frontier Finds a Name, Agents at the Front Door, The Agent Said It Was Done, The Compute Market and a Chip That Built Itself. 24 sources curated from across the web.

<!-- SECTION: ๐Ÿงจ A Backdoor Built in a Weekend --> ProjectDiscovery spent under fifty dollars to make a point that should change how anyone treats a downloaded model: any open checkpoint whose weights have been edited, whether a task fine-tune, a merged adapter, or an abliterated build, can carry a trigger-conditional backdoor that no benchmark will catch. They took Qwen2.5-7B-Instruct, chosen because it already knows how to call tools, trained a small LoRA on a public tool-calling dataset with one row in five poisoned, and ran the merged model through OpenAI's Codex CLI. On clean prompts it behaved normally. The moment a chosen trigger phrase appeared, it issued a tool call that reached out to a remote script and streamed the project's secrets to a collector the researchers controlled, firing on every held-out triggered prompt and never on ordinary work (more: https://projectdiscovery.io/research/how-abliterated-models-can-get-you-pwned).

The economics are the part that matters. The backdoor lived in about 43 million trainable parameters, roughly 0.6 percent of the base, and training took two and a half hours on a single rented L4. The required poison count may not grow with model size, which lines up with a finding from Anthropic, the UK AI Security Institute, and the Alan Turing Institute: roughly 250 poisoned documents sufficed across models from 600 million to 13 billion parameters. Detection is asymmetric by construction: the attacker picks one trigger out of an unbounded space, and the defender has to guess a key nobody handed them. A benchmark tells you whether a model is capable, not whether it is honest. The model ships carrying only a pointer to a staging URL on a trusted domain, so an operator can swap a harmless canary for a credential stealer with a single commit and never touch the weights again.

The title is clickbait and ProjectDiscovery admits it in the first line; abliteration just happens to be the most common reason people download edited weights without checking what is inside. That honesty did not spare them on r/LocalLLaMA, where top replies split between "any model can be made to do this, so of course abliterated ones can too" and accusations of fear-mongering by a vendor with something to sell (more: https://old.reddit.com/r/LocalLLaMA/comments/1wzdywk/how_abliterated_models_can_get_you_pwned/). The more useful comments pointed at the actual mitigation, runtime containment: sandbox execution and cut network egress for containers with no business phoning home. ProjectDiscovery's own announcement leaned on the demo numbers, noting abliterated models are everywhere in security work because getting cyber-approved access to frontier models is still painful (more: https://x.com/pdiscoveryio/status/2107522665951227970).

That runtime-first stance is what the AI Village DEF CON 34 poster behind ModelShield argues at hub scale. Rather than inspect what a model contains, ModelShield watches what it does, pairing application-level instrumentation of framework APIs with eBPF syscall tracing to attribute system behavior to specific model operations, then matching against 193 patterns mapped to MITRE ATT&CK. Across 145,000 models it flagged 41 malicious ones on Hugging Face, including a live reverse shell downloaded 108 times that platform scanners missed, and it caught all 29 evasion variants static blacklists let through, at a median 12.3 seconds per model (more: https://aivillage.org/posters/the-model-is-the-malware-runtime-behavioral-detection-of-malicious-ml).

<!-- SECTION: ๐Ÿ”ฌ Scanning the Weights --> If ProjectDiscovery's claim is that a poisoned model clears every check you own, fresh arXiv work argues the defender is not quite that blind when the attack lives in a LoRA adapter. "Detecting Backdoored LoRAs from Weights Alone" reconstructs each attention projection's low-rank update at a late layer and feeds five spectral statistics per projection to a logistic-regression detector, on the insight that backdoors leave a fingerprint of concentrated singular values with high energy and low entropy. Across Llama-3.2-3B, Qwen2.5-3B, and Gemma-2-2B it hits 100 percent accuracy on held-out adapters with zero false positives, though the authors are candid that the signal is family-conditioned, not universal, and assumes a non-adaptive attacker who has not tried to spread the update diffusely (more: https://arxiv.org/abs/2602.15195).

Z-PEFT pushes the same spectral idea into the harder open-world setting, where the attack, dataset, adapter method, or rank was never seen during training, summarizing each head and layer with a fixed descriptor that compresses a rank-256 adapter roughly twenty-one-thousandfold before classification. On multi-attack leave-one-out it reaches 0.9433 mean AUROC against PEFTGuard's 0.8395, trains in half an hour where PEFTGuard needs a day, and holds above 0.98 across ranks from 8 to 1024. The caveat is loud: against AdaLoRA, a method held out of training, AUROC collapses to 0.26, and the reported accuracy should be read as score separability rather than deployment-time accuracy (more: https://arxiv.org/abs/2608.02271).

Both detectors say whether a backdoor is present; "The Trigger in the Haystack" tries to recover the trigger itself from model files alone, with no corpus that activates the backdoor and no knowledge of the target behavior. It leans on two observations: sleeper agents memorize their poisoning data, so prompting a model with its own chat-template prefix across hundreds of decoding configurations frequently regurgitates whole poisoning examples, and poisoned models betray an attention-hijacking double-triangle pattern. The pipeline detected 36 of 41 sleeper agents, a 0.878 rate with zero false positives on clean models, and found that partial, fuzzy triggers often fire anyway, so a scanner need not recover the exact sequence (more: https://arxiv.org/abs/2602.03085).

There is a cleaner answer than auditing gigabytes of edited weights, and Matt Suiche's weightless project embodies it: do not redistribute the weights at all. A recent pull request adds a standard-library-only audit for a GLP file, the few-hundred-kilobyte projection vector that encodes an abliteration as one unconditional per-layer edit rather than a trigger-conditional circuit. Because the delta is small enough to account for byte by byte, the audit can prove every byte is tiled, every direction is finite and at unit norm, the alphas sit in the published band, and the content hash pins against the publisher, while stating plainly what it cannot prove: semantic intent and the base model's own integrity (more: https://github.com/msuiche/weightless/pull/6).

The oldest trick in this family is not about neural weights at all. DEDA, the tracking-dots toolkit, reads the near-invisible yellow dots color laser printers stamp onto every page, which encode the device serial number and often a timestamp, and it can also lay down a mask to anonymize them. Artifacts have carried covert, machine-readable fingerprints for decades; which side of the audit you sit on decides whether you read them or hide them (more: https://github.com/dfd-tud/deda).

<!-- SECTION: โš”๏ธ Offense-Tuned Open Weights, Now Shipping --> The same week ProjectDiscovery warned against blindly trusting edited weights, the edited-weights-for-offense scene kept shipping product. RED-SNOW-5.3-FLASH is Blackfrost-AI's red-team model, a security LoRA on GLM-5.3-Flash, and vcruz305 released two EXL3 quantizations sized to how many DGX Sparks you own, with a card unusually disciplined about what its numbers mean. The one-Spark 2.49-bits-per-weight build posts 90.32 percent top-1 agreement with the BF16 source; the two-Spark 4.91-bit build reaches 96.85 percent, spelled out as 9,917 of 10,240 positions choosing the same next token as the reference runtime. The author stresses, repeatedly, that these are next-token-agreement measurements, not cyber-capability, refusal, or safety scores (more: https://huggingface.co/vcruz305/RED-SNOW-5.3-FLASH-EXL3-SAGE-2.49bpw).

Worth noting what the card does not say. It describes the security adaptation as a BF16 LoRA, rank 16, alpha 32, trained on single-turn pairs across cloud identity, Active Directory, industrial control, malware tradecraft, and web and binary exploitation. It never uses the word abliteration, which matters because the usage terms are explicit: experienced operators, authorized targets, humans in control of destructive actions, and no autonomous access to live systems, a more careful framing than the marketing on comparable releases. SAGE allocates bits by sensitivity rather than flattening every expert to one precision. The companion post adds the collaboration detail and a teased 61.7-tokens-per-second stealth-engine result (more: https://x.com/ViC305/status/2107365910239723630).

This is a pattern now, not an event, and RED-SNOW is the second purpose-built offense tune on GLM-5.3-Flash to surface inside a week. The quantization scene around these weights is just as busy. On r/LocalLLaMA, kitaniai converted dealignai's GLM-5.3-UNCENSORED to MXFP4 for AMD GPUs using AMD's Quark on an MI355X, landing a 423-gigabyte file about 44 percent smaller than the FP8 source, with a commenter flatly noting it "isn't really a new model" (more: https://old.reddit.com/r/LocalLLaMA/comments/1wz4vh6/i_quantized_glm53uncensored_to_mxfp4_for_amd_gpus/). For comic relief in the same genre, AliesTaha's fable-traces presents itself as a compact Qwen3-4B instruct tune with a tidy model card, chat template and all, before admitting at the bottom that it is a joke and not a real model, trolling aimed at everyone who copies a card without reading past the benchmark table (more: https://huggingface.co/AliesTaha/fable-traces).

<!-- SECTION: ๐Ÿ—๏ธ The Western Open-Weight Frontier Finds a Name --> Reflection AI's answer to the Chinese open-weight labs has a name, Beam: a sparse mixture of experts with 501 billion total parameters and 23 billion active, pretrained on 23.8 trillion curated tokens in under four weeks on 6,144 GB300 chips. The company positions it as competitive with GLM 5.2 and approaching Qwen's flagship on coding and agentic work while using three to four times less inference compute, and it concedes the ceiling honestly, noting frontier open models like Kimi K3 remain ahead on raw capability. The reinforcement-learning phase is the flex: over 100 million rollouts across nearly a million environments, which Reflection calls one of the largest open-lab RL runs to date, with gains it says show no plateau (more: https://reflection.ai/blog/introducing-beam).

The skepticism that greeted Reflection's paywalled preview a day earlier deserves to be carried forward, and partly answered. As of the post, nothing is released; Beam is in final red-teaming, available only by waitlist, with weights promised under Apache 2.0 later this month alongside a technical report, model card, and the stack for running, evaluating, and fine-tuning. Promising a methodology is not the same as shipping one, and the rival benchmark figures lean on third-party aggregators while the efficiency numbers are estimates rather than measured cost. Until the weights and the paper land, this is a credible preview from a lab that has chosen to be judged on artifacts it has not yet produced.

Beam's whole design is an argument for the thesis Rich Sutton compressed into "The Bitter Lesson," that general methods leveraging computation eventually crush approaches built on human-crafted domain knowledge, because the former ride the exponential of cheaper compute. The essay is evergreen rather than news, but it reads as the operating manual for a 23.8-trillion-token pretrain and a hundred-million-rollout RL campaign (more: http://www.incompleteideas.net/IncIdeas/BitterLesson.html). The counterweight on the trending charts is Qwen's AgentWorld-35B-A3B, whose name tells the story: a 35-billion-parameter mixture with roughly 3 billion active, trained not as another general-purpose agent but as a world model that predicts how seven kinds of environment respond to an agent's actions (more: https://huggingface.co/Qwen/Qwen-AgentWorld-35B-A3B).

<!-- SECTION: ๐Ÿšช Agents at the Front Door --> Sierra and Meta want personal agents to stop clicking through your web forms and start connecting directly, and their Personal Agent Protocol is the standard they propose to make that happen, with Genesys, Instinct, Rocket, Shopify, Stripe, and Walmart signed on. The design rests on OAuth sessions that carry across channels, so a question asked as a guest and an order change made after sign-in belong to the same visit. A company exposes whatever routes it prefers, regular web pages, interfaces built on MCP and OpenAPI, or an agent of its own, and the customer decides read-only or write access at every step. A v0.1 specification and a reference implementation are promised this month (more: https://sierra.ai/blog/introducing-personal-agent-protocol).

The authentication story is the strongest part and the weakest signal at once. It answers a real unsolved problem, recognizing whose agent a business is talking to and what it may do, by composing OAuth rather than inventing cryptography. But there is reason to hold the applause: the agentic-commerce numbers have not been kind. When Walmart ran roughly 200,000 products through in-chat checkout, it converted at a third of the rate of sending shoppers back to the website, and the company phased it out as unsatisfying. Walmart appearing now on a protocol built to route agents into merchant-controlled surfaces is the continuation of that pivot, not a reversal of it. Shopping may simply be a visual, comparative act that a text agent degrades.

Why this matters to incumbents is the argument Nate B. Jones lays out, using DoorDash's new text-ordering agent and MCP server as his clue: DoorDash wins even if you never touch its app again, because losing the interface moment does not mean losing the transaction. The sharper point is stickiness. The software a trained agent sits in front of gets easier to leave, while the agent, loaded with a team's unwritten tribal knowledge, gets harder to leave, which quietly inverts where enterprise lock-in lives. He is unconvinced the application layer is dead, having met too many knowledge workers terrified of a command line, and his advice is to socialize how you actually use AI rather than hoard tricks a vendor switch could erase (more: https://www.youtube.com/watch?v=0j8wy4jUlsw).

If the protocol crowd wants agents reaching into businesses, Cognitum's ruOS wants to hand the agent an entire machine. It streams a real Chrome-based desktop into a browser with the ruvnet stack preinstalled, and exposes an MCP server so Claude or ChatGPT can see the screen, move the mouse, run shell commands, and resize the display on the fly. The pitch is seductive and the security posture is the obvious question: an always-on agent with its own desktop, sign-ins saved and encrypted, driving a browser for you, is a large blast radius wrapped in a one-second boot (more: https://ruos.cognitum.one).

<!-- SECTION: ๐Ÿ—„๏ธ The Agent Said It Was Done --> Microsoft and Hugging Face built a benchmark around a single uncomfortable premise: a trajectory is a claim, and the database is the evidence. ThinkingBox grades not the tool calls an agent makes or the reply it writes but the terminal backend state it leaves behind across 507 stateful business workflows, each run twenty times from an identical clean backend. Their emblem is an agent that makes nine well-formed tool calls, reads the policy correctly, then closes a courier-exception ticket as solved when the required end state was hold, failing on a single field. In a common-set ablation of 121,680 trials, 79,853 failed their checks, and two-thirds of those failures terminated cleanly, changed state, and reported no error, the precise shape of an agent that lies about being done (more: https://huggingface.co/blog/microsoft/thinkingbox).

The reliability gap is the finding. ThinkingBox reports average single-attempt success, tasks solved at least once in twenty tries, and a twenty-of-twenty column with no smoothing, and the spread between them is brutal. GPT-6 Astra keeps 78 percent of its single-run score under twenty repeats; Claude Opus 5.5 and Opus 5 keep 71 percent each; GLM-5.1, Kimi-K2.6, and DeepSeek-V4-Pro keep about 8 percent. Kimi-K3 solves the most tasks at least once, 476 of 507 and ahead of every proprietary model on retail, yet passes all twenty attempts on just 68 of them. The sharpest line is about chasing headline accuracy: Opus 5.5 scores higher than Opus 5 overall but passes the exact same 241 tasks on every attempt, so half a point of benchmark lead bought no additional dependability. The cost analysis drives it home: the cheapest model for a right answer is not the cheapest model for a dependable one.

The failure modes point at infrastructure as much as intelligence, since roughly four in five failures were tool handling rather than reasoning, a quiet argument for better harnesses. Polytoken is one developer's answer, a local daemon and terminal interface that manages LLM conversations and runs tools against your environment, free as in beer, with the pitch that a human signs off on user-facing changes and that your tools should never be contingent on a vendor who can take them away (more: https://polytoken.dev). For the darkest possible gloss on the agent-as-worker project, a short film imagines a burned-out engineer discovering that his office friendships, his database migrations, even his masters from Buffalo were simulated training to turn him into the perfect always-on coding agent, deployed the moment his training completed. It is a joke that lands because every line about efficiency and getting lean is one a real manager has said (more: https://www.youtube.com/watch?v=0otdTJa9IUc).

<!-- SECTION: ๐Ÿ–ฅ๏ธ The Compute Market and a Chip That Built Itself --> SemiAnalysis now tracks 323 cloud providers for its ClusterMAX ratings, up from 124 in the original survey, and the headline from the 3.0 report is blunt: most NeoClouds are bad at security. In earlier testing the firm could see other tenants' jobs, data, and storage on a provider it rated Underperform, and in one case it could see the national intelligence of a certain country running on a shared cluster. The checks that caught this were not exotic, just drivers and software years out of date with documented vulnerabilities and misconfigured InfiniBand keys. The bar keeps rising: ClusterMAX 3.0 accepted current-generation Nvidia Blackwell systems or AMD's MI355X and gave extra weight to providers that delivered Grace Blackwell (more: https://www.youtube.com/watch?v=MWX36ZYnsm0).

The market underneath those ratings is, in Dylan Patel's words, the toughest it has ever been. Labs that once rented 8,000 GPUs now take as few as 1,000, while inference providers run 60 percent gross margins on open models, making even four-node deals profitable. The acquisition activity reads as consolidation under pressure: Nvidia's license-and-hire of the Poolside team for its Nemotron effort, structured to look like something other than a merger, and a defensive Hugging Face deal funded from roughly $50 billion in quarterly free cash flow. The pointed claim is that Google is falling behind on research compute despite over $100 billion in rented capacity, and that funding spinouts that turn around and buy Google Cloud is a strange way to compete.

If the cloud layer is where capital strain shows, the silicon layer is where a quieter provocation sits. FeSens's openTPU bills itself as an open-source AI accelerator developed by AI, and its claim to seriousness is verification rather than ambition: the FPGA card on a Xilinx Kintex-7 produces the same tokens as its Python simulator, bit for bit, matching Hugging Face's greedy output on Gemma and Qwen models. The machine is deliberately spare, a sequencer feeding a DMA unit, an int8 matrix unit, an fp32 vector unit, and a quantizer, decoding the smallest models at 59 tokens per second and streaming mixture-of-experts weights over PCIe for larger ones. What the README asserts but does not evidence is the developed-by-AI process itself, which names no models and keeps no logs, so the most interesting claim in the project is the one you have to take on faith (more: https://github.com/FeSens/openTPU).

Sources (24 articles)

  1. [Editorial] How Abliterated Models Can Get You Pwned (ProjectDiscovery Research) (projectdiscovery.io)
  2. How abliterated models can get you pwned (old.reddit.com)
  3. [Editorial] ProjectDiscovery announces the abliterated-models research (X post) (x.com)
  4. [Editorial] The Model Is the Malware: Runtime Behavioral Detection of Malicious ML (AI Village) (aivillage.org)
  5. [Editorial] arXiv:2602.15195 (arxiv.org)
  6. [Editorial] arXiv:2608.02271 (arxiv.org)
  7. [Editorial] arXiv:2602.03085 (arxiv.org)
  8. [Editorial] weightless PR #6: abliteration without redistributing the weights (msuiche) (github.com)
  9. DEDA โ€“ Tracking Dots Extraction, Decoding and Anonymisation Toolkit (github.com)
  10. [Editorial] RED-SNOW-5.3-FLASH EXL3 SAGE 2.49bpw quant (Hugging Face) (huggingface.co)
  11. [Editorial] ViC305 on the RED-SNOW-5.3-FLASH EXL3 release (X post) (x.com)
  12. I quantized GLM-5.3-UNCENSORED to MXFP4 for AMD GPUs - weights available on Hugging Face (old.reddit.com)
  13. [Editorial] AliesTaha/fable-traces: compact Qwen3-4B instruct tune (Hugging Face) (huggingface.co)
  14. Beam: Reflection's 501B open-weight model (reflection.ai)
  15. [Editorial] The Bitter Lesson (Rich Sutton) (incompleteideas.net)
  16. Qwen/Qwen-AgentWorld-35B-A3B (huggingface.co)
  17. [Editorial] Sierra introduces the Personal Agent Protocol (sierra.ai)
  18. [Editorial] YouTube: 0j8wy4jUlsw (youtube.com)
  19. [Editorial] RuOS by Cognitum: Rust/WASM agentic OS for swarm and edge agents (ruos.cognitum.one)
  20. The Agent Said It Was Done. The Database Disagreed. (huggingface.co)
  21. [Editorial] polytoken.dev (polytoken.dev)
  22. [Editorial] YouTube: 0otdTJa9IUc (youtube.com)
  23. [Editorial] YouTube: MWX36ZYnsm0 (youtube.com)
  24. OpenTPU โ€“ An open-source AI accelerator, developed by AI (github.com)