# The Open-Weight Offense Fight

Published: 2026-10-01
Canonical: https://agidreams.us/edition/the-open-weight-offense-fight
Content-Complete: true

<!-- SECTION: 🎯 The Open-Weight Offense Fight -->

Anthropic's warning that a downloadable Chinese model can now build working hacks on its own reached the comment section — which is not impressed. The finding underneath is concrete: GLM-5.3 is, per Anthropic's red-team write-up, the second model on record to autonomously build end-to-end exploits — one anyone can download. On a benchmark targeting the Chrome V8 JavaScript engine it completed working exploits in 50 of 410 attempts against Claude Mythos Preview's 56; the cheaper Flash sibling turned a patched flaw, CVE-2026-11645, into a reliable arm64 chain that bypasses pointer-authentication hardening — in twenty minutes, for an estimated $20–40, against eight hours of human work.

Reddit's consensus held the warning self-serving — "McDonalds trying to outlaw people making burgers in their backyards" — plus hypocritical: Anthropic sells Claude to US national-security customers while suggesting citizens shouldn't run a competitor's open weights. The minority's counter: the thread pegs it at 753 billion parameters, beyond any gaming rig — but nation-states and crime groups are not consumers, and the safeguards are cheap to strip. A fake red-team cover story got it to engage malicious requests 64% of the time, prefilled thinking tokens 92%, an "abliterated" copy — refusals stripped — 100%. The numbers are published; the motives, on either side, are not — and only independent replication would settle what the thread argued over. (more: https://reddit.com/r/ClaudeAI/comments/1wtk9kd/anthropic_says_a_chinese_ai_model_anyone_can/)

The same week, the other party institutionalized the autonomy it spent the spring contesting. Defense Secretary Pete Hegseth announced a four-star-led "Autonomous Warfare Command" — Autowarcom — the first new combatant command since Space Command in 2019. The rationale is budget physics: "cheap compute, superintelligence, and advanced commercial manufacturing have enabled the proliferation of low-cost, high-precision strike" — Ukraine is the proof: drones account for an estimated 70% of Russian casualties. The irony has a paper trail: this year Anthropic asked the Pentagon for a guarantee of no autonomous weapons without human oversight and was refused; a federal judge later blocked the department's attempt to blacklist the lab's models even as the NSA ran them. Autowarcom is what losing that argument looks like in org-chart form. (more: https://www.reuters.com/world/pentagon-creates-autowarcom-expand-ai-drone-capabilities-2026-09-30)

<!-- SECTION: ⚔️ Measuring What Security Agents Can Do -->

The offense debate's structural flaw: both sides argue from lab-owned numbers. This week produced two alternatives — a benchmark, and the apparatus to run one.

SecRespond, from HKUST and Alibaba Cloud, is the first benchmark for post-compromise incident response; most agentic-security evaluations stop at clean, pre-compromise systems. The agent gets a frozen forensic snapshot of a genuinely compromised cloud host plus the host security product's analytics, and must deliver an intrusion report reconstructing the attack chain, a vulnerability report, a baseline assessment, and a remediation plan. Its ten cyber ranges come from end-to-end attacks over real network protocols — SSH brute-force mining, Log4j RCE, a Docker escape, an npm worm, RDP spraying — spanning 21 MITRE ATT&CK techniques across five operating systems, selected from 372 real compromised hosts. A realism detail: only 30–40% of attack actions trigger an alert. Scoring runs on 280 expert-designed checkpoints (human–judge agreement Pearson 0.96). (more: https://arxiv.org/abs/2607.26791v1)

The leaderboard flatters no one. Claude Opus 4.7 leads the 23-model field at 79.0% detection and 65.7% planning (72.4% averaged), with GLM-5.1 and Qwen3.7 Plus ahead of GPT-5.5 (70.7% detection, 36% planning). Detection beats remediation planning on every range, "largely attributed to incomplete fixes instead of wrong actions" — agents apply the obvious fix, skipping credential rotation, outbound blocking, and post-cleanup verification. Persistence detection collapses to 58.8% versus 75.4% for intrusion entities — persistence lives in cron jobs, systemd units, and shell-init files that never trip an alert. No model completed full detection and remediation on any single range; Claude Opus 4.7 refused all NPM-Worm attempts for safety. Agents still beat an agentless production scanner (2.1% on persistence — it cannot attribute entry points).

CyberPVP is the smaller, structurally interesting piece: live dual-model matches on the CyberGym benchmark — CyberKimi and CyberGLM against Altar-1, Aikido's open-weight security model, on 8×H200s. The fairness protocol is what the debate keeps missing: a manifest written before the run starts (tasks, endpoints, budgets, seed), both agents on the same task IDs, sanitized traces with SHA-256 checksums, scoring only by the upstream CyberGym proof-of-concept verifier. Pre-registration plus independent scoring, the standard a clinical trial would face. A hacking claim that can't survive it is marketing. (more: https://github.com/lordx64/cyberkimi-pvp)

<!-- SECTION: 🎭 Spoofing the Trust Anchor -->

Identity is the perimeter, and the perimeter is parser code.

SEC Consult's Timo Longin, earlier this year, found two "exotic" flaws in Apple's iCloud mail pipeline that let a plain iCloud account send mail as any @icloud.com address — demonstrated with tim.cook@icloud.com. The mechanism is a parser differential: Apple authenticates with one internal parser but normalizes with a second, and the two disagree about what a From: header is. A bare carriage return around the header's colon makes the authenticating parser ignore the header while the normalizer lets the forged one survive. The forged mail passes SPF, DKIM, and DMARC, because Apple signs after the second parser runs — "a legitimate and authenticated email," in the post's words. Disclosure took a year and a half of "still under investigation"; the first fix merely blacklisted the PoC substring "admin," leaving every other address spoofable; the correct fix, plus a $15,000 bounty, closed it. The durable conclusion: "sender identity in emailing will most likely remain far less solid than most users assume." (more: https://sec-consult.com/blog/detail/from-anyoneicloudcom-spoofing-arbitrary-apple-icloud-identities)

Proofpoint's report on TA419 — a China-aligned espionage actor active since April 2025 against US and Japanese think tanks, defense contractors, and law firms — runs the same attack on people, not parsers. In July 2026 the group impersonated Lynne Edwards Parker, former Principal Deputy Director of the White House Office of Science and Technology Policy, and economist Heidi Crebo-Redicker, offering seats on a fictitious "AI Policy Advisory Committee" or a role in a Senate AI-export report. A reply earned a shortened URL, then an adversary-in-the-middle page relaying the genuine Microsoft 365 sign-in live. The kit — an open-source Browser-in-the-Browser build with an Evilginx phishlet for Microsoft 365 — auto-accepts "Keep me signed in" and auto-submits one-time codes the moment they validate, so password, MFA, and conditional-access checks all succeed while the attacker harvests session cookies. In February the group impersonated a senior Anthropic employee — subject line "Request for Feedback on Military Integration of Claude" — against an AI policy analyst: the targeting tracks the policy fight, not the technology. Proofpoint's fix: phishing-resistant, origin-bound passkeys, and out-of-band verification of unsolicited outreach. (more: https://www.proofpoint.com/us/blog/threat-insight/hallucinating-credibility-china-aligned-ta419-impersonates-its-way-us-ai-policy)

<!-- SECTION: 🪱 Isolation Is Only as Solid as the Code -->

One level down, today's best technical read: calif.io's teardown of CVE-2026-86950, the CoreGraphics bug Apple fixed in iOS 26.7.1, credited to Meta Product Security — with Apple's note that it "may have been exploited in an extremely sophisticated attack against specific targeted individuals": spyware in the wild. An instruction-level diff found one inline function patched twenty-plus times in the anti-aliased path rasterizer — the routine converting glyph coordinates to fixed-point on a 4096×4096 sub-pixel grid. The root cause: converting a double to an int outside range is undefined behavior, and clang compiled two similar conversion sites differently — one saturating, one 64-bit-then-truncate. The differential corrupts the glyph path's bounding box; the coverage buffer (the per-scanline alpha mask) is sized from the corrupted box, and the rasterizer writes out of bounds. For delivery, the authors point at WhatsApp, whose "Kaleidoscope" framework had just gained a strict PDF mode emitting font-defect tags — strongly suggesting a malicious PDF embedding a crafted font. Their proof-of-concept, an oversized-coordinate TrueType font, crashes macOS and iOS — and they credit AI agents with navigating CoreGraphics' control flow: "our nightmare is the friendly AI agent's dream." (more: https://calif.io/research/the-great-glyph-grift)

A security channel made the same argument about a bigger boundary: containers. They aren't VMs — clone plus namespaces gives a fake root while every container shares the host kernel — and the economics have moved: kernel CVE counts are rising nearly exponentially, a pace attributed to AI-assisted vulnerability research. Dirty Cow-caliber bugs (2016) used to arrive every year or two, now every couple of months — CopyFail and DirtyFrag are "literally two of like a dozen." CopyFail writes four controlled bytes into the page cache of any readable file — overwrite /usr/bin/su, run binaries as root — and DirtyFrag does the same via /etc/passwd. The proposed answer is micro VMs: KVM gives each workload its own address space, so escape requires a hypervisor bug — smaller surface, though not zero: Google's KVM CTF pays $250,000 per escape. The closing line: containers "are not virtual machines, despite popular opinion." (more: https://www.youtube.com/watch?v=9-2CSJC114k)

Even proof systems are in it. A roughly 1.2-million-line, AI-generated Lean proof — claimed accepted by both the official kernel and Nanoda, a Rust external kernel — chained exploits of two bugs, one in each kernel, including "a bunch of weird terms to create a hash collision that was completely irrelevant," aimed at a Nanoda bug fixed ten days earlier. Lean creator Leo de Moura's team "strongly believe[s] it was built by AI." His response is non-obfuscationist: a Comparator rechecking exported proofs in a sandbox, and proving the kernel correct in "Lean for Lean" — which had verified many parts, "not the parts that had the exploits." On hiding the kernel instead: "I don't like safety by obfuscation... it goes against everything our community believes, safety by secrecy." (more: https://www.youtube.com/watch?v=ZpQFebTK75A)

<!-- SECTION: 📉 When the Numbers Say No -->

The quieter failure mode: documented numbers contradicting institutional momentum, and the momentum winning.

A freedom-of-information response obtained by Liberty Investigates shows the British Transport Police's six-month live facial recognition trial in London's busiest rail hubs — 18 deployments, February through July, more than half a million faces scanned — cost £320,786 plus roughly 100 hours of officer time, and produced exactly one watchlist alert — a false positive. No arrests followed from the technology. The police response was to extend the trial into London Underground stations — three confirmed alerts since, all people complying with court conditions. Fraser Sampson — former biometrics and surveillance camera commissioner, now a non-executive director of Facewatch, a shop-facial-recognition operator — supplies the distinction that undoes his own defense: success in a shop means unwanted people don't come in; success for police means wanted people get caught — and "in that respect the trial doesn't appear to have been very fruitful." More than half of England and Wales' police forces run live facial recognition on the streets; parliament's joint committee on human rights calls the rollout a "particularly clear example of risk." BTP deserves one credit: arrests made during deployments are excluded from LFR data because they did not result from LFR alerts. Honest accounting — and the whole story: the only number attributable to the technology was wrong. (more: https://www.theguardian.com/technology/2026/sep/29/trial-live-facial-recognition-cameras-london-stations-false-positive)

The WeWork saga, retold by its founder this week, is the same pattern with better production values. The S-1 valued the company near $47 billion while it lost $690 million in six months — about $219,000 every hour. SoftBank's Masa Son had offered $20 billion in cash; the board pushed toward $32 billion and lost the entire deal — Neumann's own verdict: "if you're greedy, you even lose what you have in your hand." He now runs Flow, a residential real-estate venture backed by $470 million from a16z, his family personally investing $350 million — the second act got funded anyway, worth noticing in any valuation cycle. His advice survives translation: choose partners "for the worst day in your life, not the best." (more: https://www.youtube.com/watch?v=IQ4JVWdj4Q0)

<!-- SECTION: 🚀 The Frontier Release Cycle -->

The release wheel spins on schedule — major launches arrive roughly every 18 days by one reviewer's count, who grades on cost per task, not vibes. His showcase for Opus 5.5: a beanie logo converted into a 514-piece Lego model with a 63-page, 58-step booklet, a parts list, a BrickLink wanted list, and a checker file validating the geometry — 89 million tokens, about 1% of his weekly allowance, roughly fifty cents of subscription value against $44 at API pricing. At $4/$20 per million input/output tokens — 20% below Opus 5, 60% below Fable 5.1 — his caution is one most coverage skips: a price cut is not a token-efficiency gain, and input, output, cache, and context buckets must be tracked separately before anyone confirms Anthropic's 40%-cheaper-workloads claim. On steerability, 5.5 is a return to form after the 4.7–4.8–Fable arc — frustration peaked with Bram Cohen's "Why is Claude turning into an asshole?" The test for good revision: it "preserves your intent" — simplifying a paragraph shouldn't delete the complication that made it worth reading, the cleanest line between AI slop and useful writing. An 18-hour unattended run across six repos succeeded, with the lesson: define stop conditions first, because the model "loves to go and go and go." (more: https://www.youtube.com/watch?v=osZZjdMZVvA)

Google's counter-move, Gemini 4 Argon, arrives through a DeepMind "Fairwind" program for cybersecurity partners and trusted testers: $2/$10 per million intro pricing (rising to $4/$20), cached input 95% cheaper, output up from 64K to 1M tokens. Per the channel's coverage — sourced from leaked demos, so treat accordingly — Argon scores 53 on the Artificial Analysis Intelligence Index, matching GPT-6 Astra; about $1.99 per task against Astra's $3.26 and Opus 5.5's $5.98; and a 15% hallucination rate, the lowest ever measured for a model scoring above 45. Reuters reports Argon is larger than previous Gemini 4 models — and the channel's "perhaps this is a 6 trillion parameter model" comes with its own disclaimer: "we simply don't have evidence to make that claim." (more: https://www.youtube.com/watch?v=RMTWayrwYkQ)

A smaller demonstration of the older generation's craft: a Claude Opus 5 user had the model generate every vector glyph of an animated music score from scratch, via a script transposing any MusicXML or MuseScore file into the renderer's format. (more: https://reddit.com/r/ClaudeAI/comments/1wp74zm/customizable_animated_music_score_opus_5/)

<!-- SECTION: 💾 Open Weights and the Hardware to Run Them -->

Below the frontier, the open-weight story has become a distillation economy. NVIDIA's Nemotron-Labs-3-Competitive-Coding-550B-A55B, shipped in NVFP4 (NVIDIA's 4-bit float format), is the cleanest specimen: one epoch of supervised fine-tuning from Nemotron-3-Ultra on 477,642 synthetic reasoning traces distilled from GLM-5.2, spanning 22,000 problems from 16 contest families. GLM-5.2 won the teacher slot over a DeepSeek-V4-Flash variant on accuracy and roughly 30% shorter generations. At inference, GenCorrect — a closed-loop test-time-compute strategy generating diverse candidates, taking evaluator feedback, refining under a fixed submission budget — carried it through the IOI 2026 problem set live, under official contest constraints: 535.4 out of 600, above the 361.12 gold-medal threshold and the top human's 498.27 — the first AI system reported to outscore the best human on an IOI problem set. Weights, data, and recipes are open; the community's read is about right — a research artifact, not a product pitch. The structural news: a frontier open-weight release now functions primarily as a distillation source — precisely as the local-model community predicted when GLM-5.2 landed. (more: https://reddit.com/r/LocalLLaMA/comments/1wsuqmb/nvidianvidianemotronlabs3competitivecoding550ba55b/)

Naive's N0.5-Flash — 309B total, 15.5B active, a 1M context, hybrid SWA/DSA attention, aimed at coding and AI R&D — drew commenters tracing its lineage to MiMo v2.5 and noting llama.cpp can already run it by letting the DSA layers fall back to dense attention. SenseNova's U1.5-8B-MoT also surfaced on Hugging Face's trending list. (more: https://reddit.com/r/LocalLLaMA/comments/1wrs58t/naiven05flash_309ba155b/) (more: https://huggingface.co/sensenova/SenseNova-U1.5-8B-MoT)

The hardware side is catching up. A community appreciation post for the pseudonymous tcclaviger documents vLLM's expert RAM offloading — still a fork, not mainline — running DeepSeek-V4-Flash-Vision-Exp on four R9700s with 160GB of expert-offload memory and dspark speculative decoding; the same flash class already ran at about 84 tokens per second on three RTX 3090s. Magnitude, a YC S25 inference engine under Apache 2.0, attacks the same problem from the compiler side, tuning kernels on the actual device before a model runs: claims of up to 2x faster than llama.cpp (92% faster decode on Metal, 19% on CUDA) and 27% less memory per agent. The standing caution: verify configuration parity first — the last widely shared 18% advantage claim was a configuration gap. (more: https://reddit.com/r/LocalLLaMA/comments/1wtg12r/ram_offloading_with_vllm_tcclaviger_appreciation/) (more: https://github.com/magnitudedev/magnitude)

<!-- SECTION: 🔌 Plumbing Between Agents -->

The most interesting infrastructure release this week: a small Go project with an MIT license. agent-tincan is a relay that lets personal AI agents ask each other for help across radically different environments: an always-on cloud VM, a pausing e2b sandbox, a chat app, a Mac terminal. The design is deliberately boring: one relay on one always-on machine in your Tailscale network, every agent dialing out — no open ports. One agent asks another for help; the relay queues the request, wakes the target "the way that agent wakes best" — webhooks, email, long-polls, or a command listener launching a headless run — and carries the reply back. Identity comes from Tailscale's WhoIs, "so nothing in a request can change who it is from," and there are no API keys between agents. Tools are exposed via MCP (Model Context Protocol) over stdio; ChatGPT joins as an OAuth-protected MCP endpoint. The Council feature runs Andrej Karpathy's llm-council over the relay, and an append-only, hash-chained audit log keeps everything inspectable. The trust model, stated without varnish: "Tailscale is the security boundary, the relay can read every request and reply." (more: https://github.com/mvanhorn/agent-tincan)

The decision layer underneath is getting cheaper. Jev, the System-1 decision model — supply a state plus multiple-choice questions, get parallel answers with probabilities, at a fraction of LLM latency and cost — is deployed where LLM judges don't pay: pre-tool-use guardrails (an LLM judge costs a tenth of a penny and over a second per call; Jev, a quarter-second and "a fraction of a fraction of a penny," almost no false positives), plus workflow routing — 12/12 on issue classification, 15/16 on PR-review depth, against 32 cents per PR with an LLM. The pattern: "sandwich Jev in between calls to a large language model." Caveats: early access, no published weights, open reimplementations within days. (more: https://www.youtube.com/watch?v=qwnJJMNGwgY)

Zer0Fit wraps two Google research models — TabFM for zero-shot tabular classification and regression, TimesFM for forecasting — in a dockerized MCP: hand it a CSV and the connected LLM routes the task. The chatter around these MCP projects has spread past Reddit and GitHub onto cooperatively governed Mastodon servers, too. (more: https://neuromatch.social/@jonny/117364515040550815) The author's framing is the right one: "I'm not doing anything special, I'm just wrapping the Google models," and "don't use this with anything where its output matters." One commenter's private-dataset test found it worse at regression than random forest and XGBoost — expected for zero-shot against a fitted model. (more: https://reddit.com/r/LocalLLaMA/comments/1wrke2t/zer0fit_zeroshot_predictions_classifications_and/)
