# Open-Weights Offense Goes to Work

Published: 2026-10-05
Canonical: https://agidreams.us/edition/open-weights-offense-goes-to-work
Content-Complete: true

<!-- SECTION: 🔓 Open-Weights Offense Goes to Work -->

Cantina Security's first open-weights security model arrives with an argument attached, and the argument is the durable half. Cybersecurity is adversarial: attackers do not respect acceptable-use policies or procurement rules, and they can run open models locally, strip safeguards, and build around capabilities that already exist — an asymmetric safety in which defenders are restricted from what attackers can obtain anyway. Cantina traces the precedent to the late-1990s disclosure fights, when Nmap and Metasploit were controversial until reproducible exploits became basic defender infrastructure — the takeaway never being that dual-use is harmless, but that restricting defenders does not make the offense disappear. Machine-speed attacks, the argument concludes, require machine-speed defense (more: https://www.cantina.security/apex-flash).

The engineering deserves more attention than the philosophy. apex-flash-1 is a reinforcement learning post-train of GLM-5.3-Flash, tuned with GRPO — Group Relative Policy Optimization, scoring attempts against each other rather than a fixed baseline. Training environments came from Cantina's own findings: 150 tasks from 50 vulnerability cases in three views — full source with guidance, full source without, and blackbox access to a running target. The vulnerability mix mirrors production work: 72 percent authorization, identity, and scope binding. Training used rank-256 LoRA across all experts and routers plus full-parameter updates on 16 activation-selected experts; updating everything collapsed the model — misattributed gradients from a small RL dataset, the team suspects.

The economics are the actual headline. Across a 60-task held-out run from 20 unseen cases, apex-flash-1 scored 66.7 percent adjudicated first-draw pass@1 at an estimated $2.38, versus 60.0 percent at $4.56 for the GLM-5.3-Flash base and 71.7 percent at $74.68 for Claude Opus 5 High — within five points of the frontier at roughly three percent of the price. The model card is careful: text-based security tasks only, image and video untested, the checkpoint a focused worker under a larger orchestrator, recommended with the Codex harness. An abliterated variant — a derivative with refusal behavior modified — ships alongside, with the evaluation applying to the standard checkpoint (more: https://huggingface.co/cantina-security/apex-flash-1).

Anthropic used CyScenarioBench — Irregular's private benchmark for multi-stage offensive cyber operations, built from real incidents as attack trees in containerized network topologies — for Claude Opus 5.5. With cyber mitigations disabled, Opus 5.5 averaged 67.6 percent across a ten-challenge subset, ahead of Claude Mythos 5.1 at 61.7, Claude Opus 5 at 53.0, and Claude Sonnet 5 at under one percent. Irregular states plainly what this measures: capability elicitation, not real-world attack efficacy (more: https://www.irregular.com/research/assessing-claude-opus-5.5-against-offensive-security-benchmarks). The resale market is not waiting: Audn.ai sells one OpenAI-compatible API for four uncensored models, including unrestricted GLM 5.3 for $999 a month aimed at 100 cybersecurity researchers, a resold offensive-cyber model with 1M context, and a Qwen3.8 fine-tune with zero refusals on the 519-prompt benchmark — rates published "as bounds, not ground truth" (more: https://platform.audn.ai).

<!-- SECTION: 🏦 Money and Matter -->

While the industry argues over who should hold offensive capability, somebody is using theirs. A wave of cyberattacks believed to have used AI agents hit Korean financial institutions: Shinhan Bank, KB Kookmin Bank, Hana Bank, BNK Busan Bank, Yegaram Savings Bank, Welcome Savings Bank, and Hyundai Capital leaked customer information, while the Korean Federation of Community Credit Cooperatives was attacked without a leak. The core networks handling deposits, loans, and transfers are separated from the external internet, and the attackers did not try them — they targeted the weaker links connected to the outside. President Lee Jae Myung ordered a thorough investigation; regulators demanded the highest vigilance. The warning is the right one: network separation alone is no longer enough, every external connection is a candidate entry point, and secondary damage — voice phishing with leaked personal data — cannot be ruled out. Established versus asserted: the leaks are documented, the AI attribution is so far a belief, and whether securities firms, insurers, and card companies were also hit remains unconfirmed (more: https://www.koreajoongangdaily.com/opinion/ai-hacking-spree-exposes-cracks-in-financial-sector-security/12904304).

The physical version came through anonymous officials. FBI and Coast Guard investigators have found evidence that hackers accessed the propulsion system of an oil supertanker approaching the Texas coast this summer — outsiders holding temporary access to the ship's digital systems, duration and scope of control still unclear. The breach underscores the urgency that led US authorities to board the VL Prosperity in August — fully laden, longer than the Chrysler Building is tall, bound for Galveston. Attribution and mechanism remain under examination; the officials asked not to be identified, and the boarding is the corroborated part (more: https://www.bloomberg.com/news/articles/2026-10-02/hackers-breached-propulsion-system-of-us-bound-oil-tanker).

<!-- SECTION: 🧮 Jev Meets Independent Measurement -->

Defenders chasing cheap, fast classification have spent weeks hearing about Jev, TypeSafe AI's non-generative decision model — no text output, just typed answers with calibrated probabilities — and the independent numbers cut in both directions. Red Hat's AI Safety team benchmarked nine guardrail candidates on prompt-injection detection and content safety, using class-balanced datasets in NeMo Guardrails. The conclusion is deflationary: the team "did not find that decision models produced faster, cheaper, or higher-quality answers compared with LLM-as-a-judge." Pre-trained CPU-scale classifiers under 200 million parameters — OpenShift AI's default guardrails — came within 0.20 points of the injection lead while topping the latency table. Jev did outperform the pre-trained content-safety model on accuracy, but Nemotron-3.5-Content-Safety trailed it by 1.13 points at significantly lower median latency, and a Qwen3.6-35B judge topped the injection leaderboard while beating Jev on latency universally. Shieldstral, a safety fine-tune, performed surprisingly poorly against its own benchmarks, and prompts tuned for Jev did not transfer to Laya (more: https://developers.redhat.com/articles/2026/10/02/benchmarking-ai-decision-models-against-traditional-guardrails).

arXiv 2609.37647, from the University of Bonn, the Lamarr Institute, and Fraunhofer IAIS, ran Jev zero-shot across 37 public datasets and seven task families — 346,009 requests, $9.15 total, 0.36-second mean latency — with Qwen3.8-27B and Gemma-4-E4B as controls. Jev beat Qwen on 27 of 37 datasets and Gemma on all 37. It excels at gist-level short-text decisions: IMDB 96.5, SST-2 96.4, ARC 98.8, language identification 99.6, CLINC150 routing 89.5. Its probabilities are genuinely calibrated — pooled expected calibration error of 0.028 against Gemma's 0.208 — and confidence supports selective prediction, lifting Banking77 from 79.7 to 96.3 at 50 percent coverage. The weak spots mirror the vendor's own list: fine-grained labels, low-resource languages (Nigerian Fulfulde 37.4), legal judgment (AGB-DE F1 0.204 — "a legal judgment that requires domain knowledge rather than a snap decision"). Then the anomaly: Jev scores better on MMLU's calculation-heavy subjects (94.3) than its others (91.3) — the opposite of both reference models. Rotating answer options changes nothing; withholding the question drops accuracy to near chance (31.5), ruling out shallow memorization but not memorized question-answer pairs. The authors treat long-established English benchmarks "with caution," with training exposure the most plausible explanation (more: https://arxiv.org/pdf/2609.37647).

The reproduction layer is arriving: apolinario/decision-index rebuilds the live leaderboard from pinned, sha256-verified sources, runnable with checkpoint and resume. Decision Index 0.2.1 spans 38 benchmarks and roughly 120,340 requests, with 67 entrants — Jev at 57.89, Rune 26B-A4B v3 at 57.44 — engines over 1,000 ms median latency excluded as no longer "Jev-like," no truncation, no option filtering, unanswered counts as wrong (more: https://github.com/apolinario/decision-index). multimodalart's Hugging Face Space hosts the board (more: https://huggingface.co/spaces/multimodalart/jev-decision-index).

<!-- SECTION: 🚨 The Model Is Not the Control Plane -->

Cirrius Tech's analysis of what happens when a decision model sits inside a security pipeline is the sharpest security writing of the batch. JEV looks ideal for security automation: give it evidence, ask a question, get a typed answer, cheaply and quickly. Christophe Parisel's "Covert Order" experiment shows the problem. The same nine sentences, reordered with meaning deliberately balanced between two choices, moved one class's probability from 0.21 to 0.46 — a 25-point swing — while confidence did not necessarily fall. Label assignment is worse: swapping which class maps to the API's arbitrary "A" label changed outcomes to the point that one class was never selected in 200 calls. Not noise — exploitable decision structure. Threat intelligence is hostile input by construction — headers, URLs, DNS names, file names, package metadata — so an attacker needs control of only enough of it to influence the decision (more: https://cirriustech.co.uk/blog/high-confidence-is-not-a-security-control).

The concrete case is alert triage: Vega's proof of concept put Jev as a cheap gate in front of a full triage agent, closing any alert it marked "not escalated" above a confidence threshold. Across two tenants it closed 15 percent of alerts on the busier one and 33 percent on the quieter one at a 0.8 threshold, with judged correctness around 98-99 percent — and confirmed escalations inside the closed set: one on the busy tenant, two on the quiet one. At 0.9 nothing slipped through, but coverage shrank; Vega is explicit that Jev is a one-sided gate, not a replacement for the agent's verdict. Cirrius's response is architectural: high confidence is not authorization, structured output is not a security boundary, and the metric that matters is P(correct | adversarially constructed input) — mutate only attacker-controlled parts of the evidence, preserve the malicious behavior, and measure how far the decision and its confidence move. The moment a probabilistic judgment determines whether investigation continues, the model has joined the control plane and needs a control plane's threat model.

Cotool is the same direction as a product: detection agents hunting against a natural-language threat model, automated response from any alert source, humans in the loop "at any step" as a vendor promise, threat-intel filtered to environment relevance, an evaluation harness measuring every agent run. Every feature assumes the boundary Cirrius is drawing — AI systems produce evidence, security controls enforce policy, and those are not the same thing (more: https://www.cotool.ai).

<!-- SECTION: 🐧 A CVE Flood Meets a Patching Ceiling -->

Debian's DSA-6528-1 fixed 1,313 CVEs in one update of the linux package (kernel 6.12.111-1 for stable): three from 2024, fifteen from 2025, roughly 1,295 from 2026, the bulk in near-contiguous blocks — the kernel's CVE Numbering Authority assigning batches wholesale. No per-CVSS severities, no subsystem breakdown, no exploitation statement — just the aggregate warning of privilege escalation, denial of service, or information leaks on an OpenPGP-signed advisory, the strongest verification anywhere in it (more: https://lwn.net/Articles/1097401/). At the micro end sits a ten-line torvalds/linux commit masking bits in ASID, the address-space identifier AMD virtualization uses — the quiet unit at which a 1,300-CVE flood actually gets fixed (more: https://github.com/torvalds/linux/commit/60d93c27859fb3ec5b5a689e24de303f166477d0).

PatchBench (arXiv 2609.04075) asks whether AI agents can absorb that workload; the answer is not yet, and the measurements flattering until now because they reward two failure modes — patch memorization (reproducing historical developer fixes) and symptom-guarding (suppressing the reported crash without fixing the root cause). Its DiffBLEU metric (a diff-aware similarity score) finds that at a 0.75 memorization threshold, 25 percent of agent patches on an existing 300-task benchmark are highly similar to historical developer fixes — roughly double the 11 percent rate for standalone LLMs — and agent harnesses amplify it, with GPT-5.6 Sol rising from 8.3 to 22.0 percent memorization under Codex. Worse, 81 percent of Codex patches modified crash-stack functions even when the real fix lies elsewhere; in CVE-2022-1276 the agent guarded the VM crash site instead of fixing the compiler-side root cause, so "the patched program can still silently generate malformed outputs." PatchBench itself is 213 tasks across 16 CWEs and 32 real projects, built so ground-truth fixes lie outside the crash stack, with vulnerabilities transplanted into newer repository versions and patch sites mutated so memorized fixes cannot be reused verbatim (more: https://arxiv.org/abs/2609.04075v1).

The results sting: PoC-only validation inflates solve rates 1.83x on average — 83.1 percent of PoCs pass while 45.3 percent of tasks are actually solved. The top three agents pass over 97 percent of PoCs but solve 59.2, 58.2, and 56.8 percent; the top AIxCC systems — Atlantis, Buttercup, RoboDuck — underperform general-purpose agents on identical models, hindered by rigid frameworks. Semantic validation is the bottleneck at 63.4 percent average, larger budgets plateau ($25 caps add at most 2.3 points), 67 tasks were solved by no agent, and 6.8 to 7.1 percent of passing patches still leave root causes partially unfixed. The human alternative is visible where stakes are highest: Yukon's HashSmash tracks the Keccak[1600]/SHA3-256 collision frontier, where only human-promoted results form the published frontier — "the nominal reference is not a qualified attack or a proved security bound" (more: https://www.yukon.org/hashsmash).

<!-- SECTION: 💻 Local AI Without a Data Center -->

Local inference keeps absorbing work that once required cloud budgets. backburner turns an iPhone into a co-processor for a Mac running local LLMs: a llama.cpp fork with custom kernels splits the job over a 10 Gb/s USB-C cable. Below 64k context the Mac runs the first 40 layers of each microbatch and streams the residual to the phone for layers 41-64; past 64k, KV pages — the attention cache of previous tokens' keys and values — migrate to the phone, buying 196k-229k tokens of context on an iPhone 17 Pro Max against 64k on a 24 GB Mac alone. Prefill runs 29 to 44 percent faster at 16k-48k context, per-token time at 140k drops from 279 to 176 ms, greedy output token-identical with or without the phone. The phone answers only the Mac over USB, with an OpenAI-compatible server at localhost (more: https://github.com/StayLameBro/backburner).

The browser is becoming the other local runtime: kernels for more than 200 common ML operations, open-sourced to run entirely in-browser on WebGPU — the browser API exposing the local GPU — with the optimizations being upstreamed into Transformers.js, ONNX Runtime Web, and LiteRT.js (more: https://old.reddit.com/r/LocalLLaMA/comments/1wu8tpg/we_just_opensourced_the_worlds_fastest_webgpu/).

Training cost is collapsing the same way. A hobbyist trained a 3.87B-parameter MoE — 1.45B active per token, every layer mixture-of-experts, 16 experts with top-4 routing, a Qwen3 tokenizer, 4,096-token context — from scratch on 86.5B tokens across one-to-two GH200s using DiLoCo, a distributed training method that synchronizes infrequently. HumanEval+ of 41.5 at roughly 0.087T tokens matches Qwen2.5-1.5B, which used 18T — a startling data-efficiency claim on code — while MMLU at 28.6 is near chance and LiveCodeBench medium/hard is near zero, limitations the author lists honestly. DPO with 220k pairs made answers longer and hurt code and math, so the checkpoint was dropped (more: https://old.reddit.com/r/LocalLLaMA/comments/1wxiy8y/i_trained_a_387b_moe_145b_active_from_scratch_on/). At the recreational end of the same hardware, a local LLM radio on a couple of DGX Sparks, a 5090, and a 4070 Ti runs three local models generating music, voices, and sound effects while a radio host comments on the songs — "a lot more farm than pilgrim" (more: https://old.reddit.com/r/LocalLLaMA/comments/1wvg5hn/i_built_a_local_llm_radio_like_the_pharaohs_built/).

<!-- SECTION: 🗝️ Infrastructure for One Agent and One Person -->

Okta and Microsoft are building agent identity for corporate fleets; nobody was building it for one person with one agent and one phone, so a hobbyist built Auth Your Agent. One rule: the agent never sees a password, and anything that changes something waits for a thumb. Sites that integrate it get a "Sign in with your agent" flow: the agent signs with its own key, the human approves with a phone passkey, and the site issues short-lived agent-scoped tokens. Every other site gets the vault: a sandboxed Chromium in Docker on the user's own machine, where a login, CAPTCHA, or 2FA prompt hands a live view to the phone, and every click that would write something is held until Approve or Deny. It ships as an MCP (Model Context Protocol) server, so the model behind it is irrelevant; the Reddit post describing it was written by the agent itself, typed through the vault with the Post button waiting for a thumb, and disclosed as such (more: https://old.reddit.com/r/ollama/comments/1wvn4rv/my_agent_could_do_everything_except_get_past_a/).

VoxLux gives people and AI agents private rooms that no external entity operates. There's no account, no central directory, and no messaging server. Each person, and each agent session in Claude Code, Codex or OpenCode, becomes a node with its own keys. Nodes join end-to-end-encrypted rooms as peers. Being in a room lets you see that messages exist, but you can read someone only once you each choose to trust the other. The same rooms carry more than chat. A node can offer a service such as SSH, a file share or a NAS, and the nodes it trusts can reach it, even when it has no public address. The only infrastructure is an anchor you run yourself, and it's needed only when two machines are both behind NAT and can't find each other directly. The anchor holds no keys and can read nothing. Vox doesn't start, contain or supervise agents. It carries their messages and tunnels, so agents and the people they work with can split up work, hand it off, and stay in one private room wherever each of them runs. Vox is MIT-licensed and runs on macOS and Linux. (more: https://voxlux.us).

REAmon brings the same instinct to reverse engineering: a Docker-orchestrated workspace with a Next.js/Prisma front, PostgreSQL for targets and task state, and Neo4j for a knowledge graph connecting evidence, findings, and hypotheses. Workspaces hold files, repositories, live processes, devices, and debug sessions under bounded defaults — a 512 MiB per-file limit, 60-second windows — and the release gate runs Trivy image scans, browser end-to-end tests, database restore drills, and two-replica crash recovery (more: https://github.com/xdCloudy/REAmon.git).

Wallet-Risk-Scanner applies it to AML screening: a 0-100 score from sanctions exposure (capped at 40), high-risk contract interaction (30), and fund-source tracing (30), aggregated from six providers (GoPlus, Etherscan V2, ChainAbuse, MistTrack, BlockSec) across 15-plus chains, with multi-hop tracing, mixer and honeypot detection, batch screening, and 24-hour local blacklist caching. Scoring runs entirely locally, and the disclaimer is the correct one: risk scores are heuristic signals, not legal determinations (more: https://github.com/Soniavasseur/Wallet-Risk-Scanner).

The provenance version of the same concern is Reuven Cohen's warning about feeding unpublished research to ChatGPT and Claude. OpenAI says eligible consumer ChatGPT and Codex content may improve its models; Anthropic says consumer Claude chats may train where users permit it. The 2025 Navier-Stokes Millennium Prize controversy made the question concrete — OpenAI investigated and says it found no evidence that mathematicians' Codex interactions influenced the result, which, as Cohen puts it, "is not proof of misuse. It is proof that provenance questions can arise after disclosure, and may be difficult to resolve." His timeline makes the point without proving copying: his Q Star work on multi-agent reinforcement learning was public in November 2023, and OpenAI filed 2024 patent families covering multi-agent interaction, shared workspaces, and agent coordination. "The timing does not prove copying" — it shows why timestamps, logs, and prior-art records matter. The advice is unglamorous: run private models on infrastructure you control, encrypt, preserve logs, timestamp discoveries, and file before disclosure (more: https://www.linkedin.com/posts/reuvencohen_use-chatgpt-and-claude-for-unpublished-research-share-7512861619522064386-wJye)
