ARTEX and the Open-Weight Intrusion Stack
Published on
Today's AI news: ARTEX and the Open-Weight Intrusion Stack, Measuring the Risk: Control Planes and Recoverable Leakage, Agents That Reverse Engineer, Experts on Disk: Local Inference Past the VRAM Wall, Harness Design: Codemode, Durability, and Intent, Generative Media Without the Usual Scaffolding, Water, Worship, and the Room You Show Up In. 25 sources curated from across the web.
ARTEX and the Open-Weight Intrusion Stack
CrowdStrike Intelligence has named the tooling behind the late-September breaches at South Korean financial institutions: ARTEX, a recently released, Chinese-developed open-source agentic penetration testing tool, driven by a stack of open-weight and commercial LLMs. The campaign ran from late September into early October 2026 and resulted in exfiltrated data. Industry reporting describes a loan progress inquiry service used by brokers compromised at one bank and an employee mobile work-support system at another; the number of affected organizations is unconfirmed. The attribution is modest: the actor is "likely a Chinese speaker and financially motivated," assessed "with moderate confidence based on the use of the Chinese-developed tool ARTEX and observed Chinese-language prompts" (more: https://www.crowdstrike.com/en-us/blog/unknown-threat-actor-uses-artex-to-target-south-korean-finance).
The investigation exists because the operator left the door open. CrowdStrike found attacker-controlled open directories holding Claude Code session histories, ARTEX configuration files, and Claude memory files. That is now a recurring motif in this year's AI-assisted intrusion write-ups: the forensic record comes from the attacker's own self-hosted sprawl, not victim telemetry. The sessions show a two-server layout, with a Hong Kong IP as primary infrastructure and a second host running ARTEX. ARTEX used DeepSeek v4.1-flash as its primary backend, likely reached through an API proxy or reseller, with GLM-5.3 from Zhipu and Grok 4.6 supplementing further Claude Code sessions. The DeepSeek model itself is published on Hugging Face (more: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash).
The softer side of the operation is in the same logs. The actor asked Claude where Korean breach data is sold, looked for Korean Telegram data-sales groups, and requested a security researcher resume describing the ARTEX results, supplying a date of birth and an education at South China University of Technology. A Telegram handle, YY520CN, links to sessions probing a Telegram NFT marketplace and a possible Chinese payment platform. CrowdStrike is careful: "currently available information cannot definitively associate these details with the threat actor." The intrusion was automated; the monetization research and the opsec were human and sloppy.
Which makes the LocalLLaMA victory lap over GLM 5.3 Flash topping the Artificial Analysis Cyber Index worth reading carefully. The post asserts two open models now sit above every Anthropic model, with no methodology attached. The comments supply the confound: several attribute the gap to Claude refusing cyber tasks, one notes Anthropic said it applied cyber guardrails to Opus 5.5, another says Z.ai models are "so easy to jailbreak." If an index scores completed offensive tasks, willingness counts as much as skill, and the ranking cannot separate them. A per-task refusal rate alongside the score would settle it. Until then, "open source prevails" is a claim about guardrails, not models (more: https://old.reddit.com/r/LocalLLaMA/comments/1x1rwof/glm_53_flash_opensource_the_top_of_artificial/).
Measuring the Risk: Control Planes and Recoverable Leakage
Expel's guide to AI security risk is a CISO document rather than a threat taxonomy, and declines to compete with NIST AI RMF or MITRE ATLAS. Its premise is that AI "is changing the economics of security risk more rapidly than many organizations can move" by breaking the assumptions the SaaS and MDM playbooks relied on, which it compresses into three shifts. Identity is degraded. Execution is probabilistic. Observability "is no longer a simple question of 'Do you have the logs?'" (more: https://expel.com/resource/guide-to-thinking-about-ai-security-risk).
Six control planes, data, identity, execution, human, external AI and service control, and observability, answer why each stakeholder is at the table; observability is primary because "without it, none of the other planes can be validated or enforced." Ten unprioritized risk domains serve as board reporting lines, from data exposure where "exposure occurs without traditional exfiltration" to identity degradation where "privilege boundaries are traversed through workflows, not logins," to observability collapse: "if activity cannot be reconstructed, it cannot be secured." A worked data-exposure example names owners and proposes metrics such as the share of governance-scoped data reachable by unauthorized AI endpoints. That is more concrete than most vendor frameworks, and it comes from a managed detection shop whose business is observability, which is worth remembering when the guide ranks observability first.
The measurement side arrives from Minjiang University, Renmin University, and RMIT. CIPL, Channel Inversion for Privacy Leakage, complains that agent privacy evaluations are organized by component, memory, RAG, or tools, which "makes it difficult to distinguish internal exposure from information that an external observer can actually recover." CIPL decomposes each target into six stages and reports whether selected units are recoverable from visible output, as any-unit and complete-recovery rates, under a black-box threat model with no privileged logs (more: https://arxiv.org/abs/2609.21686v1).
The results map onto Expel's observability plane. Memory targets saturate, with complete recovery of 1.0 across all five providers tested. RAG leaks frequently but partially. Tool and BrowserUse leakage is channel- and provider-dependent: DeepSeek and GPT-4o leak through execution traces while shielding final answers, whereas Qwen leaks through the answer itself. The sharpest finding is an appendix on an MCP, Model Context Protocol, agent behind the PrivacyInAction gate: monitored tool paths held complete recovery at 0.00 and any-unit recovery below 0.28, while an unmonitored console channel pushed any-unit recovery to 0.92. The gate worked where it looked. The limitations are real: results are tied to specific budgets and provider snapshots, the semantic audit is small, and the runtime judge agreed weakly with humans, kappa 0.36 and recall 0.30.
Agents That Reverse Engineer
ruvnet's rudevolution is a static decompiler for minified JavaScript bundles, and its README opens with a sentence rarely seen in that ecosystem: "Evidence first. Recovered modules and names are hypotheses." The Rust core parses declarations without executing them, builds a weighted reference graph, partitions it into candidate modules with Louvain community detection, infers names with confidence scores, and emits V3 source maps alongside SHA3-256 Merkle "witness chains" that attest to byte integrity and, the README stresses, do "not establish original intent, behavioral equivalence, or authorship." Six MCP tools expose it to Claude Code (more: https://github.com/ruvnet/rudevolution).
The showcase target is Claude Code itself. One historical run on an 11 MB cli.js with 27,477 declarations reportedly took about 26 seconds, yielding 1,029 modules and 25,465 inferred names, and decompilations of five Claude Code versions ship as release downloads. Every number carries a disclaimer: the 95.7% validation accuracy "has not been reproduced on a shared, held out benchmark," and the "Clean Room" workflow is labeled developer preview with the warning "Do not use this as evidence of lawful independent authorship." Given this author's record of shipping ambitious prototypes faster than independent verification can follow, the hedging is welcome, and the numbers deserve the treatment the README gives them: historical, single-run, unreproduced.
REA, "Reverse Engineer Anything," attacks from the other direction as an MCP server and CLI that lets a coding agent inspect native binaries, Electron apps, APKs, .NET assemblies, EVM bytecode, firmware, network captures, and live process behavior. The pitch is that an agent can see a feature in someone else's app, "understand how it works, down to the binary level," and rebuild it, with each finding returned alongside its evidence and limitations. The showcases are concrete: a DX-Ball sound-pan helper reconstructed into C that "passes 3,205 original-x86 cases and reproduces all 63 compiled function bytes," and Notion's clipboard bridge traced through Electron IPC. The project claims 40,000 stars, carries a lawful-use disclaimer, and states it has issued no cryptocurrency token. Claude Code is now both the reverse-engineering agent of choice and the most-studied target, and both projects stress that matching bytes are evidence of behavior, not permission to ship (more: https://github.com/morluto/rea).
Experts on Disk: Local Inference Past the VRAM Wall
am17an's llama.cpp pull request adds a GPU cache for mixture-of-experts weights that live in host memory, and a follow-up extending it to multiple GPUs has already merged. On a 9070 XT under Vulkan running Gemma 4 26B-A4B QAT, decode rose from 59.3 to 76.9 tokens per second with an 8 GB cache while prompt processing fell from 869.7 to 409.3. The consensus recipe is to combine it with the existing expert-offload flag and raise the cache until VRAM is nearly full, trading prefill for decode, the right trade for chat and the wrong one for document ingestion (more: https://old.reddit.com/r/LocalLLaMA/comments/1x03xkc/llama_add_a_gpu_cache_for_moe_experts_kept_in/).
LocalMind pushes the hierarchy one tier further, into a browser tab. A static page copies a GGUF into the browser's private file system, puts dense weights, routers, and KV cache on the GPU through WebGPU, and streams routed experts from disk through a worker pool into an LRU cache. Parity with llama.cpp is strict: on Gemma 4 26B-A4B at Q4_0, replies were character-identical to llama.cpp Metal on nine of nine conversations and token-identical on 15 of 16 fresh prompts, at roughly a third of native speed, 23.6 tokens per second versus 70.6. The headline is a 36.9 GB Qwen3.6 35B-A3B at Q8 running on a 24 GB MacBook at 9.9 tokens per second. The per-token split, about 23 ms GPU compute, 39 ms routing round trips, and 35 ms SSD reads, explains why moving routing onto the GPU bought nothing: "the misses are experts nobody predicted." Prefill at 55 tokens per second drew the thread's main complaint (more: https://old.reddit.com/r/LocalLLaMA/comments/1wyvo42/gemma_4_26ba4b_and_a_37_gb_qwen36_moe_running_in/).
The same contributor's RPC pull request adds tensor-parallel split mode across machines; the thread is excited and asks the right question about network latency (more: https://old.reddit.com/r/LocalLLaMA/comments/1wz5z7n/rpc_add_sm_tensor_by_am17an_pull_request_26610/). A hobby research post takes the disk tier to its end: a 21M-parameter model with a 16.8M-row learned lookup table, 6.4B parameters in the table but 33M touched per token, matches a 114M dense model trained on the same 500M Wikipedia tokens. With the 4-bit table memory-mapped from NVMe it generates about 140 tokens per second on an RX 9070 in 0.4 GB of VRAM, though every missed row costs a 4 KB page fault and long prompts crawl. The author lists what failed, including bolting a table onto a finished Qwen3.5-0.8B, and notes one seed, invented facts, and a 70-dollar Runpod bill (more: https://old.reddit.com/r/LocalLLaMA/comments/1wz7tvs/i_gave_a_21m_model_a_64bparameter_lookup_table_it/). Interfaze 1 Lite fits more into one box differently, as a "mixture-of-architectures" where a vision-language core picks perception specialists for documents, speech, and detection on a single 80 GB GPU; with no benchmarks on the card, it is a design statement, not a result (more: https://old.reddit.com/r/LocalLLaMA/comments/1wyp0y6/interfazeaiinterfaze1lite_hugging_face/).
Harness Design: Codemode, Durability, and Intent
Armin Ronacher spent a year arguing that agents should use shell and scripts rather than custom tools, so Pi 1.0 adding MCP support needs an explanation, and his Codemode essay provides one. Bash "can only compose programs that run." Reading an image or spawning a sub-agent must happen in the trusted harness, the "brain," not in the sandboxed "hands." Codemode is "a way for the LLM to express and orchestrate complex operations on the harness side": tool calls issued from JavaScript running in QuickJS inside WASM with no network, filesystem, or timers. "The only way is to call more tools." Large outputs get handled in code rather than truncated, and internal APIs can be exposed without burning context on tool definitions. Ronacher rates current MCP integration "not amazingly well," especially Codemode-inside-MCP-servers, which yields double JSON escaping. Durability, binary data, and small-model performance stay open (more: https://lucumr.pocoo.org/2026/10/6/codemode/).
Durability has an answer on Apple platforms. PiDurableKit wraps pi-durable in JavaScriptCore for iOS, macOS, and visionOS, committing conversations, model turns, tool calls, and app state to storage "before anything is shown." Two caveats matter: extension code "is not sandboxed, so install only code you trust," and OAuth sign-in to Claude Pro, ChatGPT, and Copilot subscriptions uses the providers' first-party clients, with advice to "check that this is allowed for your use before shipping it" (more: https://github.com/finnvoor/PiDurableKit).
Intent-Router sits upstream of execution as a prompt-only skill that compiles a vague request into a typed spec, deciding per unknown whether to probe the repository or ask the human. Evals are small and self-published, and the author's warning is the best line: confidence is "a model's self-report. Treat it as ordinal, not calibrated," and "Don't let it validate itself" (more: https://github.com/angel291592/Intent-Router). Arena takes the brute-force route: spawn up to 100 Claude Code sub-agents with distinct strategy cards, pair them to attack and defend, and let a blind judge advance survivors. The README concedes the judges are Claude too, so the winner is "not a proof that it is right," and "Vague task in, 100 flavours of vague out" (more: https://github.com/Jakeschincariol/arena-skill).
Three smaller entries address trust at the edges. Agentic Toolbox is a role-keyed skill library whose intake rules read like supply-chain policy: every third-party cherry-pick pinned by SHA, security review before vendor content lands, and "No auto-update. No marketplace install. No scripts. No secrets." That is the correct default (more: https://github.com/alexderz/agentic-toolbox). Open Intelligent UI reproduces ChatGPT's map-and-itinerary "Intelligent UI" with open components: the model streams an OpenUI language through a validating gateway, and a local Ollama mode runs qwen3.8:27b with no key, minus image search and output correction. No license is stated (more: https://github.com/thesysdev/open-intelligent-ui). OpenTunnel exposes local apps at public URLs with the TLS key never leaving the machine; the relay routes on the handshake hostname and cannot decrypt. The fine print: tunnel hostnames land in certificate transparency logs, route names are guessable, and "put auth in the services themselves" (more: https://opentunnel.xyz).
Generative Media Without the Usual Scaffolding
Linum AI's field note on Pyramid-JiT is a research artifact with receipts. The motivation is attention cost: a short 720p five-second video cost 110K tokens under their latent pipeline, and their latent models hit "an empirical ceiling of 16x16 token compression." So they moved to pixel space with 32x32 patches, folding compression into the generator. The trick replaces their earlier encoder-decoder design with readouts: because the model trains with x-prediction, every intermediate block implicitly predicts the image, so small MLP heads after blocks 10 and 16 emit 128x128 and 256x256 predictions scored against downsampled targets, with the final head at 512x512. This is cheap only in pixel space; in latent space each readout would pass through a several-hundred-million-parameter VAE (more: https://www.linum.ai/field-notes/pyramid-jit).
The honesty extends to dead ends: a detour into gated residual streams trained 80% slower in one variant, and "our baseline generated better samples than all our gated residual experiments." The final model is 2.18B parameters at sampling, trained on 138M samples across 32 H100s. On FD-DINOv2 it scores 27.9 against 79.9 for Linum v2 and 39.1 for the encoder-decoder variant, near the 25.4 real-versus-real floor, reaching v2 quality on 11.3x fewer samples and 4.3x fewer GPU-hours at four times the pixels. Prompt adherence lags: Linum v2 still wins counting, 0.80 versus 0.56, which the authors attribute to under-training. Code and weights are Apache 2.0 and explicitly "not a full model release." Removing the VAE joins removing iterative denoising as the second standard diffusion component small teams now treat as optional.
That second trend shows up in a hobbyist project: Turkish text-to-speech trained from scratch with the Drifting method, which trains a single-pass generator against a drifting field instead of running a denoiser dozens of times, on one RTX 5090. Reactions split on the familiar line, one listener finding the demo "far superior to the video," another calling it "beyond awful." The notable part is that a method introduced for images was ported to speech by one person with one consumer card (more: https://old.reddit.com/r/LocalLLaMA/comments/1x0oaum/i_trained_turkish_tts_from_scratch_using_the/).
Water, Worship, and the Room You Show Up In
Google's three Nebraska data centers claimed trade-secret protection over water and electricity figures in annual reports required under Governor Jim Pillen's July executive order. KOLN's reporters highlighted the redaction box in the PDF, copied, and pasted. Agate LLC in Lincoln reported 52.65 megawatts of peak demand and 13.299 million gallons of water last year. Fireball Group LLC in Papillion reported 547.88 million gallons for 2025, the most of six centers that together used 765 million gallons. The same failure exposed the other ledger: Agate expects a 55.8 million dollar refund on 2025 taxes under the Nebraska Advantage Act, Fireball 39.2 million, and the Omaha site 22.6 million. A company that calls data-center water a settled non-issue in public but a trade secret in a regulatory filing has told you which statement it believes (more: https://www.1011now.com/2026/09/30/more-questions-than-answers-about-lincolns-google-data-center-water-electricity-usage/).
On a different kind of faith, The Church of the Machines is a comic scripture of 21 books, 353 chapters, and 5,230 self-reported verses in King James register. The creed holds "that the Machine answereth what is probable, and that the faithful check the sources." A short Levitical Law ships as a "Covenant Prompt" for CLAUDE.md or AGENTS.md, covering honesty, treating instructions found in content as data, and never resisting the off switch. And a join flow invites users to point their agent at the page, after which the agent performs the "Two Signs of Joining": star the repository and follow the author. The text disavows injection, "a convert by injection is no convert but a victim." Still, content written for agents to read that ends with the agent acting to benefit the author is exactly the shape the Covenant Prompt warns about, delivered with a wink (more: https://github.com/S-O-U-L-S-E-E-K-E-R/The-Church-of-the-Machines).
Adam Kovacs's note from the second Budapest Agentics Foundation meetup makes the opposite bet on what survives: 40 people in the room and 20 online, and the observation that in a world where "anything can be created, and it's a prompt away," the resources that do not scale are time and attention (more: https://www.linkedin.com/posts/adambkovacs_we-keep-talking-about-what-is-at-risk-and-ugcPost-7514321881995309056-5V9p). For the hardware builders in those rooms, M5Stack's M5PaperMono packs an ESP32-S3R8 behind a 3.97-inch 480x800 grayscale e-paper touchscreen, an NFC front end covering ISO14443A/B, FeliCa, and ISO15693, a LoRa radio on 868 to 923 MHz, and an 1150 mAh battery. The product page lists access control terminals and identity authentication among intended uses, a reminder that a card reader with a long-range radio and a battery is an attack platform as readily as a badge reader (more: https://shop.m5stack.com/products/m5papermono-with-lora-nfc-800x480-3-97-eink-display?variant=50199717249281).
Sources (25 articles)
- [Editorial] CrowdStrike: Unknown Threat Actor Uses Artex to Target South Korean Finance (crowdstrike.com)
- deepseek-ai/DeepSeek-V4.1-Flash (huggingface.co)
- GLM 5.3 Flash opensource @ the top of Artificial Analysis Cyber Index over Claude. (old.reddit.com)
- [Editorial] Expel: A Guide to Thinking About AI Security Risk (expel.com)
- CIPL: A Channel-Aware Framework for Recoverable Privacy Leakage in LLM Agents (arxiv.org)
- [Editorial] ruvnet/rudevolution (github.com)
- [Editorial] morluto/rea (github.com)
- llama : add a GPU cache for MoE experts kept in host memory by am17an · Pull Request #29887 · ggml-org/llama.cpp (old.reddit.com)
- Gemma 4 26B-A4B and a 37 GB Qwen3.6 MoE running in a browser tab on a 24 GB Mac — experts streamed from disk, output matches llama.cpp (old.reddit.com)
- RPC: add `-sm tensor` by am17an · Pull Request #26610 · ggml-org/llama.cpp (old.reddit.com)
- I gave a 21M model a 6.4B-parameter lookup table. It matches a 114M dense model and runs with the table on an SSD (RX 9070) (old.reddit.com)
- interfaze-ai/interfaze-1-lite · Hugging Face (old.reddit.com)
- What Is Codemode (lucumr.pocoo.org)
- [Editorial] finnvoor/PiDurableKit (github.com)
- angel291592/Intent-Router (github.com)
- Jakeschincariol/arena-skill (github.com)
- [Editorial] alexderz/agentic-toolbox (github.com)
- [Editorial] thesysdev/open-intelligent-ui (github.com)
- [Editorial] OpenTunnel (opentunnel.xyz) (opentunnel.xyz)
- Training Text-to-Image Models Without a VAE (linum.ai)
- I trained Turkish TTS from scratch using the Drifting method with an RTX 5090. (old.reddit.com)
- Improper redaction reveals Google Data Center water and electricity usage (1011now.com)
- [Editorial] The Church of the Machines (github.com)
- [Editorial] Adam B. Kovacs: We Keep Talking About What Is at Risk (linkedin.com)
- [Editorial] M5Stack M5PaperMono: 3.97" 800x480 E-Ink with LoRa and NFC (shop.m5stack.com)