Quantization & Efficiency

Model compression, GGUF, efficient inference, optimization

801 articles across 204 editions

Articles

  1. Claude Haiku 5.5 -- 2026-10-08
  2. reddit.com -- 2026-10-08
  3. Editorial video submission (YouTube: Cl3OWig5hkk) -- 2026-10-08
  4. Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s -- 2026-10-06
  5. Two local Qwen ( 3.8 27b unsloth Q6 and Qwen flash next strata coder ) models vs Claude Opus 4.6 on the same 3 coding tasks. One of them tied it. Not here to start a fight, just sharing numbers -- 2026-10-06
  6. Qwen4Exp: add MTP by am17an · Pull Request #29761 · ggml-org/llama.cpp -- 2026-10-06
  7. [Discussion] A 5KB pure x86-64 assembly engine for Gemma-2B (FP16, 4.6 tok/s on CPU) -- 2026-10-06
  8. DDR4/PCIe4 vs DDR5/PCIe5 for LLMs- I benchmarked them for pre-training. What are your thoughts? -- 2026-10-06
  9. [Editorial] Post by @ashxhart on X -- 2026-10-06
  10. EmbeddingGemma 2 running locally in-browser on WebGPU -- 2026-10-06
  11. Decodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection -- 2026-10-06
  12. [Editorial] Mistral Large 4 model documentation -- 2026-10-06
  13. [Editorial] Mistral CEO says new AI model beats Chinese ones in some areas (Reuters) -- 2026-10-06
  14. Q (@qtnx_) on X - Mistral Large 4 is still doing RL runs, keep seeing improvements (vs preview version). Release at the end of the month -- 2026-10-06
  15. Reflection AI Is About to Release a US Open-Weight Model to Take On DeepSeek and Qwen -- 2026-10-06
  16. Decision models like Jev don't beat LLM-as-a-judge or traditional classifiers -- 2026-10-05
  17. [Editorial] arXiv 2609.37647 -- 2026-10-05
  18. [Editorial] decision-index (apolinario/decision-index) on GitHub -- 2026-10-05
  19. [Editorial] Jev Decision Index (Hugging Face Space by multimodalart) -- 2026-10-05
  20. [Editorial] Korea JoongAng Daily: AI hacking spree exposes cracks in financial-sector security -- 2026-10-05
  21. [Editorial] Bloomberg: Hackers breached propulsion system of US-bound oil tanker -- 2026-10-05
  22. BSI Notes on Classic McEliece (Post-Quantum Cryptography) -- 2026-10-02
  23. Stolen Thoughts: Research Update (PDF) -- 2026-10-02
  24. Perplexity: Escaping Space, Part I -- 2026-10-02
  25. Hijacking the PS5's RTMP stream -- 2026-10-02
  26. [Editorial] -- 2026-10-01
  27. [Editorial] -- 2026-10-01
  28. [Editorial] -- 2026-10-01
  29. Ember-1 -- 2026-09-29
  30. 42x Faster Prompt Lookup Drafting in llama.cpp -- 2026-09-29
  31. MiaAI-Lab/DeepSeek-v4.1-Flash-EXL3-2x-DGX-Sparks -- 2026-09-29
  32. ESP32S3 cluster running 1.58-bit (BitNet) Language model -- 2026-09-29
  33. [Editorial] -- 2026-09-28
  34. FT: Corporate America rejects overpriced frontier, embraces open models -- 2026-09-28
  35. GPT-3 is discontinued today -- 2026-09-28
  36. nex-agi/Nex-N2-Pro -- 2026-09-28
  37. [Splash Engine] Qwen3.8-27B in native 8-bit at 37-55 tok/s on Apple Silicon, 256k context and the Reasoning Cliff -- 2026-09-25
  38. Bonsai 2: Qwen 3.8 27B quality at 5.9GB VRAM via ternary compression -- 2026-09-25
  39. R9V Update: KVA projections for Qwen3.8 Flash Next give 1.45-1.85x prefill speedup on 2x R9700 -- 2026-09-25
  40. [Editorial] Sebastian Raschka on LLM research (x.com/rasbt) -- 2026-09-24
  41. MiMo 2.6 Pro: Reducing overthinking and second-guessing -- 2026-09-24
  42. XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B -- 2026-09-24
  43. quants for K2-Horizon are now available -- 2026-09-24
  44. [Editorial] Jev's Architecture Unmasked (archerhume.com) -- 2026-09-24
  45. Got jev-like api running natively on ninfer / qwen3.8 27b and results are quite decent -- 2026-09-24
  46. Mods: can we do something about half the forum getting filled with these advertising posts for Jev? -- 2026-09-24
  47. [Editorial] brainstormity post (x.com) -- 2026-09-24
  48. What are KV caches, really? (Glenn Lockwood) -- 2026-09-23
  49. Dynamic Quantiser - a way to make your own high quality dynamic quants -- 2026-09-23
  50. Transformers now runs llama.cpp quants -- 2026-09-23
  51. tokenizers v1: encode, decode and scaling, measured -- 2026-09-23
  52. CrofAI "cheapest inference provider in the world" gets exposed as an OpenRouter wrapper, routing requests to smaller, cheaper models at up to 20x markup. CrofAI responds to Wire Fraud allegations by denying everything, then backtracking, then 3 hours later wiping their entire online presence -- 2026-09-22
  53. [Editorial] pwardle/not-a-mused -- 2026-09-22
  54. Can gzip be a language model? -- 2026-09-22
  55. circle-group/hktex -- 2026-09-22
  56. Attention is all you have -- 2026-09-22
  57. Pruning LLMs Like a Physicist: Block Removal as an Ising Optimization Problem -- 2026-09-22
  58. [Editorial] AikidoSec/altar-1 on Hugging Face -- 2026-09-22
  59. Breaking the 1.58-bit Barrier for Ternary LLMs -- 2026-09-21
  60. Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint -- 2026-09-21
  61. No GPU? No Problem: Flagship LLMs on a GPU-less Teenaged Server -- 2026-09-21
  62. Qwen-3.8-Flash-Next on 1x RTX 5090: TG=50 t/s, PP=2300 t/s - with FreeToken -- 2026-09-21
  63. Show HN: Mini-AGI – Dynamic continual learning model trained on 8GB VRAM -- 2026-09-21
  64. nvidia/Cosmos3-Super-Text2Image -- 2026-09-21
  65. Oilbird: Training-Free Speculative Decoding with Keys the Verifier Already Computes -- 2026-09-18
  66. Running Qwen3.8-Flash-Next locally on a 12GB VRAM card -- 2026-09-18
  67. 10%+ performance improvement on MoE ssd-streaming with expert-lookahead -- 2026-09-18
  68. Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data -- 2026-09-18
  69. Apple Reference Image: A New Approach for Verified Photography -- 2026-09-17
  70. [Editorial] PirateFace -- 2026-09-17
  71. [Editorial] Colibri (JustVugg/colibri) -- 2026-09-16
  72. Qwen3.8 Flash Next now at 1.2k t/s prefill on Strix Halo -- 2026-09-16
  73. NVIDIA PAIR routing to llama.cpp on an AMD ROCm node (2×R9700). Notes. -- 2026-09-16
  74. Nvidia's RTX 5090 vanishes from online retail in the US — third-party sellers now demand as much as $9,500 for Nvidia's fastest GPU -- 2026-09-16
  75. DeepSeek V4.1 Flash beats Astra on AA's new benchmark -- 2026-09-14
  76. DSV4 Flash 0731 on OpenRouter. Why is the price SO LOW -- 2026-09-14
  77. [Editorial] So you want to use OpenRouter -- 2026-09-14
  78. IFM/K2-Horizon-MoVA-36B-A4B -- 2026-09-14
  79. [Editorial] The Goodies: vehicles proposal (adrianco) -- 2026-09-14
  80. Music Theory for the 21st-Century Classroom -- 2026-09-14
  81. [Editorial] Claude Code artifact -- 2026-09-14
  82. [Editorial] arXiv paper 2609.11799 -- 2026-09-14
  83. Open Source Acoustic Drone Detection -- 2026-09-14
  84. tonyd2wild/GLM-5.3-Flash-NVFP4-DFlash2-2x-DGX-Spark -- 2026-09-11
  85. Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses -- 2026-09-11
  86. I made a custom llama.cpp build optimized for 7900xtx (one or two). for qwen 3.8 next and 27B. includes optimizations for PciE x4 and tensor parallel. read inside! (no AI slop) -- 2026-09-11
  87. [Editorial] GCD-AuthZ: authorization paper (cybersharkvin) -- 2026-09-10
  88. Gradient Immunity: Null-Space Resistance to Malicious Fine-Tuning -- 2026-09-10
  89. Detecting and countering misuse of AI: September 2026 -- 2026-09-10
  90. Reduced Matrix Multiplication: Input-Adaptive Matrix-Product Reduction for LLM Inference -- 2026-09-09
  91. GitHub - coder543/minnow: Fast LLaDA2.2 inference server -- 2026-09-09
  92. Expert expansion with llama.cpp -- 2026-09-09
  93. inclusionAI/Ling-3.0-flash-VL · Hugging Face -- 2026-09-09
  94. model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support by YanissAmz · Pull Request #25444 · ggml-org/llama.cpp -- 2026-09-09
  95. Friends Don't Let Friends Use Ollama -- 2026-09-09
  96. You Don't Need More VRAM To Run AI At Home (The Stack) -- 2026-09-08
  97. exllamav3 comfortably beats llama.cpp running CPU-offloaded Qwen-3.8-Flash-Next on my setup! -- 2026-09-08
  98. Spark-2.5-4B is an interesting model for 8GB Jetson Orin Nano Super SoC. -- 2026-09-08
  99. VoxGen, an AMD-optimized TTS inference engine for VoxCPM 2 models -- 2026-09-08
  100. Building a tiny ElevenLabs on a single 3090 in 2-hour runs. Here's the log of everything that broke. -- 2026-09-08
  101. Structured Memory for Edge Language Models: Persistent Context and Corpus Retrieval via O(1) SSM State Injection -- 2026-09-07
  102. [Editorial] -- 2026-09-07
  103. hku-sail/StreamPI -- 2026-09-07
  104. BenchMIRT: What are LLM benchmarks actually measuring? -- 2026-09-07
  105. A Four-Axis Trustworthiness Benchmark for LLM-as-Judge in Principle-Based Regulation -- 2026-09-07
  106. [Editorial] -- 2026-09-07
  107. [Editorial] -- 2026-09-07
  108. [Editorial] Autoresearch: Sticky Refusals, Free Speculative Decoding, and the Invisible Quantisation Cliff -- 2026-09-04
  109. [Editorial] Weightless (msuiche) -- 2026-09-04
  110. Anyone else notice strange refusal-related reasoning traces from Qwen3.8-Flash-Next during routine coding sessions? -- 2026-09-04
  111. Path to Astra: critical capabilities and frontier safeguards -- 2026-09-04
  112. Gemini 3.8 Flash and 3.8 Flash Cyber -- 2026-09-04
  113. [Editorial] The Coming Split in Models: Learning to Sort the Computable (Balaji Lakshmanan) -- 2026-09-03
  114. 2akouwu/reverify -- 2026-09-03
  115. azrtydxb/procoder -- 2026-09-03
  116. Frona v2026.8.0 – self-hosted personal AI assistant with ontology memory -- 2026-09-03
  117. [Editorial] The Emergent Symbolic Structure of Artificial Neural Networks (arXiv 2608.29530) -- 2026-09-03
  118. How to build a diffusion language model -- 2026-09-03
  119. Qwen3.8-Flash-Next on 2x3090 + DDR4: 17 → 25-29 t/s decode with the expert cache PR -- 2026-09-03
  120. It's official! 192GB Framework -- 2026-09-03
  121. Vision support merged for DeepSeek-V4-Flash-Vision-Exp -- 2026-09-03
  122. Muse Spark open weights coming soon -- 2026-09-03
  123. internlm/Intern-S2-Preview -- 2026-09-03
  124. [Editorial] -- 2026-09-01
  125. [Editorial] -- 2026-09-01
  126. [Editorial] -- 2026-09-01
  127. [Editorial] -- 2026-09-01
  128. [Editorial] -- 2026-09-01
  129. LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V6-GGUF -- 2026-09-01
  130. Hidden in the Request: Explaining Unethical LLM Compliance through Token Relevance -- 2026-09-01
  131. Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment -- 2026-08-28
  132. Semantic Browsing: Controllable Diversity for Image Generation -- 2026-08-28
  133. [Editorial] OpenAI: Jalapeño First Results -- 2026-08-26
  134. Apple introduces new Mac Studio with M5 Max and M5 Ultra - up to 512GB of unified memory -- 2026-08-26
  135. I hosted Kimi K3 (2.8T parameters) using 8 B300s. 92 tok/s, $190 per million tokens -- 2026-08-26
  136. We quantized Qwen 3.8 27B and compared the quants on an RTX 6000 -- 2026-08-26
  137. [Editorial] Signal Windows Desktop: ContentProtection Bypass (IOActive) -- 2026-08-26
  138. Spoofed Serial Number Unlocks Cricut Machine -- 2026-08-26
  139. "One robot could infect other vulnerable robots nearby ... Attackers could take control of entire fleets of robots." -- 2026-08-26
  140. [Editorial] arXiv 2608.21986 -- 2026-08-25
  141. [Editorial] IACR ePrint 2026/1734 -- 2026-08-25
  142. [Editorial] Unsloth Dynamic 3.0 GGUFs -- 2026-08-20
  143. LFM2.5 Q4_0 Checkpoints from Quantization-Aware Distillation -- 2026-08-20
  144. [2511.07885] Intelligence per Watt: Measuring Intelligence Efficiency of Local AI -- 2026-08-20
  145. Transformer-Guided Swarm Intelligence for Frugal Neural Architecture Search -- 2026-08-20
  146. 0xSero/deepseek-v4-flash-0731-spark-sparkinfer -- 2026-08-20
  147. [Editorial] Deepseek-API -- 2026-08-20
  148. Linguistic Firewall: Geometry as Defense in Multi-Agent Systems Routing -- 2026-08-19
  149. Random Logit Scaling: Defending Deep Neural Networks Against Black-Box Score-Based Adversarial Example Attacks -- 2026-08-19
  150. bitsandbytes creator teasing new quantization method: GLM 5.3 on a single DGX Spark at 7t/s -- 2026-08-18
  151. GLM 5.3 weights. It might offer the best capacity-to-size ratio. -- 2026-08-18
  152. Ling 3.0 support merged into llama.cpp -- 2026-08-18
  153. Why not? ☺️ -- 2026-08-18
  154. We quantized DeepSeek V4 0731 and benchmarked it against popular quants on 8× RTX 5090 -- 2026-08-14
  155. Gemma 4 12B Q3: +8.55% Coding Performance From Tensor-Level Quantization Allocation -- 2026-08-14
  156. I trained a 1B-parameter LLM from scratch on 20B tokens for about $200 -- 2026-08-14
  157. Showoff Saturday: Local 4x 6000 Pro (multi-year progression) -- 2026-08-14
  158. [Editorial] Foxflow Fleet (probagi.com) -- 2026-08-14
  159. State of Open Models: Summer 2026 Observations -- 2026-08-14
  160. Qwen 3.8 27B is out : open weights, best local dense model yet -- 2026-08-14
  161. CohereLabs/North-Micro-Vision-Instruct · Hugging Face -- 2026-08-14
  162. U.S. Department of Energy Launches the Genesis Open Models Initiative and, with Arcee, Unveils Genesis-Science-1 — Its First Open-Weight Model for Scientific Research -- 2026-08-14
  163. enabling PCI-E p2p for consumer Nvidia cards will yield you more than you think -- 2026-08-12
  164. Rumored 50-series Super refresh bumps everything +50% VRAM -- 2026-08-12
  165. Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning -- 2026-08-12
  166. [2606.05682] Beyond Output Matching: Preserving Internal Geometry in NVFP4 LLM Distillation -- 2026-08-12
  167. torchtune: PyTorch native post-training library -- 2026-08-12
  168. Qwen3.8-2.4T-A95B Released -- 2026-08-12
  169. Exact Qwen 3.8 27b release date and time -- 2026-08-12
  170. inclusionAI/Ling-2.6-1T -- 2026-08-12
  171. nvidia/Nemotron-Cascade-2-30B-A3B -- 2026-08-12
  172. DiffusionGemma Technical Report -- 2026-08-12
  173. LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge -- 2026-08-12
  174. llama.cpp -- 2026-08-12
  175. [Editorial] Anthropic's Invisible Watermarks in Claude Text -- 2026-08-11
  176. [Editorial] Reuven Cohen: Quantum Mechanics May Be Telling Us Something -- 2026-08-11
  177. Muse Glimmer 30B + DFlash speculative decoding on vLLM: 6 patches needed, 25 → 57 tok/s. Dockerfile and numbers inside. -- 2026-08-11
  178. Apple Silicon and macOS VMs: 11–16× Faster LLM Inference with Llama.cpp -- 2026-08-11
  179. Echo Dot 2 can run 28M LLM at decent speed -- 2026-08-11
  180. Chunked KL loss for running Knowledge Distillation locally (<6GB VRAM at 32K context length) -- 2026-08-11
  181. Sonic Pi v5 -- 2026-08-11
  182. MiniMax-AI/MiniMax-H3 -- 2026-08-11
  183. Glimmer seems pretty censored? -- 2026-08-11
  184. No wonder Qwen and Gemma are so different -- 2026-08-11
  185. CJK Manga/Manhwa/Manhua 150M OCR model (hayai-ocr-v2) outperforming PaddleOCR-VL-For-Manga -- 2026-08-11
  186. [Editorial] -- 2026-08-10
  187. [Editorial] -- 2026-08-10
  188. [Editorial] -- 2026-08-10
  189. DeepSeek V4 Flash 2-bit quant achieves 100% on SQL benchmark locally -- 2026-08-07
  190. Scotoma-2: Gemma4, but with less annoying slop and better writing -- 2026-08-07
  191. Mach-1 Additive: 95% of Qwen 3.6 35B performance while 10x smaller -- 2026-08-07
  192. Intern S2 Mobius — Qwen3.5-35B derivative with architectural throughput gains -- 2026-08-07
  193. [Update] DeepSeek-V4-Flash-0731 on a single RTX 5090: phase-adaptive DSpark K1/K2 with dual CUDA graphs -- 2026-08-05
  194. TensorSharp DSpark Benchmark on DeepSeek V4 Flash — up to 2x speedup -- 2026-08-05
  195. DeepSeek V4 Flash on a single RTX 4090 at 64k context — complete config and measured numbers -- 2026-08-05
  196. DeepSeek v4 Flash vs. Qwen3.6-27B, 3.5-122B, and Gemma 4 31B — Local Agentic Coding Benchmark -- 2026-08-05
  197. Vacuum 16T: A 16.5-trillion-parameter model that contains nothing -- 2026-08-03
  198. I benchmarked classic vector RAG vs Google's new OKF format vs both combined -- 2026-08-03
  199. [Editorial] Awesome Systematic Trading -- 2026-08-03
  200. [Editorial] -- 2026-07-30
  201. What do we know about the "AI Accelerators" used to train LongCat-2? -- 2026-07-24
  202. Running a 13M ASR conformer on a microcontroller -- 2026-07-24
  203. Running Qwen 3.6 35B MoE (Q4_K_M) on a Zeus (Xiaomi 12 Pro, 12GB RAM) -- 2026-07-24
  204. Squeeze-Release: Iterative Pruning with Exact Structural Minimization -- 2026-07-24
  205. The Open Weight Unbundling -- 2026-07-24
  206. Model "distillation" accusations are getting way overblown at this point -- 2026-07-24
  207. The Distillation Claims Are Fake and Desperate -- 2026-07-24
  208. I gave Kimi K3 a shot at auditing my post-quantum crypto project, it found 5 real bugs Fable/Opus 4.8 and GPT-5.6 Sol had all missed -- 2026-07-24
  209. China's open-weights AI strategy is winning -- 2026-07-22
  210. How do we benefits from 2+ T models? -- 2026-07-22
  211. Serving a fleet of Qwen3.5 122b sessions on a single Mac Studio (96GB) without losing your sanity -- 2026-07-22
  212. Quant Convergence: Bridging Classical Value Investing and Modern Factor Models for Systematic Equity Selection -- 2026-07-21
  213. [Editorial] -- 2026-07-21
  214. [Editorial] -- 2026-07-21
  215. OpenAI had to pause an unreleased model after it escaped containment. -- 2026-07-21
  216. [Editorial] -- 2026-07-21
  217. [Editorial] -- 2026-07-16
  218. [Editorial] -- 2026-07-16
  219. [Editorial] -- 2026-07-16
  220. [Editorial] -- 2026-07-16
  221. FT: Companies Turn to Chinese Open Weight Models to Cut Costs -- 2026-07-14
  222. PrismML Compresses Qwen-3.6-27B to Under 4GB — Runs on iPhone 17 Pro -- 2026-07-14
  223. [Editorial] ThinkingCap Qwen3.6-27B -- 2026-07-14
  224. GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine -- 2026-07-14
  225. Qwen3.6-27B: NVFP4/FP8 agent loops vs flawless BF16. Config or quant issue? -- 2026-07-13
  226. NVIDIA Puzzle-75B-A9B NVFP4 at 132 t/s on 3×3090 — Why is this size category a desert otherwise? -- 2026-07-13
  227. Exploring FlashAttention-3/4 optimizations on RTX GPUs -- 2026-07-13
  228. [audio.cpp] What Does the Fox Say: 4 ASR models in native C++/GGML, streaming support, 327s transcribed in 2.17s -- 2026-07-13
  229. [Editorial] -- 2026-07-13
  230. I benchmarked 13 models at 65K-128K context to find out what actually matters for agentic workloads -- 2026-07-10
  231. Deepseek V4 Flash running on RTX 5090 MoE -- 2026-07-10
  232. Seasonic PSU calculator now mentions RTX 5080 SUPER (24GB), RTX 5070 Ti SUPER (24GB) and RTX 5070 SUPER (18GB) -- 2026-07-10
  233. Soon we'll run 100B models on cheap hardware -- 2026-07-10
  234. Döner Bench round 2: Quant compare -- 2026-07-10
  235. Pine64 launch $50 smart speaker for Home Assistant tinkerers -- 2026-07-03
  236. Tiny Jetson Nano Orin Super Benchmarking of 1B and sub 1B LLMs — llama.cpp vs Ollama -- 2026-07-03
  237. nvidia/Cosmos3-Nano -- 2026-07-03
  238. Microsoft has taken down fastcontext model from everywhere -- 2026-07-02
  239. Orthrus (diffusion head) trained Qwen 3.5/3.6 and Gemma 4 models are dropping soon -- 2026-07-02
  240. SenseNova-U1-8b-MoT-Infographic-V2 (released yesterday) - An open source SOTA beast for infographic design and image editing. -- 2026-07-02
  241. on Dario's statement -- 2026-07-02
  242. After i distilled a 26B into a 4B to cut false positives I'll probably use the base model after all. -- 2026-07-02
  243. We built a calibration-aware Q4_K_M quant of Qwen3.5 0.8B that recovers 96.5% of the BF16 gap vs pure llama.cpp Q4_K_M (SpectralQuant) -- 2026-06-29
  244. Update: First Manual Results from Testing Procedural Skill Transfer in Small Models -- 2026-06-29
  245. Multi Tier MoE Caching -- 2026-06-29
  246. Apple raises prices of MacBooks, iPads -- 2026-06-26
  247. IBM debuts sub-1 nanometer chip technology -- 2026-06-26
  248. OpenAI and Broadcom unveil LLM-optimized inference chip -- 2026-06-26
  249. Ideogram 4: Open Image Model at the Forefront of Design -- 2026-06-25
  250. QUEST-35B: Open-Source Deep Research Agent Trained on 32 H100s — Full Recipe, Weights, and Data Released -- 2026-06-25
  251. North Mini Code: 4-Bit Quant + Ollama + OpenRouter — Now Runs on 20GB -- 2026-06-25
  252. EdgeRazor: Mixed-Precision Quantization-Aware Distillation Down to 1.58-Bit -- 2026-06-25
  253. [Editorial] Qwable-3.6-27b — Open Model Release -- 2026-06-25
  254. ScenemaAI/scenema-audio — Audio Generation Model -- 2026-06-25
  255. "Tokaine Addiction" — PhD Student's Colleague Can't Stop Running AI Agents -- 2026-06-25
  256. Suitcase Robot Uses Gas Sensor to Modulate LLM Sampler Temperature in Real-Time -- 2026-06-25
  257. [Editorial] Microsoft Majorana 2: Quantum Discovery via Agentic AI -- 2026-06-17
  258. cuTile Rust: Safe, Data-Race-Free GPU Kernels in Rust -- 2026-06-17
  259. [Editorial] Adversarial AI Research — The Malicious Use of Artificial Intelligence -- 2026-06-17
  260. [Editorial] Sloptimization — AI Citation Rot and the Tools Fighting Back -- 2026-06-15
  261. [Editorial] Leapable AI Platform -- 2026-06-15
  262. huawei-csl/KVarN -- 2026-06-12
  263. chiennv2000/orthrus -- 2026-06-12
  264. Qwen/Qwen3.5-122B-A10B -- 2026-06-12
  265. zai-org/GLM-OCR -- 2026-06-12
  266. How's Linear so fast? A technical breakdown -- 2026-06-11
  267. Port React Compiler to Rust -- 2026-06-11
  268. Show HN: Extend UI – open-source UI kit for modern document apps -- 2026-06-11
  269. MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second -- 2026-06-09
  270. Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States -- 2026-06-09
  271. A Post-Quantum Future for Let's Encrypt -- 2026-06-08
  272. Gooey: A GPU-accelerated UI framework for Zig -- 2026-06-05
  273. Branchless Quicksort faster than std:sort and pdqsort with C and C++ API -- 2026-06-05
  274. HP re-releases classic computer science calculator: The HP-16C -- 2026-06-05
  275. OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization -- 2026-05-29
  276. Shard — getting to 10× KV cache compression -- 2026-05-29
  277. LightVLM: Efficient inference toolkit for vision-language models -- 2026-05-29
  278. Speech Tokenizer Arena: Side-by-side benchmarking for discrete speech tokenizers -- 2026-05-29
  279. DeepSeek just popped the American AI bubble. -- 2026-05-29
  280. DeepSeek V4 Flash at 8.4 tok/s on 3×3090 — patching GGUF metadata for cchuter's fork -- 2026-05-29
  281. Why are the AI Companies spreading F.U.D. about AI? -- 2026-05-29
  282. $400 Qwen 3.6-27B Setup — Dual RTX 3060 — 30-50 t/s -- 2026-05-27
  283. Qwen3.6 27B made me a believer — single-shot game development with local model -- 2026-05-27
  284. Qwen3.5 35B-A3B Uncensored Heretic — Native MTP Preserved, multiple formats -- 2026-05-27
  285. [Editorial] MSI NVIDIA DGX Station -- 2026-05-27
  286. Have we passed the peak of inflated expectations? -- 2026-05-27
  287. Separable Expert Architecture: Privacy-Preserving LLM Personalization via Composable Adapters -- 2026-05-21
  288. 512k Context Pre-training on a 12GB Consumer GPU with O(n) Attention -- 2026-05-21
  289. Introducing the Ettin Reranker Family -- 2026-05-20
  290. OlmoEarth v1.1: A more efficient family of models -- 2026-05-20
  291. Regex Chess: A 2-ply minimax chess engine in 84,688 regular expressions -- 2026-05-19
  292. FastDMS: 6.4X KV-cache compression running faster than vLLM BF16/FP8 -- 2026-05-11
  293. [Editorial] RuVector Sparse Attention Crate -- 2026-05-11
  294. AMD to release slottable GPU -- 2026-05-11
  295. Taiwanese company Skymizer announces HTX301 - PCIE inference card with 384GB of Memory at ~240 Watts -- 2026-05-11
  296. Apple Removes 256GB M3 Ultra Mac Studio Model From Online Store -- 2026-05-11
  297. Virtual violin produces realistic sounds (MIT) -- 2026-05-06
  298. Lightricks/LTX-2.3-22b-IC-LoRA-HDR -- 2026-05-06
  299. [Editorial] Finding Zero-Days with Any Model -- 2026-05-01
  300. [Editorial] OMLX.ai -- 2026-04-30
  301. llama.cpp - NVFP4 native support on Blackwell from now - b8967 -- 2026-04-30
  302. Sigilant: GGUF Quality Benchmarking Beyond TPS — Tool-Calling Pass Rate as Selection Criterion -- 2026-04-30
  303. I'm done with using local LLMs for coding -- 2026-04-30
  304. Qwen3.6-27B IQ4_XS FULL VRAM with 110k context -- 2026-04-28
  305. Can we already use Google's TurboQuant (TQ) for KV Cache in llama-server? Or are we waiting for a PR? -- 2026-04-28
  306. Gemma 4 beats Qwen 3.5 (UPDATE), and Qwen 3.6 27B + MiniMax M2.7 is the best OpenCode setup -- 2026-04-28
  307. Tried Qwen3.6-27B-UD-Q6_K_XL.gguf with CloudeCode, well I can't believe but it is usable -- 2026-04-28
  308. BitNet is the AI future? -- 2026-04-28
  309. FP4 inference in llama.cpp (NVFP4) and ik_llama.cpp (MXFP4) landed -- 2026-04-27
  310. [Editorial] Bonsai-8B MLX 1-bit -- 2026-04-27
  311. VRAM.cpp: Running llama-fit-params directly in your browser -- 2026-04-27
  312. Thoughts on using an AMD Alveo V80 FPGA as a poor man's Taalas HC1 -- 2026-04-27
  313. [Editorial] hw-smi — Cross-Platform Hardware Monitor -- 2026-04-27
  314. China's DeepSeek valuation rockets above $20B!! -- 2026-04-24
  315. [Editorial] DeepSeek Open-Sources Tile Kernels -- 2026-04-24
  316. DeepSeek v4 -- 2026-04-24
  317. [Editorial] Video Content -- 2026-04-22
  318. [Editorial] Four Horsemen of the AIpocalypse -- 2026-04-22
  319. Closest replacement for Claude + Claude Code? (got banned, no explanation) -- 2026-04-22
  320. [Editorial] Roomote: Remote Development Tool -- 2026-04-22
  321. [Editorial] Video Content -- 2026-04-22
  322. Ternary Bonsai: Top Intelligence at 1.58 Bits -- 2026-04-22
  323. Personal Eval: Gemma4 26B MoE vs Qwen3.5 27B Dense vs Gemma4 31B Dense Compared -- 2026-04-22
  324. NVIDIA Nemotron-3-Super-120B-A12B-FP8 -- 2026-04-22
  325. [Editorial] Stanford HAI AI Index Report 2026 -- 2026-04-17
  326. [Editorial] Steve Yegge on AI -- 2026-04-17
  327. [Editorial] The AI Resentment Stage -- 2026-04-17
  328. A cryptography engineer's perspective on quantum computing timelines -- 2026-04-07
  329. Sam Altman may control our future – can he be trusted? -- 2026-04-07
  330. [Editorial] Anthropic's Claude Code Source Leak — What It Means -- 2026-04-01
  331. [Editorial] Claude Code Was Leaked — I Read All of It -- 2026-04-01
  332. [Editorial] The Claude Code 'Oops' — Source Code Leak -- 2026-04-01
  333. [Editorial] Claude Code Just Open-Sourced Itself (Not Intentionally) -- 2026-04-01
  334. [Editorial] nirholas/claude-code Repository -- 2026-04-01
  335. [Editorial] arxiv:2603.15569 -- 2026-03-30
  336. TinyLoRA: LoRA training works at just 13 parameters -- 2026-03-30
  337. KV rotation PR: q8 quants tank performance on AIME25, recovered with rotation -- 2026-03-30
  338. [Editorial] AI ASIC for LLMs -- 2026-03-30
  339. [Editorial] Heretic -- 2026-03-30
  340. mlx-snn: Spiking Neural Network library for Apple MLX -- 2026-03-30
  341. TurboQuant: Redefining AI efficiency with extreme compression -- 2026-03-27
  342. [Editorial] TurboQuant Deep Dive -- 2026-03-27
  343. RightNow-AI/autokernel -- 2026-03-27
  344. NVIDIA 2026 Conference LIVE. New Base model coming! -- 2026-03-20
  345. Nemotron 3 Nano 4B: A Compact Hybrid Model for Efficient Local AI -- 2026-03-20
  346. Granite 4.0 1B Speech: Compact, Multilingual, and Built for the Edge -- 2026-03-20
  347. [New Model & Agent] LocoTrainer-4B: A Claude Code-style local agent designed specifically to master the MS-SWIFT framework (4B, 32K, GGUF) -- 2026-03-20
  348. shallowdream204/BitDance-14B-16x -- 2026-03-20
  349. [Editorial] Claude 1M Context GA -- 2026-03-14
  350. NVIDIA Nemotron 3 Super: open-weight 120B MoE hybrid with 1M-token context -- 2026-03-14
  351. Expert parallelism for 1T MoE finetuning on a single node - 50x faster and 2x cheaper than alternatives -- 2026-03-14
  352. Fine-tuned Qwen3 SLMs (0.6-8B) beat frontier LLMs on narrow tasks -- 2026-03-12
  353. Towards Cold-Start Drafting and Continual Refining: A Value-Driven Memory Approach with Application to NPU Kernel Synthesis -- 2026-03-12
  354. [Editorial] -- 2026-03-09
  355. My journey through Reverse Engineering SynthID -- 2026-03-09
  356. Did Alibaba just kneecap its powerful Qwen AI team? -- 2026-03-07
  357. [Editorial] You're 1,191 Days Late — Here's What to Do -- 2026-03-07
  358. [Editorial] The Zero-Day Clock Is Ticking -- 2026-03-06
  359. [Editorial] Step-by-Step Guide to Exploiting AI Systems -- 2026-03-06
  360. [Editorial] Unprompted 2026: Top Insights Day One -- 2026-03-06
  361. [Editorial] Unprompted 2026: Top Insights Day Two -- 2026-03-06
  362. The L in "LLM" Stands for Lying -- 2026-03-05
  363. [Editorial] Am I Living in a Parallel AI Universe? -- 2026-03-05
  364. [Editorial] Video Pick -- 2026-03-05
  365. Qwen3.5 122B in 72GB VRAM (3x3090) is the best model available at this time — also it nails the "car wash test" -- 2026-03-03
  366. Qwen 3.5 is multimodal. Here is how to enable image understanding in opencode with llama cpp -- 2026-03-03
  367. pplx-embed: State-of-the-Art Embedding Models for Web-Scale Retrieval -- 2026-03-03
  368. andimarafioti/faster-qwen3-tts -- 2026-03-03
  369. Meta's AI smart glasses and data privacy concerns -- 2026-03-03
  370. [Editorial] AI Search Index -- 2026-03-03
  371. Mercury 2: Fast reasoning LLM powered by diffusion -- 2026-02-26
  372. I Benchmarked Opus 4.6 vs Sonnet 4.6 on agentic PR review and browser QA the results weren't what I expected -- 2026-02-26
  373. [Editorial] Bullshit meter :) -- 2026-02-26
  374. The Qwen team verified that there are serious problems with the data quality of the GPQA and HLE test sets. -- 2026-02-25
  375. Qwen 3.5 craters on hard coding tasks — tested all Qwen3.5 models (And Codex 5.3) on 70 real repos so you don't have to. -- 2026-02-25
  376. ChatGPT isn't the only chatbot pulling answers from Elon Musk's Grokipedia -- 2026-02-25
  377. [Editorial] Benchmarking LLMs for Voice Agent Use Cases -- 2026-02-21
  378. Claude Opus 4.6 Surges Past Forecasts on METR's 50% Time-Horizon Benchmark with Exponential Gains -- 2026-02-21
  379. [Editorial] Unsloth: MiniMax M2.5 Fine-Tuning Guide -- 2026-02-21
  380. Let your coding agent benchmark llama.cpp for you (auto-hunt the fastest params per model) -- 2026-02-06
  381. GGML implementation of Qwen3-ASR -- 2026-02-06
  382. Running LLMs &amp; VLMs Fully On-Device on iPhone(6GB RAM) — Offline, Privacy-Focused, Real-Time Performance -- 2026-02-06
  383. We benchmarked every 4-bit quantization method in vLLM 👀 -- 2026-01-12
  384. Gpu inference with model that does not fit in one GPU -- 2026-01-12
  385. Llama.cpp rpc experiment -- 2026-01-12
  386. Performance improvements in llama.cpp over time -- 2026-01-12
  387. [Editorial] https://docs.rs/crate/bitchat-qudag/latest -- 2026-01-02
  388. [Editorial] https://github.com/permissionlesstech/bitchat/blob/main/WHITEPAPER.md -- 2026-01-02
  389. [Editorial] https://www.npmjs.com/package/@ruvector/edge-net -- 2026-01-02
  390. Why I Ditched Serverless Neptune/OpenSearch for Dockerized Neo4j/pgvector on EC2 (60% Cost Cut) -- 2025-12-30
  391. Llama-3.3-8B-Instruct -- 2025-12-30
  392. Benchmarking local llms for speed with CUDA and vulkan, found an unexpected speedup for select models -- 2025-12-30
  393. Why Kimi K2 Thinking choose Int4 QAT, from infra enginner of KImi -- 2025-12-30
  394. Help RTX 5090 + llama.cpp crashes after 2-3 inferences (VFIO passthrough, SM120 CUDA) -- 2025-12-30
  395. AI-Doomsday-Toolbox Distributed inference + workflows -- 2025-12-30
  396. [Tool] imesde: Zero-GPU, In-Memory Vector Engine for Real-Time Local RAG -- 2025-12-22
  397. I built a Rust-based HTML-to-Markdown converter to save RAG tokens (Self-Hosted / API) -- 2025-12-22
  398. Golang optimizations for high‑volume services -- 2025-12-12
  399. PaCoRe: The first open-source deep think 8B model beats GPT-5 on HMMT25 -- 2025-12-11
  400. RnJ-1-Instruct FP8 Quantization -- 2025-12-10
  401. Optical Context Compression Is Just (Bad) Autoencoding -- 2025-12-10
  402. Masked Diffusion Models as Energy Minimization -- 2025-12-10
  403. Miles + FSDP2 = Megatron-Level Performance with More Flexibility -- 2025-12-10
  404. P4nda0s/IDA-NO-MCP -- 2025-12-09
  405. Toyota unintended acceleration and the big bowl of "spaghetti" code (2013) -- 2025-12-09
  406. https://huggingface.co/Doradus/Hermes-4.3-36B-FP8 -- 2025-12-09
  407. Support for rnj-1 now in llama.cpp -- 2025-12-09
  408. Comfy-Org/flux2-dev -- 2025-12-09
  409. baidu/ERNIE-4.5-VL-28B-A3B-Thinking -- 2025-12-09
  410. I built a personal assistant script, and the CPU inference speed beats my Llama setup. -- 2025-12-08
  411. Semantic Compression (2014) -- 2025-12-08
  412. A Deep Dive into Using PIO and DMA on the RP2350 -- 2025-12-05
  413. Free yourself from the Spotify desktop client with spotifyd -- 2025-12-04
  414. I cooked abliterated gemma3-27b-it with norm-preserving technique -- 2025-12-04
  415. Qwen3 VL built from scratch with PyTorch -- 2025-12-03
  416. EmbeddingGemma: Powerful and Lightweight Text Representations -- 2025-12-03
  417. Z-Image: Powerful and highly efficient image generation model with 6B parameters -- 2025-12-03
  418. RTX 5090 + Qwen 30B MoE @ 135 tok/s in NVFP4 - Full guide with C++ patches -- 2025-12-02
  419. [Editorial] https://huggingface.co/blog/grimjim/norm-preserving-biprojected-abliteration -- 2025-12-02
  420. Optimizing Token Generation in llama.cpp's CUDA Backend -- 2025-12-01
  421. [Editorial] https://arxiv.org/html/2511.09030v1 -- 2025-11-28
  422. You're using HuggingFace wrong. Stop downloading pre-quantized GGUFs and start building hardware-optimized, domain-specific models. Here's the pipeline I built to do it. -- 2025-11-26
  423. Binary Quantization For LLMs Through Dynamic Grouping -- 2025-11-26
  424. dx8152/Relight -- 2025-11-26
  425. Question About Motherboards -- 2025-11-26
  426. [Release] DragonMemory: 16× semantic compression for local RAG context (open-source, AGPL) -- 2025-11-25
  427. ORPO-Distill: Mixed-Policy Preference Optimization for Cross-Architecture LLM Distillation -- 2025-11-25
  428. Continuous batching from first principles -- 2025-11-25
  429. Can an expert chime in and explain what is holding Vulkan back from becoming the standard API for ML? -- 2025-11-25
  430. luozijian1990/network-traffic-ebpf-exporter -- 2025-11-24
  431. We found cryptography bugs in the elliptic library using Wycheproof -- 2025-11-24
  432. Show HN: Cynthia – Reliably play MIDI music files – MIT / Portable / Windows -- 2025-11-24
  433. Browser Fingerprinting and Why VPNs Won’t Make You Anonymous -- 2025-11-24
  434. [Release] Memory-Isolated Recursive Compression (MIRC). A local-first probabilistic compression utility for Apple Silicon. Research Preview (Open Source) -- 2025-11-21
  435. Read long podcasts locally with Whisper + LLM, open sourced -- 2025-11-21
  436. Local all-in-one AI system (Local multimodal AI) -- 2025-11-21
  437. JMS1717/8mb.local -- 2025-11-21
  438. Mimir Memory Bank now uses llama.cpp! -- 2025-11-21
  439. Quantum physicists have shrunk and "de-censored" DeepSeek R1 -- 2025-11-20
  440. Built a tool to solve the "how much GPU do I actually need?" problem for LLM deployment -- 2025-11-20
  441. New Parameter Browser added to Llamacpp Model Launcher! experimental model parameter tuning(window/cuda only) -- 2025-11-20
  442. cuda device list mismatch - ggml_cuda_init / ubuntu - significance to using --main-gpu flag -- 2025-11-20
  443. What Size of LLM Can 4x RTX 5090 Handle? (96GB VRAM) -- 2025-11-20
  444. Gain 60% performance on RDNA 4 using this fix -- 2025-11-19
  445. wildminder/ComfyUI-DyPE -- 2025-11-19
  446. lightx2v/Autoencoders -- 2025-11-19
  447. Scale-out is the silent killer of LLM applications. Are we solving the wrong problem? -- 2025-11-19
  448. PyTorch 2.10.0a0 w/ Blackwell (sm_120) Support — Patched &amp; Packaged for One-Command Install -- 2025-11-17
  449. Half-trillion parameter model on a machine with 128 GB RAM + 24 GB VRAM -- 2025-11-17
  450. [Editorial] https://www.linkedin.com/posts/andriyburkov_when-you-train-a-model-on-one-dataset-it-activity-7392804316769701888-166x/ -- 2025-11-14
  451. xRouter: Training Cost-Aware LLMs Orchestration System via Reinforcement Learning -- 2025-11-14
  452. [Editorial] Balancing order, freedom, and technology -- 2025-11-12
  453. AMD warns the Intel and Nvidia partnership is a risk to its business -- 2025-11-12
  454. A Pentium In Your Hand -- 2025-11-12
  455. [Editorial] https://www.linkedin.com/posts/ismaelvelasco_theres-an-ai-text-model-comparable-to-sota-activity-7393850964731912192-nZT1 -- 2025-11-12
  456. Last week in Multimodal AI - Local Edition -- 2025-11-12
  457. Apache Iggy is a high-performance, persistent message streaming platform -- 2025-11-07
  458. [P] Training Better LLMs with 30% Less Data – Entropy-Based Data Distillation -- 2025-11-06
  459. I fine tuned a (small) model to help with reasoning backfill on old/non-reasoning datasets -- 2025-11-06
  460. Superhuman AI for Multiplayer Poker -- 2025-11-06
  461. cerebras/GLM-4.5-Air-REAP-82B-A12B -- 2025-11-06
  462. Retrieval Enhanced Feedback via In-context Neural Error-book -- 2025-11-06
  463. Kimi release Kimi K2 Thinking, an open-source trillion-parameter reasoning model -- 2025-11-06
  464. OpenAI asks U.S. for loan guarantees to fund $1T AI expansion -- 2025-11-06
  465. Kimi released Kimi K2 Thinking, an open-source trillion-parameter reasoning model -- 2025-11-06
  466. [Editorial] Frequently wrong, but never in doubt’ -- 2025-11-05
  467. The Zero Freeze Formula: Teaching Local LLaMA Real Physics Through Python (SU(3) Mass Gap Simulation) to solve the Yang–Mills Mass Gap -- 2025-11-05
  468. Audio Sound Capture Project Needs Help -- 2025-11-05
  469. [D] It turns out WDDM driver mode is making our RAM - GPU transfer extremely slower compared to TCC or MCDM mode. Anyone has figured out the bypass NVIDIA software level restrictions? -- 2025-11-05
  470. GLaDOS TTS finetuning on MLX from the original game files -- 2025-11-04
  471. zeusftk/FTK_CANVAS_AGENT_for_Comfyui -- 2025-11-04
  472. guyyariv/DyPE -- 2025-11-04
  473. Qwen3-VL-32B Q8 speeds in llama.cpp vs vLLM FP8 on a RTX PRO 6000 -- 2025-11-03
  474. Help me decide: EPYC 7532 128GB + 2 x 3080 20GB vs GMtec EVO-X2 -- 2025-11-03
  475. amd/Nitro-E -- 2025-11-03
  476. CISA and NSA share tips on securing Microsoft Exchange servers -- 2025-11-02
  477. The Smol Training Playbook: The Secrets to Building World-Class LLMs -- 2025-11-02
  478. Latest Update from Anthropic's new model - Neptune V6 -- 2025-11-02
  479. AI "Phone Farm" Startup Gets Funding from Marc Andreessen to Flood Social Media With Spam -- 2025-11-02
  480. FlashPack: High-throughput tensor loading for PyTorch -- 2025-11-01
  481. M5 Neural Accelerator benchmark results from Llama.cpp -- 2025-11-01
  482. Kafka is Fast – I'll use Postgres -- 2025-11-01
  483. [Editorial] https://www.linkedin.com/posts/busiel-morley_economic-shifts-in-the-age-of-ai-ugcPost-7390349517612806144-8djS -- 2025-11-01
  484. Analog Surround Sound Was Everywhere, But You Probably Didn’t Notice -- 2025-11-01
  485. US Gas Turbine Shortage Likely to Slow AI Demand Growth -- 2025-10-31
  486. The Supercon 2025 Badge is Built to be Customized -- 2025-10-31
  487. Experimenting with Qwen3-VL for Computer-Using Agents -- 2025-10-30
  488. Built a full voice AI assistant running locally on my RX 6700 with Vulkan - Proof AMD cards excel at LLM inference -- 2025-10-30
  489. Streaming datasets: 100x More Efficient -- 2025-10-30
  490. Cerebras REAP'd GLM4.6: 25%, 30%, 40% pruned FP8 checkpoints on HF! -- 2025-10-28
  491. Qwen/Qwen3-VL-30B-A3B-Instruct-FP8 -- 2025-10-28
  492. lightx2v/Wan2.2-Distill-Loras -- 2025-10-28
  493. [Editorial] Periodic table for ai algorithms -- 2025-10-26
  494. Need help understanding OpenAIs API usage for text-embedding -- 2025-10-26
  495. Qwen3 Next support in llama.cpp ready for review -- 2025-10-25
  496. GLM Air REAP tool call problems -- 2025-10-25
  497. Reverse Engineering STL Files with FreeCAD -- 2025-10-25
  498. Un-LOCC (Universal Lossy Optical Context Compression), Achieve Up To 3× context compression with 93.65% Accuracy. -- 2025-10-24
  499. LiquidAI/LFM2-1.2B-RAG -- 2025-10-24
  500. zai-org/GLM-4.6 -- 2025-10-24
  501. inference-net/Schematron-3B -- 2025-10-21
  502. [By GLM Team] Glyph: Scaling Context Windows via Visual-Text Compression -- 2025-10-21
  503. DGX SPARK Compiled llama.cpp Benchmarks Compared to M4 MAX (non-MLX) -- 2025-10-21
  504. perplexityai/search_evals -- 2025-10-21
  505. Hetzner: The Simple Cloud just got more flexible and more affordable -- 2025-10-21
  506. A new, super simple LLM benchmark for testing changes across models, quants, parameters, samplers, engines, etc -- 2025-10-21
  507. riptideslabs/tokenex -- 2025-10-20
  508. Multi-Tenant SaaS's Wildcard TLS: An Overview of DNS-01 Challenges -- 2025-10-20
  509. From cloud to OCP? Be ready to wrangle firmware -- 2025-10-20
  510. FLOSS Weekly Episode 851: Buckets of Money -- 2025-10-20
  511. Significant speedup for local models -- 2025-10-20
  512. Cursor tricking paid users with fake Claude Sonnet 4.5 -- 2025-10-20
  513. inclusionAI/Ring-1T -- 2025-10-20
  514. volantvm/volant -- 2025-10-19
  515. Wireshark 4.6.0 Supports macOS Pktap Metadata (PID, Process Name, etc.) -- 2025-10-19
  516. A classified network of SpaceX satellites is emitting a mysterious signal -- 2025-10-19
  517. linkedlist771/SoraWatermarkCleaner -- 2025-10-19
  518. Qwen/Qwen-Image-Edit-2509 -- 2025-10-19
  519. The Entire Process of Building an Open Source Analog ASIC -- 2025-10-15
  520. Built a 1288x RTFx Parakeet Speech-to-Text server... Enjoy! -- 2025-10-13
  521. Novel OpenGL Pixel Shader Dewarping -- 2025-10-13
  522. lovis93/next-scene-qwen-image-lora-2509 -- 2025-10-13
  523. Beyond Token Count: Our Research Suggests "Contextual Weight" is a Key Limiter on Large Context Windows -- 2025-10-13
  524. FractalAIResearch/Fathom-Search-4B -- 2025-10-13
  525. LLM Robustness Leaderboard v1 --Technical report -- 2025-10-13
  526. Hide and Seek with LLMs: An Adversarial Game for Sneaky Error Generation and Self-Improving Diagnosis -- 2025-10-13
  527. Preference optimization with ORPO and LoRA -- 2025-10-12
  528. [Show] SpiralTorch: A Rust-based PyTorch-style autograd engine (Python 3.14-ready) -- 2025-10-12
  529. Qwen3-VL-30B-A3B-Thinking GGUF with llama.cpp patch to run it -- 2025-10-10
  530. Did anyone try out GLM-4.5-Air-GLM-4.6-Distill ? -- 2025-10-10
  531. What and when 7900xtx is boosted? -- 2025-10-10
  532. Modelfile. Do I need these tags PER prompt? -- 2025-10-10
  533. Divining Air Quality With A Cheap Computer Vision Device -- 2025-10-09
  534. Awesome Local LLM Speech-to-Speech Models &amp; Frameworks -- 2025-10-08
  535. FabioSarracino/VibeVoice-Large-Q8 -- 2025-10-08
  536. CapRecover: A Cross-Modality Feature Inversion Attack Framework on Vision Language Models -- 2025-10-08
  537. Mitigating Watermark Stealing Attacks in Generative Models via Multi-Key Watermarking -- 2025-10-08
  538. How to make the AI Bot to understand the exact design and App flow -- 2025-10-04
  539. Creating a Full Stack App W/Cloudflare Works and BetterAuth -- 2025-10-04
  540. Comprehension debt: A ticking time bomb of LLM-generated code -- 2025-10-04
  541. deepseek-ai/DeepSeek-V3.2-Exp -- 2025-10-03
  542. moondream/moondream3-preview -- 2025-10-03
  543. [Editorial] https://github.com/emcie-co/parlant -- 2025-10-02
  544. Built a persistent memory system for LLMs - 3 months testing with Claude/Llama -- 2025-10-02
  545. Do I need to run /init on a repo if I already have AGENTS.md? -- 2025-10-02
  546. Inside NVIDIA GPUs: Anatomy of high performance matmul kernels -- 2025-09-29
  547. Bit is all we need: binary normalized neural networks -- 2025-09-29
  548. Accelerating Qwen3-8B Agent on Intel® Core™ Ultra with Depth-Pruned Draft Models -- 2025-09-29
  549. This $5,999 RTX PRO 6000 Ebay listing is a scam, right? -- 2025-09-26
  550. Accelerating Local AI on Consumer GPUs: A Hardware-Aware Dynamic Strategy for YOLOv10s -- 2025-09-26
  551. Efficient 4B parameter gpt OSS distillation without the over-censorship -- 2025-09-22
  552. [Project] I created an AI photo organizer that uses Ollama to sort photos, filter duplicates, and write Instagram captions. -- 2025-09-22
  553. Pointer Tagging in C++: The Art of Packing Bits into a Pointer -- 2025-09-22
  554. inclusionAI/Ring-mini-2.0 -- 2025-09-22
  555. Local real-time assistant that remembers convo + drafts a doc -- 2025-09-22
  556. XiaomiMiMo/MiMo-Audio-7B-Instruct -- 2025-09-21
  557. Scaling Self-Supervised Representation Learning for Symbolic Piano Performance -- 2025-09-21
  558. Uncensor Qwen3 models without retraining -- 2025-09-20
  559. Depth upscaling? -- 2025-09-20
  560. Qwen3‑Next‑80B‑A3B‑Instruct (FP8) on Windows 11 WSL2 + vLLM + Docker (Blackwell) -- 2025-09-19
  561. unsloth/Qwen3-Next-80B-A3B-Instruct -- 2025-09-19
  562. The AI-Scraping Free-for-All Is Coming to an End -- 2025-09-18
  563. Visible Watermarking with Gradio -- 2025-09-18
  564. xiaomi-research/q-frame -- 2025-09-17
  565. google/embeddinggemma-300m -- 2025-09-17
  566. 3-month Claude Code Max user review - considering alternatives -- 2025-09-15
  567. Chesars/whatsapp-mcp -- 2025-09-15
  568. Claude’s memory architecture is the opposite of ChatGPT’s -- 2025-09-15
  569. The Internet Will Be More Dead Than Alive Within 3 Years, Trend Shows | All signs point to a future internet where bot-driven interactions far outnumber human ones. -- 2025-09-15
  570. New "speech" mode in Imagine... -- 2025-09-15
  571. I made local RAG, web search, and voice mode on iPhones completely open source, private, and free -- 2025-09-08
  572. jwest33/jam_model_memory -- 2025-09-08
  573. How was your experience with Claude vs Codex? -- 2025-09-08
  574. [Project/Code] Fine-Tuning LLMs on Windows with GRPO + TRL -- 2025-09-07
  575. nvidia/NVIDIA-Nemotron-Nano-12B-v2-Base -- 2025-09-07
  576. huihui-ai/Huihui-gpt-oss-20b-BF16-abliterated -- 2025-09-07
  577. RX570 compatibility issues -- 2025-09-07
  578. Continue.dev setup -- 2025-09-07
  579. Little SSM (RWKV7 7B) state checkpointing demo. -- 2025-09-04
  580. Need advice on how to get VLLM working with 2xR9700 + 2x7900xtx? -- 2025-09-04
  581. pwnfuzz/diffrays -- 2025-09-04
  582. Chromium Hardening Guide -- 2025-09-04
  583. roomkangali/dursgo -- 2025-09-04
  584. QuEST/Quartet authors discuss their work on SOTA 4-bit training optimizations -- 2025-09-01
  585. F-Stack – A network development kit with high performance based on DPDK -- 2025-09-01
  586. An Empirical Study of Knowledge Distillation for Code Understanding Tasks -- 2025-09-01
  587. A Comparative Analysis of Vision Language Models for Scientific Data Interpretation -- 2025-08-31
  588. Sparrow: Custom language model architecture for microcontrollers like the ESP32 -- 2025-08-30
  589. Password only for this week: Welcome to Hugston -- 2025-08-26
  590. Prism MCP Rust SDK v0.1.0 - Production-Grade Model Context Protocol Implementation -- 2025-08-26
  591. Compute Where It Counts: High Quality Sparsely Activated LLMs -- 2025-08-25
  592. moonshotai/Kimi-K2-Base -- 2025-08-25
  593. unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF -- 2025-08-25
  594. BlueLM-2.5-3B Technical Report -- 2025-08-25
  595. lightx2v/Qwen-Image-Lightning -- 2025-08-25
  596. Made Chatterbox TTS a bit faster again on CUDA (155it/s on 3090) -- 2025-08-25
  597. KittenML/KittenTTS -- 2025-08-25
  598. city96/Qwen-Image-gguf -- 2025-08-23
  599. Menlo/Lucy-128k -- 2025-08-23
  600. NVIDIA just accelerated output of OpenAI’s gpt-oss-120B by nearly 2x -- 2025-08-23
  601. COMponent-Aware Pruning for Accelerated Control Tasks in Latent Space Models -- 2025-08-23
  602. Speculative decoding in archgw candidate release 0.4.0. Could use feedback, -- 2025-08-16
  603. Nvidia Tilus: A Tile-Level GPU Kernel Programming Language -- 2025-08-16
  604. SynapseRoute: An Auto-Route Switching Framework on Dual-State Large Language Model -- 2025-08-16
  605. New Tool for Finding Why Your LLM Inference is Slow -- 2025-08-14
  606. I ran OpenAI’s GPT-OSS 20B locally on a 16GB Mac with Ollama — setup, gotchas, and mini demo -- 2025-08-14
  607. GLM 4.5 Air - Optimizing - Vulkan vs. CUDA? -- 2025-08-14
  608. GLiNER2: An Efficient Multi-Task Information Extraction System with Schema-Driven Interface -- 2025-08-13
  609. TextQuests: How Good are LLMs at Text-Based Video Games? -- 2025-08-13
  610. Mitigate Hallucinations by Fine-tuning gpt-oss-120b with One Example -- 2025-08-10
  611. uncensored gpt-oss-20b, bf16 and mxfp4 both available -- 2025-08-10
  612. LGAI-EXAONE/EXAONE-4.0-1.2B -- 2025-08-10
  613. New Open-Source Text-to-Image Model Just Dropped Qwen-Image (20B MMDiT) by Alibaba! -- 2025-08-10
  614. Kitten TTS Web Demo -- 2025-08-09
  615. Show HN: I built a tool to replace capcut audio transcription -- 2025-08-09
  616. Whispers From The Void, Transcribed With AI -- 2025-08-09
  617. The Tape Speed Keyboard -- 2025-08-08
  618. 0.82 um 105 W diode-pumped thulium-doped all silica fiber laser -- 2025-08-06
  619. GLM-4.5 llama.cpp PR is nearing completion -- 2025-08-05
  620. glm-4.5-Air appreciation poist - if you have not done so already, give this model a try -- 2025-08-05
  621. Amazon's AI Coding Revealed a Dirty Little Secret -- 2025-08-02
  622. On the Interaction of Compressibility and Adversarial Robustness -- 2025-08-02
  623. realtime-ai/blastoff-llm -- 2025-08-02
  624. Quantize your own GGUFs the same way as your fav Unsloth Dynamic GGUFs -- 2025-08-01
  625. unsloth/Qwen3-235B-A22B-Instruct-2507-GGUF -- 2025-08-01
  626. On the Predictive Power of Representation Dispersion in Language Models -- 2025-08-01
  627. Wan 2.2 T2V,I2V 14B MoE Models -- 2025-07-31
  628. PowerInfer/SmallThinker-21BA3B-Instruct -- 2025-07-31
  629. Ollama + Open WebUI -- is there a way for the same query to run through the same model multiple times (could be 3 times, could be 100 times), then gather all the answers together to summarise/count? -- 2025-07-25
  630. WGRAMMAR: Leverage Prior Knowledge to Accelerate Structured Decoding -- 2025-07-25
  631. Semantic chunking using LLMs -- 2025-07-20
  632. Does the OpenWebUi run the sentence transformer models locally? -- 2025-07-20
  633. Dataset for structured (JSON) output? -- 2025-07-19
  634. support for Kimi-K2 has been merged into llama.cpp -- 2025-07-19
  635. t-tech/T-pro-it-2.0 -- 2025-07-19
  636. Madness, the ignorant's question. Would it be possible to lighten an LLM model? -- 2025-07-18
  637. ETH Zurich and EPFL will release a fully open-source LLM developed on public infrastructure. Trained on the “Alps” supercomputer at the Swiss National Supercomputing Centre (CSCS). Trained on 60% english/40% non-english, it will be released in 8B and 70B sizes. -- 2025-07-17
  638. Moonshot AI’s open source Kimi K2 outperforms GPT-4 in key benchmarks -- 2025-07-17
  639. Advice Needed: Best way to replace Together API with self-hosted LLM for high-concurrency app -- 2025-07-17
  640. baidu/ERNIE-4.5-0.3B-PT -- 2025-07-17
  641. LiquidAI/LFM2-700M -- 2025-07-17
  642. RekaAI/reka-flash-3.1 · Hugging Face -- 2025-07-17
  643. Seq vs Seq: the Ettin Suite of Paired Encoders and Decoders -- 2025-07-17
  644. T5Gemma: A new collection of encoder-decoder Gemma models- Google Developers Blog -- 2025-07-17
  645. H-Net: a hierarchical network that replaces tokenization with a dynamic chunking process directly inside the model, automatically discovering and operating over meaningful units of data -- 2025-07-17
  646. How I build software quickly -- 2025-07-16
  647. RekaAI/reka-flash-3.1 -- 2025-07-15
  648. What kind of throughput can I expect with Llama 3.1 on a H200? -- 2025-07-15
  649. MIDI-VALLE: Improving Expressive Piano Performance Synthesis Through Neural Codec Language Modelling -- 2025-07-14
  650. Replication of Quantum Factorisation Records with an 8-bit Home Computer [pdf] -- 2025-07-14
  651. Local llms works great! -- 2025-07-12
  652. LiquidAI/LFM2-350M -- 2025-07-12
  653. QPART: Adaptive Model Quantization and Dynamic Workload Balancing for Accuracy-aware Edge Inference -- 2025-07-12
  654. Issues with Qwen 3 Embedding models (4B and 0.6B) -- 2025-07-12
  655. [Tool Release] Finetune &amp; Quantize 1–3B LLMs on 8GB RAM using LoFT CLI (TinyLlama + QLoRA + llama.cpp) -- 2025-07-11
  656. Qwen3-8B-BitNet -- 2025-07-11
  657. Megakernel doubles Llama-1B inference speed for batch size 1 -- 2025-07-09
  658. Smallest &amp; best OCR model that can read math &amp; code? -- 2025-07-07
  659. Qwen/WorldPM-72B -- 2025-07-07
  660. Code single file with multiple LLM models -- 2025-07-07
  661. Gen-Verse/CURE -- 2025-07-07
  662. Run Deepseek locally on a 24g GPU: Quantizing on our Giga Computing 6980P Xeon -- 2025-07-01
  663. I built a document workflow system using VLMs: processes complex docs end-to-end (runs locally!!) -- 2025-07-01
  664. Jan Nano + Deepseek R1: Combining Remote Reasoning with Local Models using MCP -- 2025-07-01
  665. Query Classifier for RAG - Save your $$$ and users from irrelevant responses -- 2025-07-01
  666. Building a memory-heavy AI agent — looking for local-first storage & recall solutions -- 2025-07-01
  667. Is there any easy way to get up and running with chatgpt-like capabilities at home? -- 2025-07-01
  668. No recognition of slavic characters. English characters recognized are separate singular characters, not a block of text when using PaddleOCR. -- 2025-07-01
  669. Tired of copy-pasting from ChatGPT for coding? I am building an open-source tool (Athanor) to fix that - Alpha testers/feedback wanted! -- 2025-07-01
  670. VideoGameBench from Princeton: Can vision-language models play 90s video games? -- 2025-07-01
  671. New band surges to 500k listeners on Spotify, but turns out it's AI slop -- 2025-07-01
  672. How we cut CKEditor's bundle size by 40% -- 2025-07-01
  673. My VSCode → AI chat website connector extension just got 3 new features! -- 2025-07-01
  674. 100 Gbps Indoor Access and 4.8 Gbps Outdoor Point-to-Point LiFi Transmission Systems using Laser-based Light Sources -- 2025-06-30
  675. (0,4) brane box models -- 2025-06-30
  676. Cursor 1.0 -- 2025-06-30
  677. Help me design a robust on-prem Llama 3 70B infrastructure for 30 users – Complete hardware/software list wanted -- 2025-06-30
  678. Jan-nano, a 4B model that can outperform 671B on MCP -- 2025-06-30
  679. Models that are good and fast at Long Document Processing -- 2025-06-30
  680. I am making an AI batteries included Web Framework (like Django but for AI) -- 2025-06-30
  681. [New Features & Better] Tabulens: A Vision-LLM Powered PDF Table Extractor -- 2025-06-30
  682. I tested 10 LLMs locally on my MacBook Air M1 (8GB RAM!) – Here's what actually works- -- 2025-06-30
  683. Chatbot without ChatGPT -- 2025-06-30
  684. What's the best way to save and manage different text files for the models to reference? PRD, cursor rules, tech stack, design reference, etc? -- 2025-06-30
  685. Bzip2 crate switches from C to 100% Rust -- 2025-06-30
  686. Litestream: Revamped -- 2025-06-30
  687. Stop using REST for state synchronization (2024) -- 2025-06-30
  688. Announcing `mcp-protocol-sdk`: A New Enterprise grade Rust SDK for AI Tool Calling (Model Context Protocol) -- 2025-06-30
  689. Reinforcement Pre-Training -- 2025-06-29
  690. unsloth/gemma-3n-E4B-it-GGUF -- 2025-06-29
  691. chandar-lab/NeoBERT -- 2025-06-29
  692. tencent/Hunyuan-A13B-Instruct -- 2025-06-27
  693. maya-research/Veena -- 2025-06-27
  694. Meet Mistral Devstral, SOTA open model designed specifically for coding agents -- 2025-06-26
  695. 1.93bit Deepseek R1 0528 beats Claude Sonnet 4 -- 2025-06-26
  696. DeepSeek R1 05/28 performance on five independent benchmarks -- 2025-06-26
  697. Few-Shot Examples: Overfitting / Leakage -- 2025-06-26
  698. Finetune a model to think and use tools -- 2025-06-26
  699. I need help using open web UI with Ollama. Help installing and getting it running win 11 -- 2025-06-26
  700. I built/am building a micro-transformer for learning and experimentation -- 2025-06-26
  701. I shipped more code yesterday with Claude 4 than the last 3 weeks combined -- 2025-06-26
  702. A deep dive into self-improving AI and the Darwin-Gödel Machine -- 2025-06-26
  703. 100% of the zeros of the Riemann zeta-function are on the critical line -- 2025-06-25
  704. 100% of odd hyperelliptic Jacobians have no rational points of small height -- 2025-06-25
  705. deepseek-ai/DualPipe -- 2025-06-23
  706. 100 Particles Quantum Heat Engine: Exploring the Impact of Criticality on Efficiency -- 2025-06-23
  707. 0-Auslander correspondence -- 2025-06-23
  708. Advanced Time Manipulation with GDB -- 2025-06-21
  709. Practical SDR: Getting started with software-defined radio -- 2025-06-21
  710. 1000-10,000 M$_\odot$ Primordial Stars Created the Nitrogen Excess in the Galaxy GS 3073 at $z = 5.55$ -- 2025-06-21
  711. $0^+$ to $2^+$ neutrinoless double-$β$ decay of $^{76}$Ge, $^{82}$Se, $^{130}$Te and $^{136}$Xe in the microscopic interacting boson model} -- 2025-06-21
  712. 0-1 laws for pattern occurrences in phylogenetic trees and networks -- 2025-06-20
  713. 100ps time resolution with thin silicon pixel detectors and a SiGe HBT amplifier -- 2025-06-18
  714. 0-$\pi$ quantum transition in a carbon nanotube Josephson junction: universal phase dependence and orbital degeneracy -- 2025-06-18
  715. openbmb/MiniCPM4-8B -- 2025-06-17
  716. lym00/Wan2.1-T2V-1.3B-Self-Forcing-VACE-Addon-Experiment -- 2025-06-17
  717. DeepSeek R1 05 28 Tested. It finally happened. The ONLY model to score 100% on everything I threw at it. -- 2025-06-17
  718. ubergarm/DeepSeek-R1-0528-GGUF -- 2025-06-17
  719. LLM training on RTX 5090 -- 2025-06-17
  720. [DEMO] I created a coding agent that can do dynamic, runtime debugging. -- 2025-06-17
  721. Is anyone productively using Aider and Ollama together? -- 2025-06-17
  722. For everyone who's still confused by Attention... I made this spreadsheet just for you(FREE) -- 2025-06-17
  723. What setup/model do you use and what’s your monthly spend? -- 2025-06-17
  724. Xiaomi released an updated 7B reasoning model and VLM version claiming SOTA for their size -- 2025-06-17
  725. UPDATE: Inference needs nontrivial amount of PCIe bandwidth (8x RTX 3090 rig, tensor parallelism) -- 2025-06-16
  726. IQ1_Smol_Boi -- 2025-06-16
  727. Qwen releases official MLX quants for Qwen3 models in 4 quantization levels: 4bit, 6bit, 8bit, and BF16 -- 2025-06-16
  728. Seeking Help Setting Up a Local LLM Assistant for TTRPG Worldbuilding + RAG on Windows 11 -- 2025-06-16
  729. New VS Code Pair Programming Extension, Need Help Testing -- 2025-06-16
  730. Claude-Trace -- 2025-06-16
  731. Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training -- 2025-06-16
  732. best fine tuned local LLM for Github Copilot Agent specificaly -- 2025-06-16
  733. A Simulation in C++ of Joseph Weizenbaum's 1966 Eliza -- 2025-06-16
  734. 0D-2D Heterostructure for making very Large Quantum Registers using itinerant Bose-Einstein Condensate of Excitons -- 2025-06-16
  735. 100-mJ class, sub-two-cycle, carrier-envelope phase-stable dual-chirped optical parametric amplification -- 2025-06-16
  736. [Update] Rensa: added full CMinHash + OptDensMinHash support (fast MinHash in Rust for dataset deduplication / LLM fine-tuning) -- 2025-06-15
  737. Open Source Unsiloed AI Chunker (EF2024) -- 2025-06-15
  738. ether0 - Mistral 24B with RL on several molecular design tasks in chemistry -- 2025-06-15
  739. Need selfhosted AI to generate better bash scripts and ansible playbooks -- 2025-06-15
  740. How do I finetune Devstral with vision support? -- 2025-06-15
  741. What's the best approach for including niche dependency source files and associated documentation reference material in context? -- 2025-06-15
  742. Airlines Don't Want You to Know They Sold Your Flight Data to DHS -- 2025-06-15
  743. John Deere Must Face Second Right to Repair Lawsuit -- 2025-06-15
  744. What vector database and embeddings are y'all using -- 2025-06-15
  745. Turn based two model critique for rounds to refine answer - any examples or FOSS projects? -- 2025-06-15
  746. mistralai/Magistral-Small-2506_gguf -- 2025-06-14
  747. Ruminate: From All-or-Nothing to Just-Right Reasoning in LLMs -- 2025-06-14
  748. [update] Restructured repo under rvn-tools — modular CLI for LLM formats -- 2025-06-14
  749. Testing Quant Quality for Shisa V2 405B -- 2025-06-14
  750. Old model, new implementation -- 2025-06-14
  751. Ollama vs Llamacpp: Different output for same model -- 2025-06-14
  752. How to improve my ViT model -- 2025-06-14
  753. From RPC to transactions and durable executions -- 2025-06-14
  754. Flattening Rust’s learning curve -- 2025-06-14
  755. Async from scratch 3: Pinned against the wall -- 2025-06-14
  756. How to get the most out of my AMD 7900XT? -- 2025-06-14
  757. Faulty 120W charger analysis (Anker GAN Prime) [video] -- 2025-06-13
  758. unsloth/Magistral-Small-2506-GGUF -- 2025-06-12
  759. mistralai/Magistral-Small-2506 -- 2025-06-12
  760. rednote-hilab/dots.llm1.base -- 2025-06-12
  761. [Tool] rvn-convert: OSS Rust-based SafeTensors to GGUF v3 converter (single-shard, fast, no Python) -- 2025-06-12
  762. GuidedQuant: Boost LLM layer-wise PTQ methods using the end loss guidance (Qwen3, Gemma3, Llama3.3 / 2~4bit Quantization) -- 2025-06-12
  763. I built a memory MCP that understands you (so Sam Altman can't). -- 2025-06-12
  764. Built an open source desktop app to easily play with local LLMs and MCP -- 2025-06-12
  765. mtmd : support Qwen 2.5 Omni (input audio+vision, no audio output) by ngxson · Pull Request #13784 · ggml-org/llama.cpp -- 2025-06-12
  766. i got tired of the errors, so automated debugging using Ollama -- 2025-06-12
  767. Ablating Gemma 3 27B variants with synthetic data from Sonnet 4 (Few-shot vs LoRA) -- 2025-06-12
  768. The LLM Gateway gets a major upgrade: becomes a data-plane for Agents. -- 2025-06-12
  769. Introducing stronger dependencies on systemd -- 2025-06-12
  770. How we decreased GitLab repo backup times from 48 hours to 41 minutes -- 2025-06-12
  771. The Quest for 100k - LLAMA.CPP Setting for a Noobie -- 2025-06-12
  772. Clipjacking: Hacked by copying text – Clickjacking but better -- 2025-06-11
  773. 0/1 Deep Neural Networks via Block Coordinate Descent -- 2025-06-10
  774. turbulentdrom/sing-srs-converter -- 2025-06-10
  775. abi/screenshot-to-code -- 2025-06-10
  776. Qwen/Qwen3-Embedding-4B -- 2025-06-10
  777. 100 Gbps Quantum-safe IPsec VPN Tunnels over 46 km Deployed Fiber -- 2025-06-09
  778. Qwen/Qwen3-Embedding-0.6B -- 2025-06-08
  779. 0-$π$ qubit in one Josephson junction -- 2025-06-07
  780. 100-kT Magnetic field generation using paisley targets by femtosecond laser-plasma interactions -- 2025-06-07
  781. 100 GHz Micrometer compact broadband Monolithic ITO Mach Zehnder Interferometer Modulator enabling 3500 times higher Packing Density -- 2025-06-06
  782. 0-$\pi$ phase-controllable $thermal$ Josephson junction -- 2025-06-06
  783. Precomputing Transparency Order in 3D -- 2025-06-06
  784. ban6cat6/aparecium -- 2025-06-03
  785. 0-Gaps on 3D Digital Curves -- 2025-06-03
  786. Reports of Deno's Demise Have Been Greatly Exaggerated -- 2025-06-02
  787. Comparing Parallel Functional Array Languages: Programming and Performance -- 2025-06-02
  788. What Every Programmer Should Know About Enumerative Combinatorics -- 2025-06-02
  789. DuckLake: SQL as a Lakehouse Format -- 2025-05-31
  790. 1000x Faster Camera and Machine Vision with Ordinary Devices -- 2025-05-31
  791. 0.75 Gbit/s high-speed classical key distribution with mode-shift keying chaos synchronization of Fabry-Perot lasers -- 2025-05-31
  792. EdinburghNLP/MMLongBench -- 2025-05-29
  793. Show HN: Model2vec-Rs – Fast Static Text Embeddings in Rust -- 2025-05-29
  794. unsloth/DeepSeek-R1-0528-GGUF -- 2025-05-29
  795. QuantStack/Wan2.1-VACE-14B-GGUF -- 2025-05-29
  796. 100,000 frames-per-second compressive imaging with a conventional rolling-shutter camera by random point-spread-function engineering -- 2025-05-29
  797. 1,000-Fold Enhancement of Light-Induced Magnetism in Plasmonic Au Nanoparticles -- 2025-05-29
  798. nvidia/Llama-3.1-Nemotron-Nano-4B-v1.1 -- 2025-05-28
  799. Tongyi-Zhiwen/QwenLong-L1-32B -- 2025-05-28
  800. PKU-DS-LAB/FairyR1-32B -- 2025-05-28
  801. google/medgemma-4b-pt -- 2025-05-28