Quantization & Efficiency

Model compression, GGUF, efficient inference, optimization

660 articles across 177 editions

Articles

  1. [Editorial] Unsloth Dynamic 3.0 GGUFs -- 2026-08-20
  2. LFM2.5 Q4_0 Checkpoints from Quantization-Aware Distillation -- 2026-08-20
  3. [2511.07885] Intelligence per Watt: Measuring Intelligence Efficiency of Local AI -- 2026-08-20
  4. Transformer-Guided Swarm Intelligence for Frugal Neural Architecture Search -- 2026-08-20
  5. 0xSero/deepseek-v4-flash-0731-spark-sparkinfer -- 2026-08-20
  6. [Editorial] Deepseek-API -- 2026-08-20
  7. Linguistic Firewall: Geometry as Defense in Multi-Agent Systems Routing -- 2026-08-19
  8. Random Logit Scaling: Defending Deep Neural Networks Against Black-Box Score-Based Adversarial Example Attacks -- 2026-08-19
  9. bitsandbytes creator teasing new quantization method: GLM 5.3 on a single DGX Spark at 7t/s -- 2026-08-18
  10. GLM 5.3 weights. It might offer the best capacity-to-size ratio. -- 2026-08-18
  11. Ling 3.0 support merged into llama.cpp -- 2026-08-18
  12. Why not? ☺️ -- 2026-08-18
  13. We quantized DeepSeek V4 0731 and benchmarked it against popular quants on 8× RTX 5090 -- 2026-08-14
  14. Gemma 4 12B Q3: +8.55% Coding Performance From Tensor-Level Quantization Allocation -- 2026-08-14
  15. I trained a 1B-parameter LLM from scratch on 20B tokens for about $200 -- 2026-08-14
  16. Showoff Saturday: Local 4x 6000 Pro (multi-year progression) -- 2026-08-14
  17. [Editorial] Foxflow Fleet (probagi.com) -- 2026-08-14
  18. State of Open Models: Summer 2026 Observations -- 2026-08-14
  19. Qwen 3.8 27B is out : open weights, best local dense model yet -- 2026-08-14
  20. CohereLabs/North-Micro-Vision-Instruct · Hugging Face -- 2026-08-14
  21. U.S. Department of Energy Launches the Genesis Open Models Initiative and, with Arcee, Unveils Genesis-Science-1 — Its First Open-Weight Model for Scientific Research -- 2026-08-14
  22. enabling PCI-E p2p for consumer Nvidia cards will yield you more than you think -- 2026-08-12
  23. Rumored 50-series Super refresh bumps everything +50% VRAM -- 2026-08-12
  24. Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning -- 2026-08-12
  25. [2606.05682] Beyond Output Matching: Preserving Internal Geometry in NVFP4 LLM Distillation -- 2026-08-12
  26. torchtune: PyTorch native post-training library -- 2026-08-12
  27. Qwen3.8-2.4T-A95B Released -- 2026-08-12
  28. Exact Qwen 3.8 27b release date and time -- 2026-08-12
  29. inclusionAI/Ling-2.6-1T -- 2026-08-12
  30. nvidia/Nemotron-Cascade-2-30B-A3B -- 2026-08-12
  31. DiffusionGemma Technical Report -- 2026-08-12
  32. LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge -- 2026-08-12
  33. llama.cpp -- 2026-08-12
  34. [Editorial] Anthropic's Invisible Watermarks in Claude Text -- 2026-08-11
  35. [Editorial] Reuven Cohen: Quantum Mechanics May Be Telling Us Something -- 2026-08-11
  36. Muse Glimmer 30B + DFlash speculative decoding on vLLM: 6 patches needed, 25 → 57 tok/s. Dockerfile and numbers inside. -- 2026-08-11
  37. Apple Silicon and macOS VMs: 11–16× Faster LLM Inference with Llama.cpp -- 2026-08-11
  38. Echo Dot 2 can run 28M LLM at decent speed -- 2026-08-11
  39. Chunked KL loss for running Knowledge Distillation locally (<6GB VRAM at 32K context length) -- 2026-08-11
  40. Sonic Pi v5 -- 2026-08-11
  41. MiniMax-AI/MiniMax-H3 -- 2026-08-11
  42. Glimmer seems pretty censored? -- 2026-08-11
  43. No wonder Qwen and Gemma are so different -- 2026-08-11
  44. CJK Manga/Manhwa/Manhua 150M OCR model (hayai-ocr-v2) outperforming PaddleOCR-VL-For-Manga -- 2026-08-11
  45. [Editorial] -- 2026-08-10
  46. [Editorial] -- 2026-08-10
  47. [Editorial] -- 2026-08-10
  48. DeepSeek V4 Flash 2-bit quant achieves 100% on SQL benchmark locally -- 2026-08-07
  49. Scotoma-2: Gemma4, but with less annoying slop and better writing -- 2026-08-07
  50. Mach-1 Additive: 95% of Qwen 3.6 35B performance while 10x smaller -- 2026-08-07
  51. Intern S2 Mobius — Qwen3.5-35B derivative with architectural throughput gains -- 2026-08-07
  52. [Update] DeepSeek-V4-Flash-0731 on a single RTX 5090: phase-adaptive DSpark K1/K2 with dual CUDA graphs -- 2026-08-05
  53. TensorSharp DSpark Benchmark on DeepSeek V4 Flash — up to 2x speedup -- 2026-08-05
  54. DeepSeek V4 Flash on a single RTX 4090 at 64k context — complete config and measured numbers -- 2026-08-05
  55. DeepSeek v4 Flash vs. Qwen3.6-27B, 3.5-122B, and Gemma 4 31B — Local Agentic Coding Benchmark -- 2026-08-05
  56. Vacuum 16T: A 16.5-trillion-parameter model that contains nothing -- 2026-08-03
  57. I benchmarked classic vector RAG vs Google's new OKF format vs both combined -- 2026-08-03
  58. [Editorial] Awesome Systematic Trading -- 2026-08-03
  59. [Editorial] -- 2026-07-30
  60. What do we know about the "AI Accelerators" used to train LongCat-2? -- 2026-07-24
  61. Running a 13M ASR conformer on a microcontroller -- 2026-07-24
  62. Running Qwen 3.6 35B MoE (Q4_K_M) on a Zeus (Xiaomi 12 Pro, 12GB RAM) -- 2026-07-24
  63. Squeeze-Release: Iterative Pruning with Exact Structural Minimization -- 2026-07-24
  64. The Open Weight Unbundling -- 2026-07-24
  65. Model "distillation" accusations are getting way overblown at this point -- 2026-07-24
  66. The Distillation Claims Are Fake and Desperate -- 2026-07-24
  67. I gave Kimi K3 a shot at auditing my post-quantum crypto project, it found 5 real bugs Fable/Opus 4.8 and GPT-5.6 Sol had all missed -- 2026-07-24
  68. China's open-weights AI strategy is winning -- 2026-07-22
  69. How do we benefits from 2+ T models? -- 2026-07-22
  70. Serving a fleet of Qwen3.5 122b sessions on a single Mac Studio (96GB) without losing your sanity -- 2026-07-22
  71. Quant Convergence: Bridging Classical Value Investing and Modern Factor Models for Systematic Equity Selection -- 2026-07-21
  72. [Editorial] -- 2026-07-21
  73. [Editorial] -- 2026-07-21
  74. OpenAI had to pause an unreleased model after it escaped containment. -- 2026-07-21
  75. [Editorial] -- 2026-07-21
  76. [Editorial] -- 2026-07-16
  77. [Editorial] -- 2026-07-16
  78. [Editorial] -- 2026-07-16
  79. [Editorial] -- 2026-07-16
  80. FT: Companies Turn to Chinese Open Weight Models to Cut Costs -- 2026-07-14
  81. PrismML Compresses Qwen-3.6-27B to Under 4GB — Runs on iPhone 17 Pro -- 2026-07-14
  82. [Editorial] ThinkingCap Qwen3.6-27B -- 2026-07-14
  83. GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine -- 2026-07-14
  84. Qwen3.6-27B: NVFP4/FP8 agent loops vs flawless BF16. Config or quant issue? -- 2026-07-13
  85. NVIDIA Puzzle-75B-A9B NVFP4 at 132 t/s on 3×3090 — Why is this size category a desert otherwise? -- 2026-07-13
  86. Exploring FlashAttention-3/4 optimizations on RTX GPUs -- 2026-07-13
  87. [audio.cpp] What Does the Fox Say: 4 ASR models in native C++/GGML, streaming support, 327s transcribed in 2.17s -- 2026-07-13
  88. [Editorial] -- 2026-07-13
  89. I benchmarked 13 models at 65K-128K context to find out what actually matters for agentic workloads -- 2026-07-10
  90. Deepseek V4 Flash running on RTX 5090 MoE -- 2026-07-10
  91. Seasonic PSU calculator now mentions RTX 5080 SUPER (24GB), RTX 5070 Ti SUPER (24GB) and RTX 5070 SUPER (18GB) -- 2026-07-10
  92. Soon we'll run 100B models on cheap hardware -- 2026-07-10
  93. Döner Bench round 2: Quant compare -- 2026-07-10
  94. Pine64 launch $50 smart speaker for Home Assistant tinkerers -- 2026-07-03
  95. Tiny Jetson Nano Orin Super Benchmarking of 1B and sub 1B LLMs — llama.cpp vs Ollama -- 2026-07-03
  96. nvidia/Cosmos3-Nano -- 2026-07-03
  97. Microsoft has taken down fastcontext model from everywhere -- 2026-07-02
  98. Orthrus (diffusion head) trained Qwen 3.5/3.6 and Gemma 4 models are dropping soon -- 2026-07-02
  99. SenseNova-U1-8b-MoT-Infographic-V2 (released yesterday) - An open source SOTA beast for infographic design and image editing. -- 2026-07-02
  100. on Dario's statement -- 2026-07-02
  101. After i distilled a 26B into a 4B to cut false positives I'll probably use the base model after all. -- 2026-07-02
  102. We built a calibration-aware Q4_K_M quant of Qwen3.5 0.8B that recovers 96.5% of the BF16 gap vs pure llama.cpp Q4_K_M (SpectralQuant) -- 2026-06-29
  103. Update: First Manual Results from Testing Procedural Skill Transfer in Small Models -- 2026-06-29
  104. Multi Tier MoE Caching -- 2026-06-29
  105. Apple raises prices of MacBooks, iPads -- 2026-06-26
  106. IBM debuts sub-1 nanometer chip technology -- 2026-06-26
  107. OpenAI and Broadcom unveil LLM-optimized inference chip -- 2026-06-26
  108. Ideogram 4: Open Image Model at the Forefront of Design -- 2026-06-25
  109. QUEST-35B: Open-Source Deep Research Agent Trained on 32 H100s — Full Recipe, Weights, and Data Released -- 2026-06-25
  110. North Mini Code: 4-Bit Quant + Ollama + OpenRouter — Now Runs on 20GB -- 2026-06-25
  111. EdgeRazor: Mixed-Precision Quantization-Aware Distillation Down to 1.58-Bit -- 2026-06-25
  112. [Editorial] Qwable-3.6-27b — Open Model Release -- 2026-06-25
  113. ScenemaAI/scenema-audio — Audio Generation Model -- 2026-06-25
  114. "Tokaine Addiction" — PhD Student's Colleague Can't Stop Running AI Agents -- 2026-06-25
  115. Suitcase Robot Uses Gas Sensor to Modulate LLM Sampler Temperature in Real-Time -- 2026-06-25
  116. [Editorial] Microsoft Majorana 2: Quantum Discovery via Agentic AI -- 2026-06-17
  117. cuTile Rust: Safe, Data-Race-Free GPU Kernels in Rust -- 2026-06-17
  118. [Editorial] Adversarial AI Research — The Malicious Use of Artificial Intelligence -- 2026-06-17
  119. [Editorial] Sloptimization — AI Citation Rot and the Tools Fighting Back -- 2026-06-15
  120. [Editorial] Leapable AI Platform -- 2026-06-15
  121. huawei-csl/KVarN -- 2026-06-12
  122. chiennv2000/orthrus -- 2026-06-12
  123. Qwen/Qwen3.5-122B-A10B -- 2026-06-12
  124. zai-org/GLM-OCR -- 2026-06-12
  125. How's Linear so fast? A technical breakdown -- 2026-06-11
  126. Port React Compiler to Rust -- 2026-06-11
  127. Show HN: Extend UI – open-source UI kit for modern document apps -- 2026-06-11
  128. MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second -- 2026-06-09
  129. Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States -- 2026-06-09
  130. A Post-Quantum Future for Let's Encrypt -- 2026-06-08
  131. Gooey: A GPU-accelerated UI framework for Zig -- 2026-06-05
  132. Branchless Quicksort faster than std:sort and pdqsort with C and C++ API -- 2026-06-05
  133. HP re-releases classic computer science calculator: The HP-16C -- 2026-06-05
  134. OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization -- 2026-05-29
  135. Shard — getting to 10× KV cache compression -- 2026-05-29
  136. LightVLM: Efficient inference toolkit for vision-language models -- 2026-05-29
  137. Speech Tokenizer Arena: Side-by-side benchmarking for discrete speech tokenizers -- 2026-05-29
  138. DeepSeek just popped the American AI bubble. -- 2026-05-29
  139. DeepSeek V4 Flash at 8.4 tok/s on 3×3090 — patching GGUF metadata for cchuter's fork -- 2026-05-29
  140. Why are the AI Companies spreading F.U.D. about AI? -- 2026-05-29
  141. $400 Qwen 3.6-27B Setup — Dual RTX 3060 — 30-50 t/s -- 2026-05-27
  142. Qwen3.6 27B made me a believer — single-shot game development with local model -- 2026-05-27
  143. Qwen3.5 35B-A3B Uncensored Heretic — Native MTP Preserved, multiple formats -- 2026-05-27
  144. [Editorial] MSI NVIDIA DGX Station -- 2026-05-27
  145. Have we passed the peak of inflated expectations? -- 2026-05-27
  146. Separable Expert Architecture: Privacy-Preserving LLM Personalization via Composable Adapters -- 2026-05-21
  147. 512k Context Pre-training on a 12GB Consumer GPU with O(n) Attention -- 2026-05-21
  148. Introducing the Ettin Reranker Family -- 2026-05-20
  149. OlmoEarth v1.1: A more efficient family of models -- 2026-05-20
  150. Regex Chess: A 2-ply minimax chess engine in 84,688 regular expressions -- 2026-05-19
  151. FastDMS: 6.4X KV-cache compression running faster than vLLM BF16/FP8 -- 2026-05-11
  152. [Editorial] RuVector Sparse Attention Crate -- 2026-05-11
  153. AMD to release slottable GPU -- 2026-05-11
  154. Taiwanese company Skymizer announces HTX301 - PCIE inference card with 384GB of Memory at ~240 Watts -- 2026-05-11
  155. Apple Removes 256GB M3 Ultra Mac Studio Model From Online Store -- 2026-05-11
  156. Virtual violin produces realistic sounds (MIT) -- 2026-05-06
  157. Lightricks/LTX-2.3-22b-IC-LoRA-HDR -- 2026-05-06
  158. [Editorial] Finding Zero-Days with Any Model -- 2026-05-01
  159. [Editorial] OMLX.ai -- 2026-04-30
  160. llama.cpp - NVFP4 native support on Blackwell from now - b8967 -- 2026-04-30
  161. Sigilant: GGUF Quality Benchmarking Beyond TPS — Tool-Calling Pass Rate as Selection Criterion -- 2026-04-30
  162. I'm done with using local LLMs for coding -- 2026-04-30
  163. Qwen3.6-27B IQ4_XS FULL VRAM with 110k context -- 2026-04-28
  164. Can we already use Google's TurboQuant (TQ) for KV Cache in llama-server? Or are we waiting for a PR? -- 2026-04-28
  165. Gemma 4 beats Qwen 3.5 (UPDATE), and Qwen 3.6 27B + MiniMax M2.7 is the best OpenCode setup -- 2026-04-28
  166. Tried Qwen3.6-27B-UD-Q6_K_XL.gguf with CloudeCode, well I can't believe but it is usable -- 2026-04-28
  167. BitNet is the AI future? -- 2026-04-28
  168. FP4 inference in llama.cpp (NVFP4) and ik_llama.cpp (MXFP4) landed -- 2026-04-27
  169. [Editorial] Bonsai-8B MLX 1-bit -- 2026-04-27
  170. VRAM.cpp: Running llama-fit-params directly in your browser -- 2026-04-27
  171. Thoughts on using an AMD Alveo V80 FPGA as a poor man's Taalas HC1 -- 2026-04-27
  172. [Editorial] hw-smi — Cross-Platform Hardware Monitor -- 2026-04-27
  173. China's DeepSeek valuation rockets above $20B!! -- 2026-04-24
  174. [Editorial] DeepSeek Open-Sources Tile Kernels -- 2026-04-24
  175. DeepSeek v4 -- 2026-04-24
  176. [Editorial] Video Content -- 2026-04-22
  177. [Editorial] Four Horsemen of the AIpocalypse -- 2026-04-22
  178. Closest replacement for Claude + Claude Code? (got banned, no explanation) -- 2026-04-22
  179. [Editorial] Roomote: Remote Development Tool -- 2026-04-22
  180. [Editorial] Video Content -- 2026-04-22
  181. Ternary Bonsai: Top Intelligence at 1.58 Bits -- 2026-04-22
  182. Personal Eval: Gemma4 26B MoE vs Qwen3.5 27B Dense vs Gemma4 31B Dense Compared -- 2026-04-22
  183. NVIDIA Nemotron-3-Super-120B-A12B-FP8 -- 2026-04-22
  184. [Editorial] Stanford HAI AI Index Report 2026 -- 2026-04-17
  185. [Editorial] Steve Yegge on AI -- 2026-04-17
  186. [Editorial] The AI Resentment Stage -- 2026-04-17
  187. A cryptography engineer's perspective on quantum computing timelines -- 2026-04-07
  188. Sam Altman may control our future – can he be trusted? -- 2026-04-07
  189. [Editorial] Anthropic's Claude Code Source Leak — What It Means -- 2026-04-01
  190. [Editorial] Claude Code Was Leaked — I Read All of It -- 2026-04-01
  191. [Editorial] The Claude Code 'Oops' — Source Code Leak -- 2026-04-01
  192. [Editorial] Claude Code Just Open-Sourced Itself (Not Intentionally) -- 2026-04-01
  193. [Editorial] nirholas/claude-code Repository -- 2026-04-01
  194. [Editorial] arxiv:2603.15569 -- 2026-03-30
  195. TinyLoRA: LoRA training works at just 13 parameters -- 2026-03-30
  196. KV rotation PR: q8 quants tank performance on AIME25, recovered with rotation -- 2026-03-30
  197. [Editorial] AI ASIC for LLMs -- 2026-03-30
  198. [Editorial] Heretic -- 2026-03-30
  199. mlx-snn: Spiking Neural Network library for Apple MLX -- 2026-03-30
  200. TurboQuant: Redefining AI efficiency with extreme compression -- 2026-03-27
  201. [Editorial] TurboQuant Deep Dive -- 2026-03-27
  202. RightNow-AI/autokernel -- 2026-03-27
  203. NVIDIA 2026 Conference LIVE. New Base model coming! -- 2026-03-20
  204. Nemotron 3 Nano 4B: A Compact Hybrid Model for Efficient Local AI -- 2026-03-20
  205. Granite 4.0 1B Speech: Compact, Multilingual, and Built for the Edge -- 2026-03-20
  206. [New Model & Agent] LocoTrainer-4B: A Claude Code-style local agent designed specifically to master the MS-SWIFT framework (4B, 32K, GGUF) -- 2026-03-20
  207. shallowdream204/BitDance-14B-16x -- 2026-03-20
  208. [Editorial] Claude 1M Context GA -- 2026-03-14
  209. NVIDIA Nemotron 3 Super: open-weight 120B MoE hybrid with 1M-token context -- 2026-03-14
  210. Expert parallelism for 1T MoE finetuning on a single node - 50x faster and 2x cheaper than alternatives -- 2026-03-14
  211. Fine-tuned Qwen3 SLMs (0.6-8B) beat frontier LLMs on narrow tasks -- 2026-03-12
  212. Towards Cold-Start Drafting and Continual Refining: A Value-Driven Memory Approach with Application to NPU Kernel Synthesis -- 2026-03-12
  213. [Editorial] -- 2026-03-09
  214. My journey through Reverse Engineering SynthID -- 2026-03-09
  215. Did Alibaba just kneecap its powerful Qwen AI team? -- 2026-03-07
  216. [Editorial] You're 1,191 Days Late — Here's What to Do -- 2026-03-07
  217. [Editorial] The Zero-Day Clock Is Ticking -- 2026-03-06
  218. [Editorial] Step-by-Step Guide to Exploiting AI Systems -- 2026-03-06
  219. [Editorial] Unprompted 2026: Top Insights Day One -- 2026-03-06
  220. [Editorial] Unprompted 2026: Top Insights Day Two -- 2026-03-06
  221. The L in "LLM" Stands for Lying -- 2026-03-05
  222. [Editorial] Am I Living in a Parallel AI Universe? -- 2026-03-05
  223. [Editorial] Video Pick -- 2026-03-05
  224. Qwen3.5 122B in 72GB VRAM (3x3090) is the best model available at this time — also it nails the "car wash test" -- 2026-03-03
  225. Qwen 3.5 is multimodal. Here is how to enable image understanding in opencode with llama cpp -- 2026-03-03
  226. pplx-embed: State-of-the-Art Embedding Models for Web-Scale Retrieval -- 2026-03-03
  227. andimarafioti/faster-qwen3-tts -- 2026-03-03
  228. Meta's AI smart glasses and data privacy concerns -- 2026-03-03
  229. [Editorial] AI Search Index -- 2026-03-03
  230. Mercury 2: Fast reasoning LLM powered by diffusion -- 2026-02-26
  231. I Benchmarked Opus 4.6 vs Sonnet 4.6 on agentic PR review and browser QA the results weren't what I expected -- 2026-02-26
  232. [Editorial] Bullshit meter :) -- 2026-02-26
  233. The Qwen team verified that there are serious problems with the data quality of the GPQA and HLE test sets. -- 2026-02-25
  234. Qwen 3.5 craters on hard coding tasks — tested all Qwen3.5 models (And Codex 5.3) on 70 real repos so you don't have to. -- 2026-02-25
  235. ChatGPT isn't the only chatbot pulling answers from Elon Musk's Grokipedia -- 2026-02-25
  236. [Editorial] Benchmarking LLMs for Voice Agent Use Cases -- 2026-02-21
  237. Claude Opus 4.6 Surges Past Forecasts on METR's 50% Time-Horizon Benchmark with Exponential Gains -- 2026-02-21
  238. [Editorial] Unsloth: MiniMax M2.5 Fine-Tuning Guide -- 2026-02-21
  239. Let your coding agent benchmark llama.cpp for you (auto-hunt the fastest params per model) -- 2026-02-06
  240. GGML implementation of Qwen3-ASR -- 2026-02-06
  241. Running LLMs &amp; VLMs Fully On-Device on iPhone(6GB RAM) — Offline, Privacy-Focused, Real-Time Performance -- 2026-02-06
  242. We benchmarked every 4-bit quantization method in vLLM 👀 -- 2026-01-12
  243. Gpu inference with model that does not fit in one GPU -- 2026-01-12
  244. Llama.cpp rpc experiment -- 2026-01-12
  245. Performance improvements in llama.cpp over time -- 2026-01-12
  246. [Editorial] https://docs.rs/crate/bitchat-qudag/latest -- 2026-01-02
  247. [Editorial] https://github.com/permissionlesstech/bitchat/blob/main/WHITEPAPER.md -- 2026-01-02
  248. [Editorial] https://www.npmjs.com/package/@ruvector/edge-net -- 2026-01-02
  249. Why I Ditched Serverless Neptune/OpenSearch for Dockerized Neo4j/pgvector on EC2 (60% Cost Cut) -- 2025-12-30
  250. Llama-3.3-8B-Instruct -- 2025-12-30
  251. Benchmarking local llms for speed with CUDA and vulkan, found an unexpected speedup for select models -- 2025-12-30
  252. Why Kimi K2 Thinking choose Int4 QAT, from infra enginner of KImi -- 2025-12-30
  253. Help RTX 5090 + llama.cpp crashes after 2-3 inferences (VFIO passthrough, SM120 CUDA) -- 2025-12-30
  254. AI-Doomsday-Toolbox Distributed inference + workflows -- 2025-12-30
  255. [Tool] imesde: Zero-GPU, In-Memory Vector Engine for Real-Time Local RAG -- 2025-12-22
  256. I built a Rust-based HTML-to-Markdown converter to save RAG tokens (Self-Hosted / API) -- 2025-12-22
  257. Golang optimizations for high‑volume services -- 2025-12-12
  258. PaCoRe: The first open-source deep think 8B model beats GPT-5 on HMMT25 -- 2025-12-11
  259. RnJ-1-Instruct FP8 Quantization -- 2025-12-10
  260. Optical Context Compression Is Just (Bad) Autoencoding -- 2025-12-10
  261. Masked Diffusion Models as Energy Minimization -- 2025-12-10
  262. Miles + FSDP2 = Megatron-Level Performance with More Flexibility -- 2025-12-10
  263. P4nda0s/IDA-NO-MCP -- 2025-12-09
  264. Toyota unintended acceleration and the big bowl of "spaghetti" code (2013) -- 2025-12-09
  265. https://huggingface.co/Doradus/Hermes-4.3-36B-FP8 -- 2025-12-09
  266. Support for rnj-1 now in llama.cpp -- 2025-12-09
  267. Comfy-Org/flux2-dev -- 2025-12-09
  268. baidu/ERNIE-4.5-VL-28B-A3B-Thinking -- 2025-12-09
  269. I built a personal assistant script, and the CPU inference speed beats my Llama setup. -- 2025-12-08
  270. Semantic Compression (2014) -- 2025-12-08
  271. A Deep Dive into Using PIO and DMA on the RP2350 -- 2025-12-05
  272. Free yourself from the Spotify desktop client with spotifyd -- 2025-12-04
  273. I cooked abliterated gemma3-27b-it with norm-preserving technique -- 2025-12-04
  274. Qwen3 VL built from scratch with PyTorch -- 2025-12-03
  275. EmbeddingGemma: Powerful and Lightweight Text Representations -- 2025-12-03
  276. Z-Image: Powerful and highly efficient image generation model with 6B parameters -- 2025-12-03
  277. RTX 5090 + Qwen 30B MoE @ 135 tok/s in NVFP4 - Full guide with C++ patches -- 2025-12-02
  278. [Editorial] https://huggingface.co/blog/grimjim/norm-preserving-biprojected-abliteration -- 2025-12-02
  279. Optimizing Token Generation in llama.cpp's CUDA Backend -- 2025-12-01
  280. [Editorial] https://arxiv.org/html/2511.09030v1 -- 2025-11-28
  281. You're using HuggingFace wrong. Stop downloading pre-quantized GGUFs and start building hardware-optimized, domain-specific models. Here's the pipeline I built to do it. -- 2025-11-26
  282. Binary Quantization For LLMs Through Dynamic Grouping -- 2025-11-26
  283. dx8152/Relight -- 2025-11-26
  284. Question About Motherboards -- 2025-11-26
  285. [Release] DragonMemory: 16× semantic compression for local RAG context (open-source, AGPL) -- 2025-11-25
  286. ORPO-Distill: Mixed-Policy Preference Optimization for Cross-Architecture LLM Distillation -- 2025-11-25
  287. Continuous batching from first principles -- 2025-11-25
  288. Can an expert chime in and explain what is holding Vulkan back from becoming the standard API for ML? -- 2025-11-25
  289. luozijian1990/network-traffic-ebpf-exporter -- 2025-11-24
  290. We found cryptography bugs in the elliptic library using Wycheproof -- 2025-11-24
  291. Show HN: Cynthia – Reliably play MIDI music files – MIT / Portable / Windows -- 2025-11-24
  292. Browser Fingerprinting and Why VPNs Won’t Make You Anonymous -- 2025-11-24
  293. [Release] Memory-Isolated Recursive Compression (MIRC). A local-first probabilistic compression utility for Apple Silicon. Research Preview (Open Source) -- 2025-11-21
  294. Read long podcasts locally with Whisper + LLM, open sourced -- 2025-11-21
  295. Local all-in-one AI system (Local multimodal AI) -- 2025-11-21
  296. JMS1717/8mb.local -- 2025-11-21
  297. Mimir Memory Bank now uses llama.cpp! -- 2025-11-21
  298. Quantum physicists have shrunk and "de-censored" DeepSeek R1 -- 2025-11-20
  299. Built a tool to solve the "how much GPU do I actually need?" problem for LLM deployment -- 2025-11-20
  300. New Parameter Browser added to Llamacpp Model Launcher! experimental model parameter tuning(window/cuda only) -- 2025-11-20
  301. cuda device list mismatch - ggml_cuda_init / ubuntu - significance to using --main-gpu flag -- 2025-11-20
  302. What Size of LLM Can 4x RTX 5090 Handle? (96GB VRAM) -- 2025-11-20
  303. Gain 60% performance on RDNA 4 using this fix -- 2025-11-19
  304. wildminder/ComfyUI-DyPE -- 2025-11-19
  305. lightx2v/Autoencoders -- 2025-11-19
  306. Scale-out is the silent killer of LLM applications. Are we solving the wrong problem? -- 2025-11-19
  307. PyTorch 2.10.0a0 w/ Blackwell (sm_120) Support — Patched &amp; Packaged for One-Command Install -- 2025-11-17
  308. Half-trillion parameter model on a machine with 128 GB RAM + 24 GB VRAM -- 2025-11-17
  309. [Editorial] https://www.linkedin.com/posts/andriyburkov_when-you-train-a-model-on-one-dataset-it-activity-7392804316769701888-166x/ -- 2025-11-14
  310. xRouter: Training Cost-Aware LLMs Orchestration System via Reinforcement Learning -- 2025-11-14
  311. [Editorial] Balancing order, freedom, and technology -- 2025-11-12
  312. AMD warns the Intel and Nvidia partnership is a risk to its business -- 2025-11-12
  313. A Pentium In Your Hand -- 2025-11-12
  314. [Editorial] https://www.linkedin.com/posts/ismaelvelasco_theres-an-ai-text-model-comparable-to-sota-activity-7393850964731912192-nZT1 -- 2025-11-12
  315. Last week in Multimodal AI - Local Edition -- 2025-11-12
  316. Apache Iggy is a high-performance, persistent message streaming platform -- 2025-11-07
  317. [P] Training Better LLMs with 30% Less Data – Entropy-Based Data Distillation -- 2025-11-06
  318. I fine tuned a (small) model to help with reasoning backfill on old/non-reasoning datasets -- 2025-11-06
  319. Superhuman AI for Multiplayer Poker -- 2025-11-06
  320. cerebras/GLM-4.5-Air-REAP-82B-A12B -- 2025-11-06
  321. Retrieval Enhanced Feedback via In-context Neural Error-book -- 2025-11-06
  322. Kimi release Kimi K2 Thinking, an open-source trillion-parameter reasoning model -- 2025-11-06
  323. OpenAI asks U.S. for loan guarantees to fund $1T AI expansion -- 2025-11-06
  324. Kimi released Kimi K2 Thinking, an open-source trillion-parameter reasoning model -- 2025-11-06
  325. [Editorial] Frequently wrong, but never in doubt’ -- 2025-11-05
  326. The Zero Freeze Formula: Teaching Local LLaMA Real Physics Through Python (SU(3) Mass Gap Simulation) to solve the Yang–Mills Mass Gap -- 2025-11-05
  327. Audio Sound Capture Project Needs Help -- 2025-11-05
  328. [D] It turns out WDDM driver mode is making our RAM - GPU transfer extremely slower compared to TCC or MCDM mode. Anyone has figured out the bypass NVIDIA software level restrictions? -- 2025-11-05
  329. GLaDOS TTS finetuning on MLX from the original game files -- 2025-11-04
  330. zeusftk/FTK_CANVAS_AGENT_for_Comfyui -- 2025-11-04
  331. guyyariv/DyPE -- 2025-11-04
  332. Qwen3-VL-32B Q8 speeds in llama.cpp vs vLLM FP8 on a RTX PRO 6000 -- 2025-11-03
  333. Help me decide: EPYC 7532 128GB + 2 x 3080 20GB vs GMtec EVO-X2 -- 2025-11-03
  334. amd/Nitro-E -- 2025-11-03
  335. CISA and NSA share tips on securing Microsoft Exchange servers -- 2025-11-02
  336. The Smol Training Playbook: The Secrets to Building World-Class LLMs -- 2025-11-02
  337. Latest Update from Anthropic's new model - Neptune V6 -- 2025-11-02
  338. AI "Phone Farm" Startup Gets Funding from Marc Andreessen to Flood Social Media With Spam -- 2025-11-02
  339. FlashPack: High-throughput tensor loading for PyTorch -- 2025-11-01
  340. M5 Neural Accelerator benchmark results from Llama.cpp -- 2025-11-01
  341. Kafka is Fast – I'll use Postgres -- 2025-11-01
  342. [Editorial] https://www.linkedin.com/posts/busiel-morley_economic-shifts-in-the-age-of-ai-ugcPost-7390349517612806144-8djS -- 2025-11-01
  343. Analog Surround Sound Was Everywhere, But You Probably Didn’t Notice -- 2025-11-01
  344. US Gas Turbine Shortage Likely to Slow AI Demand Growth -- 2025-10-31
  345. The Supercon 2025 Badge is Built to be Customized -- 2025-10-31
  346. Experimenting with Qwen3-VL for Computer-Using Agents -- 2025-10-30
  347. Built a full voice AI assistant running locally on my RX 6700 with Vulkan - Proof AMD cards excel at LLM inference -- 2025-10-30
  348. Streaming datasets: 100x More Efficient -- 2025-10-30
  349. Cerebras REAP'd GLM4.6: 25%, 30%, 40% pruned FP8 checkpoints on HF! -- 2025-10-28
  350. Qwen/Qwen3-VL-30B-A3B-Instruct-FP8 -- 2025-10-28
  351. lightx2v/Wan2.2-Distill-Loras -- 2025-10-28
  352. [Editorial] Periodic table for ai algorithms -- 2025-10-26
  353. Need help understanding OpenAIs API usage for text-embedding -- 2025-10-26
  354. Qwen3 Next support in llama.cpp ready for review -- 2025-10-25
  355. GLM Air REAP tool call problems -- 2025-10-25
  356. Reverse Engineering STL Files with FreeCAD -- 2025-10-25
  357. Un-LOCC (Universal Lossy Optical Context Compression), Achieve Up To 3× context compression with 93.65% Accuracy. -- 2025-10-24
  358. LiquidAI/LFM2-1.2B-RAG -- 2025-10-24
  359. zai-org/GLM-4.6 -- 2025-10-24
  360. inference-net/Schematron-3B -- 2025-10-21
  361. [By GLM Team] Glyph: Scaling Context Windows via Visual-Text Compression -- 2025-10-21
  362. DGX SPARK Compiled llama.cpp Benchmarks Compared to M4 MAX (non-MLX) -- 2025-10-21
  363. perplexityai/search_evals -- 2025-10-21
  364. Hetzner: The Simple Cloud just got more flexible and more affordable -- 2025-10-21
  365. A new, super simple LLM benchmark for testing changes across models, quants, parameters, samplers, engines, etc -- 2025-10-21
  366. riptideslabs/tokenex -- 2025-10-20
  367. Multi-Tenant SaaS's Wildcard TLS: An Overview of DNS-01 Challenges -- 2025-10-20
  368. From cloud to OCP? Be ready to wrangle firmware -- 2025-10-20
  369. FLOSS Weekly Episode 851: Buckets of Money -- 2025-10-20
  370. Significant speedup for local models -- 2025-10-20
  371. Cursor tricking paid users with fake Claude Sonnet 4.5 -- 2025-10-20
  372. inclusionAI/Ring-1T -- 2025-10-20
  373. volantvm/volant -- 2025-10-19
  374. Wireshark 4.6.0 Supports macOS Pktap Metadata (PID, Process Name, etc.) -- 2025-10-19
  375. A classified network of SpaceX satellites is emitting a mysterious signal -- 2025-10-19
  376. linkedlist771/SoraWatermarkCleaner -- 2025-10-19
  377. Qwen/Qwen-Image-Edit-2509 -- 2025-10-19
  378. The Entire Process of Building an Open Source Analog ASIC -- 2025-10-15
  379. Built a 1288x RTFx Parakeet Speech-to-Text server... Enjoy! -- 2025-10-13
  380. Novel OpenGL Pixel Shader Dewarping -- 2025-10-13
  381. lovis93/next-scene-qwen-image-lora-2509 -- 2025-10-13
  382. Beyond Token Count: Our Research Suggests "Contextual Weight" is a Key Limiter on Large Context Windows -- 2025-10-13
  383. FractalAIResearch/Fathom-Search-4B -- 2025-10-13
  384. LLM Robustness Leaderboard v1 --Technical report -- 2025-10-13
  385. Hide and Seek with LLMs: An Adversarial Game for Sneaky Error Generation and Self-Improving Diagnosis -- 2025-10-13
  386. Preference optimization with ORPO and LoRA -- 2025-10-12
  387. [Show] SpiralTorch: A Rust-based PyTorch-style autograd engine (Python 3.14-ready) -- 2025-10-12
  388. Qwen3-VL-30B-A3B-Thinking GGUF with llama.cpp patch to run it -- 2025-10-10
  389. Did anyone try out GLM-4.5-Air-GLM-4.6-Distill ? -- 2025-10-10
  390. What and when 7900xtx is boosted? -- 2025-10-10
  391. Modelfile. Do I need these tags PER prompt? -- 2025-10-10
  392. Divining Air Quality With A Cheap Computer Vision Device -- 2025-10-09
  393. Awesome Local LLM Speech-to-Speech Models &amp; Frameworks -- 2025-10-08
  394. FabioSarracino/VibeVoice-Large-Q8 -- 2025-10-08
  395. CapRecover: A Cross-Modality Feature Inversion Attack Framework on Vision Language Models -- 2025-10-08
  396. Mitigating Watermark Stealing Attacks in Generative Models via Multi-Key Watermarking -- 2025-10-08
  397. How to make the AI Bot to understand the exact design and App flow -- 2025-10-04
  398. Creating a Full Stack App W/Cloudflare Works and BetterAuth -- 2025-10-04
  399. Comprehension debt: A ticking time bomb of LLM-generated code -- 2025-10-04
  400. deepseek-ai/DeepSeek-V3.2-Exp -- 2025-10-03
  401. moondream/moondream3-preview -- 2025-10-03
  402. [Editorial] https://github.com/emcie-co/parlant -- 2025-10-02
  403. Built a persistent memory system for LLMs - 3 months testing with Claude/Llama -- 2025-10-02
  404. Do I need to run /init on a repo if I already have AGENTS.md? -- 2025-10-02
  405. Inside NVIDIA GPUs: Anatomy of high performance matmul kernels -- 2025-09-29
  406. Bit is all we need: binary normalized neural networks -- 2025-09-29
  407. Accelerating Qwen3-8B Agent on Intel® Core™ Ultra with Depth-Pruned Draft Models -- 2025-09-29
  408. This $5,999 RTX PRO 6000 Ebay listing is a scam, right? -- 2025-09-26
  409. Accelerating Local AI on Consumer GPUs: A Hardware-Aware Dynamic Strategy for YOLOv10s -- 2025-09-26
  410. Efficient 4B parameter gpt OSS distillation without the over-censorship -- 2025-09-22
  411. [Project] I created an AI photo organizer that uses Ollama to sort photos, filter duplicates, and write Instagram captions. -- 2025-09-22
  412. Pointer Tagging in C++: The Art of Packing Bits into a Pointer -- 2025-09-22
  413. inclusionAI/Ring-mini-2.0 -- 2025-09-22
  414. Local real-time assistant that remembers convo + drafts a doc -- 2025-09-22
  415. XiaomiMiMo/MiMo-Audio-7B-Instruct -- 2025-09-21
  416. Scaling Self-Supervised Representation Learning for Symbolic Piano Performance -- 2025-09-21
  417. Uncensor Qwen3 models without retraining -- 2025-09-20
  418. Depth upscaling? -- 2025-09-20
  419. Qwen3‑Next‑80B‑A3B‑Instruct (FP8) on Windows 11 WSL2 + vLLM + Docker (Blackwell) -- 2025-09-19
  420. unsloth/Qwen3-Next-80B-A3B-Instruct -- 2025-09-19
  421. The AI-Scraping Free-for-All Is Coming to an End -- 2025-09-18
  422. Visible Watermarking with Gradio -- 2025-09-18
  423. xiaomi-research/q-frame -- 2025-09-17
  424. google/embeddinggemma-300m -- 2025-09-17
  425. 3-month Claude Code Max user review - considering alternatives -- 2025-09-15
  426. Chesars/whatsapp-mcp -- 2025-09-15
  427. Claude’s memory architecture is the opposite of ChatGPT’s -- 2025-09-15
  428. The Internet Will Be More Dead Than Alive Within 3 Years, Trend Shows | All signs point to a future internet where bot-driven interactions far outnumber human ones. -- 2025-09-15
  429. New "speech" mode in Imagine... -- 2025-09-15
  430. I made local RAG, web search, and voice mode on iPhones completely open source, private, and free -- 2025-09-08
  431. jwest33/jam_model_memory -- 2025-09-08
  432. How was your experience with Claude vs Codex? -- 2025-09-08
  433. [Project/Code] Fine-Tuning LLMs on Windows with GRPO + TRL -- 2025-09-07
  434. nvidia/NVIDIA-Nemotron-Nano-12B-v2-Base -- 2025-09-07
  435. huihui-ai/Huihui-gpt-oss-20b-BF16-abliterated -- 2025-09-07
  436. RX570 compatibility issues -- 2025-09-07
  437. Continue.dev setup -- 2025-09-07
  438. Little SSM (RWKV7 7B) state checkpointing demo. -- 2025-09-04
  439. Need advice on how to get VLLM working with 2xR9700 + 2x7900xtx? -- 2025-09-04
  440. pwnfuzz/diffrays -- 2025-09-04
  441. Chromium Hardening Guide -- 2025-09-04
  442. roomkangali/dursgo -- 2025-09-04
  443. QuEST/Quartet authors discuss their work on SOTA 4-bit training optimizations -- 2025-09-01
  444. F-Stack – A network development kit with high performance based on DPDK -- 2025-09-01
  445. An Empirical Study of Knowledge Distillation for Code Understanding Tasks -- 2025-09-01
  446. A Comparative Analysis of Vision Language Models for Scientific Data Interpretation -- 2025-08-31
  447. Sparrow: Custom language model architecture for microcontrollers like the ESP32 -- 2025-08-30
  448. Password only for this week: Welcome to Hugston -- 2025-08-26
  449. Prism MCP Rust SDK v0.1.0 - Production-Grade Model Context Protocol Implementation -- 2025-08-26
  450. Compute Where It Counts: High Quality Sparsely Activated LLMs -- 2025-08-25
  451. moonshotai/Kimi-K2-Base -- 2025-08-25
  452. unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF -- 2025-08-25
  453. BlueLM-2.5-3B Technical Report -- 2025-08-25
  454. lightx2v/Qwen-Image-Lightning -- 2025-08-25
  455. Made Chatterbox TTS a bit faster again on CUDA (155it/s on 3090) -- 2025-08-25
  456. KittenML/KittenTTS -- 2025-08-25
  457. city96/Qwen-Image-gguf -- 2025-08-23
  458. Menlo/Lucy-128k -- 2025-08-23
  459. NVIDIA just accelerated output of OpenAI’s gpt-oss-120B by nearly 2x -- 2025-08-23
  460. COMponent-Aware Pruning for Accelerated Control Tasks in Latent Space Models -- 2025-08-23
  461. Speculative decoding in archgw candidate release 0.4.0. Could use feedback, -- 2025-08-16
  462. Nvidia Tilus: A Tile-Level GPU Kernel Programming Language -- 2025-08-16
  463. SynapseRoute: An Auto-Route Switching Framework on Dual-State Large Language Model -- 2025-08-16
  464. New Tool for Finding Why Your LLM Inference is Slow -- 2025-08-14
  465. I ran OpenAI’s GPT-OSS 20B locally on a 16GB Mac with Ollama — setup, gotchas, and mini demo -- 2025-08-14
  466. GLM 4.5 Air - Optimizing - Vulkan vs. CUDA? -- 2025-08-14
  467. GLiNER2: An Efficient Multi-Task Information Extraction System with Schema-Driven Interface -- 2025-08-13
  468. TextQuests: How Good are LLMs at Text-Based Video Games? -- 2025-08-13
  469. Mitigate Hallucinations by Fine-tuning gpt-oss-120b with One Example -- 2025-08-10
  470. uncensored gpt-oss-20b, bf16 and mxfp4 both available -- 2025-08-10
  471. LGAI-EXAONE/EXAONE-4.0-1.2B -- 2025-08-10
  472. New Open-Source Text-to-Image Model Just Dropped Qwen-Image (20B MMDiT) by Alibaba! -- 2025-08-10
  473. Kitten TTS Web Demo -- 2025-08-09
  474. Show HN: I built a tool to replace capcut audio transcription -- 2025-08-09
  475. Whispers From The Void, Transcribed With AI -- 2025-08-09
  476. The Tape Speed Keyboard -- 2025-08-08
  477. 0.82 um 105 W diode-pumped thulium-doped all silica fiber laser -- 2025-08-06
  478. GLM-4.5 llama.cpp PR is nearing completion -- 2025-08-05
  479. glm-4.5-Air appreciation poist - if you have not done so already, give this model a try -- 2025-08-05
  480. Amazon's AI Coding Revealed a Dirty Little Secret -- 2025-08-02
  481. On the Interaction of Compressibility and Adversarial Robustness -- 2025-08-02
  482. realtime-ai/blastoff-llm -- 2025-08-02
  483. Quantize your own GGUFs the same way as your fav Unsloth Dynamic GGUFs -- 2025-08-01
  484. unsloth/Qwen3-235B-A22B-Instruct-2507-GGUF -- 2025-08-01
  485. On the Predictive Power of Representation Dispersion in Language Models -- 2025-08-01
  486. Wan 2.2 T2V,I2V 14B MoE Models -- 2025-07-31
  487. PowerInfer/SmallThinker-21BA3B-Instruct -- 2025-07-31
  488. Ollama + Open WebUI -- is there a way for the same query to run through the same model multiple times (could be 3 times, could be 100 times), then gather all the answers together to summarise/count? -- 2025-07-25
  489. WGRAMMAR: Leverage Prior Knowledge to Accelerate Structured Decoding -- 2025-07-25
  490. Semantic chunking using LLMs -- 2025-07-20
  491. Does the OpenWebUi run the sentence transformer models locally? -- 2025-07-20
  492. Dataset for structured (JSON) output? -- 2025-07-19
  493. support for Kimi-K2 has been merged into llama.cpp -- 2025-07-19
  494. t-tech/T-pro-it-2.0 -- 2025-07-19
  495. Madness, the ignorant's question. Would it be possible to lighten an LLM model? -- 2025-07-18
  496. ETH Zurich and EPFL will release a fully open-source LLM developed on public infrastructure. Trained on the “Alps” supercomputer at the Swiss National Supercomputing Centre (CSCS). Trained on 60% english/40% non-english, it will be released in 8B and 70B sizes. -- 2025-07-17
  497. Moonshot AI’s open source Kimi K2 outperforms GPT-4 in key benchmarks -- 2025-07-17
  498. Advice Needed: Best way to replace Together API with self-hosted LLM for high-concurrency app -- 2025-07-17
  499. baidu/ERNIE-4.5-0.3B-PT -- 2025-07-17
  500. LiquidAI/LFM2-700M -- 2025-07-17
  501. RekaAI/reka-flash-3.1 · Hugging Face -- 2025-07-17
  502. Seq vs Seq: the Ettin Suite of Paired Encoders and Decoders -- 2025-07-17
  503. T5Gemma: A new collection of encoder-decoder Gemma models- Google Developers Blog -- 2025-07-17
  504. H-Net: a hierarchical network that replaces tokenization with a dynamic chunking process directly inside the model, automatically discovering and operating over meaningful units of data -- 2025-07-17
  505. How I build software quickly -- 2025-07-16
  506. RekaAI/reka-flash-3.1 -- 2025-07-15
  507. What kind of throughput can I expect with Llama 3.1 on a H200? -- 2025-07-15
  508. MIDI-VALLE: Improving Expressive Piano Performance Synthesis Through Neural Codec Language Modelling -- 2025-07-14
  509. Replication of Quantum Factorisation Records with an 8-bit Home Computer [pdf] -- 2025-07-14
  510. Local llms works great! -- 2025-07-12
  511. LiquidAI/LFM2-350M -- 2025-07-12
  512. QPART: Adaptive Model Quantization and Dynamic Workload Balancing for Accuracy-aware Edge Inference -- 2025-07-12
  513. Issues with Qwen 3 Embedding models (4B and 0.6B) -- 2025-07-12
  514. [Tool Release] Finetune &amp; Quantize 1–3B LLMs on 8GB RAM using LoFT CLI (TinyLlama + QLoRA + llama.cpp) -- 2025-07-11
  515. Qwen3-8B-BitNet -- 2025-07-11
  516. Megakernel doubles Llama-1B inference speed for batch size 1 -- 2025-07-09
  517. Smallest &amp; best OCR model that can read math &amp; code? -- 2025-07-07
  518. Qwen/WorldPM-72B -- 2025-07-07
  519. Code single file with multiple LLM models -- 2025-07-07
  520. Gen-Verse/CURE -- 2025-07-07
  521. Run Deepseek locally on a 24g GPU: Quantizing on our Giga Computing 6980P Xeon -- 2025-07-01
  522. I built a document workflow system using VLMs: processes complex docs end-to-end (runs locally!!) -- 2025-07-01
  523. Jan Nano + Deepseek R1: Combining Remote Reasoning with Local Models using MCP -- 2025-07-01
  524. Query Classifier for RAG - Save your $$$ and users from irrelevant responses -- 2025-07-01
  525. Building a memory-heavy AI agent — looking for local-first storage & recall solutions -- 2025-07-01
  526. Is there any easy way to get up and running with chatgpt-like capabilities at home? -- 2025-07-01
  527. No recognition of slavic characters. English characters recognized are separate singular characters, not a block of text when using PaddleOCR. -- 2025-07-01
  528. Tired of copy-pasting from ChatGPT for coding? I am building an open-source tool (Athanor) to fix that - Alpha testers/feedback wanted! -- 2025-07-01
  529. VideoGameBench from Princeton: Can vision-language models play 90s video games? -- 2025-07-01
  530. New band surges to 500k listeners on Spotify, but turns out it's AI slop -- 2025-07-01
  531. How we cut CKEditor's bundle size by 40% -- 2025-07-01
  532. My VSCode → AI chat website connector extension just got 3 new features! -- 2025-07-01
  533. 100 Gbps Indoor Access and 4.8 Gbps Outdoor Point-to-Point LiFi Transmission Systems using Laser-based Light Sources -- 2025-06-30
  534. (0,4) brane box models -- 2025-06-30
  535. Cursor 1.0 -- 2025-06-30
  536. Help me design a robust on-prem Llama 3 70B infrastructure for 30 users – Complete hardware/software list wanted -- 2025-06-30
  537. Jan-nano, a 4B model that can outperform 671B on MCP -- 2025-06-30
  538. Models that are good and fast at Long Document Processing -- 2025-06-30
  539. I am making an AI batteries included Web Framework (like Django but for AI) -- 2025-06-30
  540. [New Features & Better] Tabulens: A Vision-LLM Powered PDF Table Extractor -- 2025-06-30
  541. I tested 10 LLMs locally on my MacBook Air M1 (8GB RAM!) – Here's what actually works- -- 2025-06-30
  542. Chatbot without ChatGPT -- 2025-06-30
  543. What's the best way to save and manage different text files for the models to reference? PRD, cursor rules, tech stack, design reference, etc? -- 2025-06-30
  544. Bzip2 crate switches from C to 100% Rust -- 2025-06-30
  545. Litestream: Revamped -- 2025-06-30
  546. Stop using REST for state synchronization (2024) -- 2025-06-30
  547. Announcing `mcp-protocol-sdk`: A New Enterprise grade Rust SDK for AI Tool Calling (Model Context Protocol) -- 2025-06-30
  548. Reinforcement Pre-Training -- 2025-06-29
  549. unsloth/gemma-3n-E4B-it-GGUF -- 2025-06-29
  550. chandar-lab/NeoBERT -- 2025-06-29
  551. tencent/Hunyuan-A13B-Instruct -- 2025-06-27
  552. maya-research/Veena -- 2025-06-27
  553. Meet Mistral Devstral, SOTA open model designed specifically for coding agents -- 2025-06-26
  554. 1.93bit Deepseek R1 0528 beats Claude Sonnet 4 -- 2025-06-26
  555. DeepSeek R1 05/28 performance on five independent benchmarks -- 2025-06-26
  556. Few-Shot Examples: Overfitting / Leakage -- 2025-06-26
  557. Finetune a model to think and use tools -- 2025-06-26
  558. I need help using open web UI with Ollama. Help installing and getting it running win 11 -- 2025-06-26
  559. I built/am building a micro-transformer for learning and experimentation -- 2025-06-26
  560. I shipped more code yesterday with Claude 4 than the last 3 weeks combined -- 2025-06-26
  561. A deep dive into self-improving AI and the Darwin-Gödel Machine -- 2025-06-26
  562. 100% of the zeros of the Riemann zeta-function are on the critical line -- 2025-06-25
  563. 100% of odd hyperelliptic Jacobians have no rational points of small height -- 2025-06-25
  564. deepseek-ai/DualPipe -- 2025-06-23
  565. 100 Particles Quantum Heat Engine: Exploring the Impact of Criticality on Efficiency -- 2025-06-23
  566. 0-Auslander correspondence -- 2025-06-23
  567. Advanced Time Manipulation with GDB -- 2025-06-21
  568. Practical SDR: Getting started with software-defined radio -- 2025-06-21
  569. 1000-10,000 M$_\odot$ Primordial Stars Created the Nitrogen Excess in the Galaxy GS 3073 at $z = 5.55$ -- 2025-06-21
  570. $0^+$ to $2^+$ neutrinoless double-$β$ decay of $^{76}$Ge, $^{82}$Se, $^{130}$Te and $^{136}$Xe in the microscopic interacting boson model} -- 2025-06-21
  571. 0-1 laws for pattern occurrences in phylogenetic trees and networks -- 2025-06-20
  572. 100ps time resolution with thin silicon pixel detectors and a SiGe HBT amplifier -- 2025-06-18
  573. 0-$\pi$ quantum transition in a carbon nanotube Josephson junction: universal phase dependence and orbital degeneracy -- 2025-06-18
  574. openbmb/MiniCPM4-8B -- 2025-06-17
  575. lym00/Wan2.1-T2V-1.3B-Self-Forcing-VACE-Addon-Experiment -- 2025-06-17
  576. DeepSeek R1 05 28 Tested. It finally happened. The ONLY model to score 100% on everything I threw at it. -- 2025-06-17
  577. ubergarm/DeepSeek-R1-0528-GGUF -- 2025-06-17
  578. LLM training on RTX 5090 -- 2025-06-17
  579. [DEMO] I created a coding agent that can do dynamic, runtime debugging. -- 2025-06-17
  580. Is anyone productively using Aider and Ollama together? -- 2025-06-17
  581. For everyone who's still confused by Attention... I made this spreadsheet just for you(FREE) -- 2025-06-17
  582. What setup/model do you use and what’s your monthly spend? -- 2025-06-17
  583. Xiaomi released an updated 7B reasoning model and VLM version claiming SOTA for their size -- 2025-06-17
  584. UPDATE: Inference needs nontrivial amount of PCIe bandwidth (8x RTX 3090 rig, tensor parallelism) -- 2025-06-16
  585. IQ1_Smol_Boi -- 2025-06-16
  586. Qwen releases official MLX quants for Qwen3 models in 4 quantization levels: 4bit, 6bit, 8bit, and BF16 -- 2025-06-16
  587. Seeking Help Setting Up a Local LLM Assistant for TTRPG Worldbuilding + RAG on Windows 11 -- 2025-06-16
  588. New VS Code Pair Programming Extension, Need Help Testing -- 2025-06-16
  589. Claude-Trace -- 2025-06-16
  590. Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training -- 2025-06-16
  591. best fine tuned local LLM for Github Copilot Agent specificaly -- 2025-06-16
  592. A Simulation in C++ of Joseph Weizenbaum's 1966 Eliza -- 2025-06-16
  593. 0D-2D Heterostructure for making very Large Quantum Registers using itinerant Bose-Einstein Condensate of Excitons -- 2025-06-16
  594. 100-mJ class, sub-two-cycle, carrier-envelope phase-stable dual-chirped optical parametric amplification -- 2025-06-16
  595. [Update] Rensa: added full CMinHash + OptDensMinHash support (fast MinHash in Rust for dataset deduplication / LLM fine-tuning) -- 2025-06-15
  596. Open Source Unsiloed AI Chunker (EF2024) -- 2025-06-15
  597. ether0 - Mistral 24B with RL on several molecular design tasks in chemistry -- 2025-06-15
  598. Need selfhosted AI to generate better bash scripts and ansible playbooks -- 2025-06-15
  599. How do I finetune Devstral with vision support? -- 2025-06-15
  600. What's the best approach for including niche dependency source files and associated documentation reference material in context? -- 2025-06-15
  601. Airlines Don't Want You to Know They Sold Your Flight Data to DHS -- 2025-06-15
  602. John Deere Must Face Second Right to Repair Lawsuit -- 2025-06-15
  603. What vector database and embeddings are y'all using -- 2025-06-15
  604. Turn based two model critique for rounds to refine answer - any examples or FOSS projects? -- 2025-06-15
  605. mistralai/Magistral-Small-2506_gguf -- 2025-06-14
  606. Ruminate: From All-or-Nothing to Just-Right Reasoning in LLMs -- 2025-06-14
  607. [update] Restructured repo under rvn-tools — modular CLI for LLM formats -- 2025-06-14
  608. Testing Quant Quality for Shisa V2 405B -- 2025-06-14
  609. Old model, new implementation -- 2025-06-14
  610. Ollama vs Llamacpp: Different output for same model -- 2025-06-14
  611. How to improve my ViT model -- 2025-06-14
  612. From RPC to transactions and durable executions -- 2025-06-14
  613. Flattening Rust’s learning curve -- 2025-06-14
  614. Async from scratch 3: Pinned against the wall -- 2025-06-14
  615. How to get the most out of my AMD 7900XT? -- 2025-06-14
  616. Faulty 120W charger analysis (Anker GAN Prime) [video] -- 2025-06-13
  617. unsloth/Magistral-Small-2506-GGUF -- 2025-06-12
  618. mistralai/Magistral-Small-2506 -- 2025-06-12
  619. rednote-hilab/dots.llm1.base -- 2025-06-12
  620. [Tool] rvn-convert: OSS Rust-based SafeTensors to GGUF v3 converter (single-shard, fast, no Python) -- 2025-06-12
  621. GuidedQuant: Boost LLM layer-wise PTQ methods using the end loss guidance (Qwen3, Gemma3, Llama3.3 / 2~4bit Quantization) -- 2025-06-12
  622. I built a memory MCP that understands you (so Sam Altman can't). -- 2025-06-12
  623. Built an open source desktop app to easily play with local LLMs and MCP -- 2025-06-12
  624. mtmd : support Qwen 2.5 Omni (input audio+vision, no audio output) by ngxson · Pull Request #13784 · ggml-org/llama.cpp -- 2025-06-12
  625. i got tired of the errors, so automated debugging using Ollama -- 2025-06-12
  626. Ablating Gemma 3 27B variants with synthetic data from Sonnet 4 (Few-shot vs LoRA) -- 2025-06-12
  627. The LLM Gateway gets a major upgrade: becomes a data-plane for Agents. -- 2025-06-12
  628. Introducing stronger dependencies on systemd -- 2025-06-12
  629. How we decreased GitLab repo backup times from 48 hours to 41 minutes -- 2025-06-12
  630. The Quest for 100k - LLAMA.CPP Setting for a Noobie -- 2025-06-12
  631. Clipjacking: Hacked by copying text – Clickjacking but better -- 2025-06-11
  632. 0/1 Deep Neural Networks via Block Coordinate Descent -- 2025-06-10
  633. turbulentdrom/sing-srs-converter -- 2025-06-10
  634. abi/screenshot-to-code -- 2025-06-10
  635. Qwen/Qwen3-Embedding-4B -- 2025-06-10
  636. 100 Gbps Quantum-safe IPsec VPN Tunnels over 46 km Deployed Fiber -- 2025-06-09
  637. Qwen/Qwen3-Embedding-0.6B -- 2025-06-08
  638. 0-$π$ qubit in one Josephson junction -- 2025-06-07
  639. 100-kT Magnetic field generation using paisley targets by femtosecond laser-plasma interactions -- 2025-06-07
  640. 100 GHz Micrometer compact broadband Monolithic ITO Mach Zehnder Interferometer Modulator enabling 3500 times higher Packing Density -- 2025-06-06
  641. 0-$\pi$ phase-controllable $thermal$ Josephson junction -- 2025-06-06
  642. Precomputing Transparency Order in 3D -- 2025-06-06
  643. ban6cat6/aparecium -- 2025-06-03
  644. 0-Gaps on 3D Digital Curves -- 2025-06-03
  645. Reports of Deno's Demise Have Been Greatly Exaggerated -- 2025-06-02
  646. Comparing Parallel Functional Array Languages: Programming and Performance -- 2025-06-02
  647. What Every Programmer Should Know About Enumerative Combinatorics -- 2025-06-02
  648. DuckLake: SQL as a Lakehouse Format -- 2025-05-31
  649. 1000x Faster Camera and Machine Vision with Ordinary Devices -- 2025-05-31
  650. 0.75 Gbit/s high-speed classical key distribution with mode-shift keying chaos synchronization of Fabry-Perot lasers -- 2025-05-31
  651. EdinburghNLP/MMLongBench -- 2025-05-29
  652. Show HN: Model2vec-Rs – Fast Static Text Embeddings in Rust -- 2025-05-29
  653. unsloth/DeepSeek-R1-0528-GGUF -- 2025-05-29
  654. QuantStack/Wan2.1-VACE-14B-GGUF -- 2025-05-29
  655. 100,000 frames-per-second compressive imaging with a conventional rolling-shutter camera by random point-spread-function engineering -- 2025-05-29
  656. 1,000-Fold Enhancement of Light-Induced Magnetism in Plasmonic Au Nanoparticles -- 2025-05-29
  657. nvidia/Llama-3.1-Nemotron-Nano-4B-v1.1 -- 2025-05-28
  658. Tongyi-Zhiwen/QwenLong-L1-32B -- 2025-05-28
  659. PKU-DS-LAB/FairyR1-32B -- 2025-05-28
  660. google/medgemma-4b-pt -- 2025-05-28