Open-Weight Models

Open-source model releases, community models, licensing

1346 articles across 277 editions

Articles

  1. Editorial video submission (YouTube: _yrw6c5gw3E) -- 2026-10-08
  2. reddit.com -- 2026-10-08
  3. reddit.com -- 2026-10-08
  4. Editorial slide deck submission (Canva: DAHXJy7MerE) -- 2026-10-08
  5. [Editorial] arXiv:2602.15195 -- 2026-10-07
  6. [Editorial] arXiv:2608.02271 -- 2026-10-07
  7. [Editorial] arXiv:2602.03085 -- 2026-10-07
  8. [Editorial] weightless PR #6: abliteration without redistributing the weights (msuiche) -- 2026-10-07
  9. DEDA – Tracking Dots Extraction, Decoding and Anonymisation Toolkit -- 2026-10-07
  10. [Editorial] RED-SNOW-5.3-FLASH EXL3 SAGE 2.49bpw quant (Hugging Face) -- 2026-10-07
  11. [Editorial] ViC305 on the RED-SNOW-5.3-FLASH EXL3 release (X post) -- 2026-10-07
  12. I quantized GLM-5.3-UNCENSORED to MXFP4 for AMD GPUs - weights available on Hugging Face -- 2026-10-07
  13. [Editorial] AliesTaha/fable-traces: compact Qwen3-4B instruct tune (Hugging Face) -- 2026-10-07
  14. Beam: Reflection's 501B open-weight model -- 2026-10-07
  15. [Editorial] The Bitter Lesson (Rich Sutton) -- 2026-10-07
  16. Qwen/Qwen-AgentWorld-35B-A3B -- 2026-10-07
  17. [Editorial] Mistral Large 4 model documentation -- 2026-10-06
  18. [Editorial] Mistral CEO says new AI model beats Chinese ones in some areas (Reuters) -- 2026-10-06
  19. Q (@qtnx_) on X - Mistral Large 4 is still doing RL runs, keep seeing improvements (vs preview version). Release at the end of the month -- 2026-10-06
  20. Reflection AI Is About to Release a US Open-Weight Model to Take On DeepSeek and Qwen -- 2026-10-06
  21. [Editorial] arXiv 2610.06783: editor-flagged research paper -- 2026-10-06
  22. Microsoft confirms OpenAI has been using Looped Transformers in the GPT-6 series -- 2026-10-06
  23. [Editorial] Cantina Security: Apex Flash -- 2026-10-05
  24. [Editorial] cantina-security/apex-flash-1 on Hugging Face -- 2026-10-05
  25. [Editorial] Irregular: Assessing Claude Opus 5.5 against offensive security benchmarks -- 2026-10-05
  26. [Editorial] Audn.ai platform -- 2026-10-05
  27. [Editorial] Korea JoongAng Daily: AI hacking spree exposes cracks in financial-sector security -- 2026-10-05
  28. [Editorial] Bloomberg: Hackers breached propulsion system of US-bound oil tanker -- 2026-10-05
  29. [Editorial] Cirrius Tech: High confidence is not a security control -- 2026-10-05
  30. [Editorial] Cotool.ai -- 2026-10-05
  31. Decision models like Jev don't beat LLM-as-a-judge or traditional classifiers -- 2026-10-05
  32. [Editorial] arXiv 2609.37647 -- 2026-10-05
  33. [Editorial] decision-index (apolinario/decision-index) on GitHub -- 2026-10-05
  34. [Editorial] Jev Decision Index (Hugging Face Space by multimodalart) -- 2026-10-05
  35. Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs -- 2026-10-02
  36. AMA about K2 Horizon, Meet our team from IFM -- 2026-10-02
  37. facebookresearch/context-language-models -- 2026-10-02
  38. We benchmarked 18 RAG pipelines against an agent loop on Google's FRAMES. The best pipeline hit 78.9%. The agent loop hit 92.7%. -- 2026-10-02
  39. Amazon Unveils Strands Decider 2B: Free, Fast, Open-Source Decision Model -- 2026-10-02
  40. Mapika/decider on GitHub -- 2026-10-02
  41. nokia-applied-research/AnyJev -- 2026-10-02
  42. llama, server: add /v1/systemone API (models: laya, julia-1, lev, openjev, kev) by ngxson · Pull Request #29818 · ggml-org/llama.cpp -- 2026-10-02
  43. reddit.com -- 2026-10-01
  44. reddit.com -- 2026-10-01
  45. sensenova/SenseNova-U1.5-8B-MoT -- 2026-10-01
  46. reddit.com -- 2026-10-01
  47. Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents -- 2026-10-01
  48. reddit.com -- 2026-10-01
  49. [Editorial] -- 2026-10-01
  50. Splash fork optimised for M5 Max: ~1.5× faster (1.25× single request) -- 2026-09-30
  51. Accelerating vision-language models with LFM2.5-VL-DSpark -- 2026-09-30
  52. [NeurIPS 2026] Concurrent Image Understanding and Generation:Self-Correcting Coupled Markov Jump Processes -- 2026-09-30
  53. Qwen-family LLMs are quietly becoming the backbone of modern audio models; One chart for the architectures of 100+ audio models -- 2026-09-30
  54. lodestones/Kroma -- 2026-09-30
  55. bosonai/higgs-audio-v3-tts-4b -- 2026-09-30
  56. Oral history of John Chowning, inventor of FM synthesis [video] -- 2026-09-28
  57. FlagOpen/InsertAny3D -- 2026-09-28
  58. [Editorial] -- 2026-09-28
  59. FT: Corporate America rejects overpriced frontier, embraces open models -- 2026-09-28
  60. GPT-3 is discontinued today -- 2026-09-28
  61. nex-agi/Nex-N2-Pro -- 2026-09-28
  62. Forging 1024-bit RSA signatures in nearly SNFS time [pdf] -- 2026-09-28
  63. [Editorial] -- 2026-09-28
  64. Show HN: Air-gapped file encryption as self-decrypting HTML page -- 2026-09-28
  65. Updated from 3x3090(2x3090, 1x3090TI) to 2x5090 -- 2026-09-28
  66. Gem16 - custom engine for Gemma4 12B & 26B on Blackwell 16GB GPUs -- 2026-09-28
  67. Splash 1.1.0 released, GGUF quants support, MLX import and more -- 2026-09-28
  68. efficient fine tune storage -- 2026-09-28
  69. [Editorial] Agentic hacks, real proofs: inside Google's Pagebreak project -- 2026-09-25
  70. [Editorial] Sakana AI releases Fugu-Cyber -- 2026-09-25
  71. [Splash Engine] Qwen3.8-27B in native 8-bit at 37-55 tok/s on Apple Silicon, 256k context and the Reasoning Cliff -- 2026-09-25
  72. Bonsai 2: Qwen 3.8 27B quality at 5.9GB VRAM via ternary compression -- 2026-09-25
  73. R9V Update: KVA projections for Qwen3.8 Flash Next give 1.45-1.85x prefill speedup on 2x R9700 -- 2026-09-25
  74. [Editorial] YouTube: N1rjtDs8blY -- 2026-09-24
  75. letorig/video-generator-client -- 2026-09-24
  76. [Editorial] dgreenheck/tidewater (GitHub) -- 2026-09-24
  77. Qwen 4 Announced at Apsara Conference -- 2026-09-24
  78. z-lab/Qwen3.8-27B-DFlash2 -- 2026-09-24
  79. Open WebUI 0.11.4 is out: slim image at ~175 MB, skills from terminals, per-language model names -- 2026-09-24
  80. hkqr/my-free-code -- 2026-09-24
  81. Pruning LLMs Like a Physicist: Block Removal as an Ising Optimization Problem -- 2026-09-22
  82. [Editorial] AikidoSec/altar-1 on Hugging Face -- 2026-09-22
  83. CrofAI "cheapest inference provider in the world" gets exposed as an OpenRouter wrapper, routing requests to smaller, cheaper models at up to 20x markup. CrofAI responds to Wire Fraud allegations by denying everything, then backtracking, then 3 hours later wiping their entire online presence -- 2026-09-22
  84. [Editorial] pwardle/not-a-mused -- 2026-09-22
  85. BLOOM-WILT: Logit Tilting for Behaviour Elicitation in Automated LLM Auditing -- 2026-09-22
  86. [Editorial] lordx64/phantom-kv -- 2026-09-22
  87. The Proxy That Made No Sense (Phrack 73) -- 2026-09-21
  88. Editorial video pick (YouTube, 5vbl5FL-nsI) -- 2026-09-21
  89. Breaking the 1.58-bit Barrier for Ternary LLMs -- 2026-09-21
  90. Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint -- 2026-09-21
  91. No GPU? No Problem: Flagship LLMs on a GPU-less Teenaged Server -- 2026-09-21
  92. Qwen-3.8-Flash-Next on 1x RTX 5090: TG=50 t/s, PP=2300 t/s - with FreeToken -- 2026-09-21
  93. Show HN: Mini-AGI – Dynamic continual learning model trained on 8GB VRAM -- 2026-09-21
  94. nvidia/Cosmos3-Super-Text2Image -- 2026-09-21
  95. Cactus Needle 3: A Sliceable 8-29MB Automation Foundation Model That Matches DeepSeek v4 Flash -- 2026-09-18
  96. [Editorial] GLiNER2.5: span-free information extraction -- 2026-09-18
  97. Made the horizontal open-source model for Jev with RLCD, and it surpasses all the Jev benchmarks. HF space, benchmark, model, repo -- 2026-09-18
  98. Oilbird: Training-Free Speculative Decoding with Keys the Verifier Already Computes -- 2026-09-18
  99. Running Qwen3.8-Flash-Next locally on a 12GB VRAM card -- 2026-09-18
  100. 10%+ performance improvement on MoE ssd-streaming with expert-lookahead -- 2026-09-18
  101. Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data -- 2026-09-18
  102. HP ZGX Fury Is Now Orderable: GB300 Superchip, 748GB Unified Memory -- 2026-09-18
  103. GPT Live clone on an RTX 3060 -- 2026-09-18
  104. Connected a local model (Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-Q4_K_M) to GIMP via MCP tools using llama.cpp - and here's the image result from my first prompt "can you draw a picture of a flower in gimp?". Needs work. Setup follows. -- 2026-09-18
  105. perseval-BLR/DLSS5-NeuralScreen -- 2026-09-18
  106. [Editorial] Event Horizon Observatory: an Astra one-shot test -- 2026-09-18
  107. XingChen-AGI/Xing4.0-29B-A4B MoE -- 2026-09-17
  108. Qwen3.5 4B + grabbing logits is almost "Jev"? Or even just Qwen Reranker? -- 2026-09-17
  109. oboroge0/hayamimi -- 2026-09-17
  110. Show HN: Kinesis – Control your Mac with the Meta Neural Band -- 2026-09-17
  111. Apple Reference Image: A New Approach for Verified Photography -- 2026-09-17
  112. [Editorial] PirateFace -- 2026-09-17
  113. 1-bit 27B in the browser: 25–30 tok/s on a 6 GB RTX 3060 Laptop (WebGPU, no install) -- 2026-09-15
  114. CPU Only Experimental Sloppy Deepseek V4.1 Flash -- 2026-09-15
  115. 2×RTX 3090 + EPYC box running qwen3.8-flash-next at ~38 tok/s -- 2026-09-15
  116. Koboldcpp v1.121 released -- 2026-09-15
  117. amap-cvlab/ABot-Recon -- 2026-09-15
  118. OUI-1: a model that generates bespoke UI elements -- 2026-09-15
  119. Comfy-Org/Ideogram-4 -- 2026-09-15
  120. ScottStevenson/SuperAstra -- 2026-09-15
  121. A super fast, non-expensive alternative to motion capture - [ft. Sara Silkin] -- 2026-09-15
  122. DeepSeek V4.1 Flash beats Astra on AA's new benchmark -- 2026-09-14
  123. DSV4 Flash 0731 on OpenRouter. Why is the price SO LOW -- 2026-09-14
  124. [Editorial] So you want to use OpenRouter -- 2026-09-14
  125. IFM/K2-Horizon-MoVA-36B-A4B -- 2026-09-14
  126. [Editorial] -- 2026-09-11
  127. [Editorial] -- 2026-09-11
  128. [Editorial] -- 2026-09-11
  129. yebiguo/ProofRun -- 2026-09-11
  130. tonyd2wild/GLM-5.3-Flash-NVFP4-DFlash2-2x-DGX-Spark -- 2026-09-11
  131. Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses -- 2026-09-11
  132. I made a custom llama.cpp build optimized for 7900xtx (one or two). for qwen 3.8 next and 27B. includes optimizations for PciE x4 and tensor parallel. read inside! (no AI slop) -- 2026-09-11
  133. [Editorial] GCD-AuthZ: authorization paper (cybersharkvin) -- 2026-09-10
  134. Gradient Immunity: Null-Space Resistance to Malicious Fine-Tuning -- 2026-09-10
  135. [Editorial] Has AI Resolved the Navier-Stokes Problem? -- 2026-09-09
  136. On the Value of Human Ideas: What data poisoning research reveals about "autonomous" AI breakthroughs -- 2026-09-09
  137. Harnessing the Universal Geometry of Embeddings -- 2026-09-09
  138. Zulwatha/content-parity -- 2026-09-09
  139. WSJ: Unregulated Open-Weight AI Is an Invitation to Disaster -- 2026-09-09
  140. [Editorial] Post by @hilbertspaess on X -- 2026-09-09
  141. Tristan Buckmaster's statement on the AI-assisted Navier-Stokes proof and the OpenAI/Anthropic dispute -- 2026-09-08
  142. Is mathematics about to enter the conservatory? -- 2026-09-08
  143. MTP released for Qwen3.8-Flash-Next-GGUF -- 2026-09-04
  144. I pushed Qwen3.8-27B to 2.000 prefill per second and 132 decode per second on A RTX 3090. -- 2026-09-04
  145. Qwen 3.8 27B NVFP4, in a single 5090, using nInfer above 200tps at 180K contexts -- 2026-09-04
  146. R9V: A designer set of kernels I've been working on for R9700s/RDNA4. Qwen3.8-Flash-Next Unsloth IQ4_XS (w/ TP on 2 R9700s, MTP, SSD n-gram, 128k ctx, vision): TG256 of *78 tok/s* (~3x increase), PP8192 of *1510 tok/s* (~30x increase). -- 2026-09-04
  147. [Editorial] Autoresearch: Sticky Refusals, Free Speculative Decoding, and the Invisible Quantisation Cliff -- 2026-09-04
  148. [Editorial] Weightless (msuiche) -- 2026-09-04
  149. Anyone else notice strange refusal-related reasoning traces from Qwen3.8-Flash-Next during routine coding sessions? -- 2026-09-04
  150. [Editorial] NEEDLE: Headless Deterministic Agent Orchestrator for LLM CLIs (jedarden) -- 2026-09-03
  151. [Editorial] Dreadnode Alfred (GitHub) -- 2026-09-03
  152. [Editorial] The Emergent Symbolic Structure of Artificial Neural Networks (arXiv 2608.29530) -- 2026-09-03
  153. How to build a diffusion language model -- 2026-09-03
  154. Qwen3.8-Flash-Next on 2x3090 + DDR4: 17 → 25-29 t/s decode with the expert cache PR -- 2026-09-03
  155. It's official! 192GB Framework -- 2026-09-03
  156. Vision support merged for DeepSeek-V4-Flash-Vision-Exp -- 2026-09-03
  157. Muse Spark open weights coming soon -- 2026-09-03
  158. internlm/Intern-S2-Preview -- 2026-09-03
  159. Editor's video pick (YouTube SJ92rOuk9Xc) -- 2026-09-02
  160. Editor's video pick (YouTube QNPwKMOQIKM) -- 2026-09-02
  161. AI-Driven Hacking Boom Fuels Cybersecurity Burnout at Hospitals and Banks (Bloomberg) -- 2026-09-02
  162. AI-Driven Hacking Boom Fuels Cybersecurity Burnout (Bloomberg, archive mirror) -- 2026-09-02
  163. How Leadership Anxiety Derails Transformation (MIT Sloan Management Review) -- 2026-09-02
  164. GLM-5.3 weights will be released tomorrow -- 2026-09-01
  165. GLM 5.3 and GLM 5.3 Flash ran locally on RTX PRO 6000 WS and built a penthouse using BlenderMCP -- 2026-09-01
  166. friendly reminder you can legally torrent ai models. -- 2026-09-01
  167. [Editorial] -- 2026-09-01
  168. [Editorial] -- 2026-09-01
  169. Natural Ungrokking: Asymmetric Control of Which Rules Survive Pretraining -- 2026-08-31
  170. Claude Plays DOOM -- 2026-08-31
  171. [Editorial] Anthropic warns infostealer malware is hijacking Claude sessions to drain usage -- 2026-08-31
  172. Cheap AI Token Resellers: The Secret Ingredient is Fraud -- 2026-08-31
  173. No, Engrams won't let you run 1T models locally. It does something even better. -- 2026-08-31
  174. What is Qwen 3.8 Next Engram usage? -- 2026-08-31
  175. Qwen3.8-Flash-Next turns 4xR9700 into a local AI powerhouse! 120 t/s TG and 12k t/s PP single request with optimized vLLM -- 2026-08-31
  176. Qwen3.8-27b q8 KV cache does seem to actually hurt model performance -- 2026-08-31
  177. sakamakismile/Qwen3-VL-32B-Heretic-MiniMax-H3-NVFP4 -- 2026-08-31
  178. Perplexity and Nvidia partner for local-first AI platform -- 2026-08-31
  179. Open WebUI 0.11.1: Streaming rebuilt, HITL Tool approval, 303 Changes! -- 2026-08-31
  180. V0.3.0 of LifeOS is out! End to end runnable on 12GB of vram. -- 2026-08-31
  181. OpenShot 4.0: Record, Edit, and Color Like Never Before -- 2026-08-31
  182. [My Experience] Claude Code recognizes when a user regularly performs safety research and dynamically reduces safety guardrails -- 2026-08-28
  183. [Editorial] Featured video (YouTube: uyzqxIoiobU) -- 2026-08-28
  184. [Editorial] Featured video (YouTube: Lf5oqGOCRCM) -- 2026-08-28
  185. Extremely Fast and Accurate Transcription with Granite Speech 5.0 Turbo CTC -- 2026-08-28
  186. I implemented a modern LLM in 700 lines of C -- 2026-08-28
  187. [Editorial] -- 2026-08-27
  188. [Editorial] -- 2026-08-27
  189. opaxial/CVE-2026-9830 -- 2026-08-27
  190. [Editorial] -- 2026-08-27
  191. Laion Big Video Dataset -- 2026-08-27
  192. FireRedTeam/FireRedTTS3 -- 2026-08-27
  193. Fully quantized NVFP4 Qwen3.8-27B with QUASAR QAD -- 2026-08-27
  194. Qwen3.8-Flash-Next better then DeepSeek V4 Pro -- 2026-08-27
  195. I really want DeepSeek V4 to work as a local coding agent, but the tool calling keeps falling apart. Has anyone solved this? -- 2026-08-27
  196. [Editorial] -- 2026-08-27
  197. The Hugging Face incident and the road ahead -- 2026-08-27
  198. How Complex Systems Fail (1998) -- 2026-08-27
  199. [Editorial] Z.ai Confirms OX-Alpha GLM Model Weight Release -- 2026-08-26
  200. [Editorial] GLM-5.3-Flash Lands on Hugging Face -- 2026-08-26
  201. [Editorial] Qwen3.8-Flash-Next on Hugging Face -- 2026-08-26
  202. Qwen 3.8 27B vs Gemini 3.7 Flash (High) for real coding: open-source 27B model did a much better job -- 2026-08-26
  203. [Editorial] Signal Windows Desktop: ContentProtection Bypass (IOActive) -- 2026-08-26
  204. Spoofed Serial Number Unlocks Cricut Machine -- 2026-08-26
  205. "One robot could infect other vulnerable robots nearby ... Attackers could take control of entire fleets of robots." -- 2026-08-26
  206. [Editorial] Scaling Memory Safety (Google Bug Hunters) -- 2026-08-25
  207. Gonna be huge for US open source -- 2026-08-25
  208. Ling-3.0 releases six base checkpoints: tiny and flash across three training stages -- 2026-08-25
  209. you can now use MTP in GLM-Air -- 2026-08-25
  210. Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original -- 2026-08-25
  211. Fun BF16 checkpoint gotcha, +1 followed by -1 isn’t always a round trip -- 2026-08-25
  212. [Editorial] Researchers Complain That OpenAI Revoked Their Access to Limited Cyber Program -- 2026-08-24
  213. [Editorial] Smolbox — Live Site -- 2026-08-24
  214. [Editorial] Smolbox — remyhax Writeup -- 2026-08-24
  215. [Editorial] gadievron/raptor — Issue #889 -- 2026-08-24
  216. Every Model Cheats -- 2026-08-24
  217. [Editorial] DreamLab-AI Loom — Research Paper v4 -- 2026-08-24
  218. [Editorial] Editor-Curated Video Pick #2 -- 2026-08-24
  219. [Editorial] Editor-Curated Video Pick #1 -- 2026-08-24
  220. [Editorial] sw30labs/singularity-atlas -- 2026-08-24
  221. [Editorial] Ox Alpha: Stealth AI Model With 1M-Token Context -- 2026-08-24
  222. [Editorial] Coding Model Ox Alpha Retains Every Prompt — And You Can't Name the Company Holding Them -- 2026-08-24
  223. Ox Alpha stealth model: GLM5 Air, Mimo V3 or ? -- 2026-08-24
  224. Artificial Analysis' Qwen3.8-27B benchmarks put it neck and neck with DeepSeek V4 and GPT-5.6 Luna Max -- 2026-08-21
  225. [Editorial] YouTube feature -- 2026-08-21
  226. DeepSeek-V4-Flash-Vision-Exp -- 2026-08-21
  227. Microsoft ran a 100B BitNet test on one CPU. Here is the detail the headline misses. -- 2026-08-21
  228. The August 17 outage -- 2026-08-21
  229. [Editorial] RepoRadar -- 2026-08-21
  230. [Editorial] Event-Horizon (ruvnet) -- 2026-08-21
  231. GIMP Development Update -- 2026-08-21
  232. [Editorial] music.cognitum.one -- 2026-08-21
  233. It's actually crazy how good DSv4 Flash 0731 is -- 2026-08-19
  234. I ran DeepSeek V4 Flash 284B + DSpark on one RTX PRO 6000. The drafter was faster in RAM than VRAM. -- 2026-08-19
  235. Deepseek Harnness - why is feels better -- 2026-08-19
  236. [Editorial] pawaca/dsh-edge -- 2026-08-19
  237. Stolen LLM Reasoning: How come OpenAI, Anthrophic, Google have the same vulnerabilities? -- 2026-08-19
  238. AI-Generated GitHub Copilot “Autofix” Allowed Compromise of Snowflake's Jira -- 2026-08-19
  239. Israel creates fake think tank in likely attempt to dupe AI chatbots -- 2026-08-19
  240. The AI Credit Resale Economy -- 2026-08-19
  241. Simon Willison: Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things -- 2026-08-19
  242. Ollama's MTP variant of Qwen3.8-27B is 2x slower than the non-MTP one, measured, with a negative control -- 2026-08-19
  243. Taking Qwen3.5-9B quants to SOTA. New lineup incoming :) -- 2026-08-19
  244. [Paper] Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning -- 2026-08-19
  245. dealignai/Gemma-4-31B-JANG_4M-CRACK -- 2026-08-19
  246. bitsandbytes creator teasing new quantization method: GLM 5.3 on a single DGX Spark at 7t/s -- 2026-08-18
  247. GLM 5.3 weights. It might offer the best capacity-to-size ratio. -- 2026-08-18
  248. Ling 3.0 support merged into llama.cpp -- 2026-08-18
  249. Why not? ☺️ -- 2026-08-18
  250. A Preview of DuckDB v2.0 -- 2026-08-18
  251. Reticulum – Decentralized Mesh Network -- 2026-08-18
  252. Design 3D-printable parts by talking -- 2026-08-18
  253. Store Tunes on Paper and Stream Them Over LoRA -- 2026-08-18
  254. RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM -- 2026-08-18
  255. I built an open source local memory engine (Hillock v0.2) to ingest docs in sub-seconds alongside Ollama -- 2026-08-18
  256. Open Web UI my usage (local LLM for a compagny) -- 2026-08-18
  257. [Editorial] GLM-5.3 Release (z.ai) -- 2026-08-14
  258. [Editorial] LinkedIn Post (gMy2vbnc) -- 2026-08-14
  259. State of Open Models: Summer 2026 Observations -- 2026-08-14
  260. Qwen 3.8 27B is out : open weights, best local dense model yet -- 2026-08-14
  261. CohereLabs/North-Micro-Vision-Instruct · Hugging Face -- 2026-08-14
  262. U.S. Department of Energy Launches the Genesis Open Models Initiative and, with Arcee, Unveils Genesis-Science-1 — Its First Open-Weight Model for Scientific Research -- 2026-08-14
  263. We quantized DeepSeek V4 0731 and benchmarked it against popular quants on 8× RTX 5090 -- 2026-08-14
  264. Gemma 4 12B Q3: +8.55% Coding Performance From Tensor-Level Quantization Allocation -- 2026-08-14
  265. I trained a 1B-parameter LLM from scratch on 20B tokens for about $200 -- 2026-08-14
  266. Showoff Saturday: Local 4x 6000 Pro (multi-year progression) -- 2026-08-14
  267. [Editorial] Foxflow Fleet (probagi.com) -- 2026-08-14
  268. Mojo 1.0 -- 2026-08-13
  269. DeepSeek Harness -- 2026-08-13
  270. [Editorial] Unsloth Desktop -- 2026-08-13
  271. I built a local realtime voice stack for Ollama: Parakeet STT → Qwen 2.5 7B → Qwen3-TTS -- 2026-08-13
  272. I put Gemma 4 E4B and E2B into an e-reader so I can ask my weird questions and share my thoughts in private directly in app. -- 2026-08-13
  273. I wanted to understand Transformers below the PyTorch abstraction layer, so I built one from scratch in CuPy -- 2026-08-13
  274. [Editorial] DeepSWE by Datacurve -- 2026-08-13
  275. Anthropic: Introducing The Conceptual Reasoning Index -- 2026-08-13
  276. pathwaycom/arc-task-gen -- 2026-08-13
  277. [Editorial] Infisical Agent Vault -- 2026-08-13
  278. [Editorial] jedarden/seam -- 2026-08-13
  279. [Editorial] LinkedIn feature -- 2026-08-13
  280. OpenAI on upcoming model "Astra" (GPT-6): "We're treating it as our first "critical" model for cybersecurity" -- 2026-08-12
  281. Everything you do is being recorded -- 2026-08-12
  282. MiniMax-AI/MiniMax-H3 -- 2026-08-11
  283. Glimmer seems pretty censored? -- 2026-08-11
  284. No wonder Qwen and Gemma are so different -- 2026-08-11
  285. CJK Manga/Manhwa/Manhua 150M OCR model (hayai-ocr-v2) outperforming PaddleOCR-VL-For-Manga -- 2026-08-11
  286. [Editorial] Anthropic's Invisible Watermarks in Claude Text -- 2026-08-11
  287. [Editorial] Reuven Cohen: Quantum Mechanics May Be Telling Us Something -- 2026-08-11
  288. Underestimated budget solution: radeon 780m iGPU -- 2026-08-10
  289. Building a zero-dependency C inference engine for BitNet (1.58-bit) - lessons from hitting 36 tok/s on a Xeon CPU -- 2026-08-10
  290. I got tired of my 300GB model loads taking 5min on RPC. PR 26291 speeds it 300% to 1min30sec (4060ti+ddr4) + (4060ti+ddr5) -- 2026-08-10
  291. Early signs that Muse-Glimmer-30B might quantize *very* well? Share your experiences. -- 2026-08-10
  292. Best Local LLMs - August 2026 -- 2026-08-10
  293. Qwen3-TTS voice cloning is now in mainline llama.cpp — the old demo finally became real support -- 2026-08-10
  294. [Editorial] -- 2026-08-10
  295. [Editorial] -- 2026-08-10
  296. [Editorial] -- 2026-08-10
  297. DeepSeek V4 Flash 2-bit quant achieves 100% on SQL benchmark locally -- 2026-08-07
  298. Scotoma-2: Gemma4, but with less annoying slop and better writing -- 2026-08-07
  299. Mach-1 Additive: 95% of Qwen 3.6 35B performance while 10x smaller -- 2026-08-07
  300. Intern S2 Mobius — Qwen3.5-35B derivative with architectural throughput gains -- 2026-08-07
  301. A 2.6B model with tool calling and 128K context now runs at 30 tok/s on a phone -- 2026-08-06
  302. inclusionAI/Ling-3.0-flash weights are up on Hugging Face — MIT, BF16 plus an official FP8 -- 2026-08-06
  303. They almost catched up on Frontier performance, so now catching up on prices -- 2026-08-06
  304. Deepseek v4 flash 0731 still not holding up. -- 2026-08-06
  305. Maple-Preview — Ternary 20B MoE running at 120 tok/s on an iPhone -- 2026-08-05
  306. Qwen Developers' AMA: 3.8-27B coming soon, 2.4T params (95B active) -- 2026-08-05
  307. AI9Stars released G9v3-39A5B — Apache 2.0, 39B/5B active MoE -- 2026-08-05
  308. [Editorial] Unsloth Kimi K3 Model Documentation -- 2026-08-05
  309. [Editorial] Series of Model Tests and Results -- 2026-08-05
  310. Smaller, faster, safer: running Kimi and GLM at scale on Cloudflare -- 2026-08-04
  311. Huawei open-sources openPangu-2.0-Pro: 505B-A18B MoE on Ascend -- 2026-08-04
  312. NousResearch ships Hermes 0.20 agent framework -- 2026-08-04
  313. DMARC Has Been Public Since 2012. 68.4% of Domains Still Don't Enforce It -- 2026-07-31
  314. Kimi K3 Architecture Overview and Notes -- 2026-07-31
  315. DeepSeek-V4-Flash-0731 on HuggingFace -- 2026-07-31
  316. [Editorial] Poolside Laguna-S-2.1 — Open Coding Model -- 2026-07-31
  317. LFM2.5-Encoders: Fast at Long Context, Even on CPU -- 2026-07-31
  318. Dario Amodei: closed-weights models are worse than open-weights ones? -- 2026-07-31
  319. Tested proven orchestration techniques on small local models — 90% failed, the 10% that survived roughly doubled task completion -- 2026-07-31
  320. LiteRT-LM is up to 3.5x faster than llama.cpp on Intel Arc iGPU -- 2026-07-31
  321. First Kimi K3 results on home lab — ~4 t/s -- 2026-07-31
  322. SK Hynix stock fell 40% in 30 days — cheap RAM and GPUs again? -- 2026-07-31
  323. [Editorial] Video Content -- 2026-07-31
  324. SciCodePile: A 128GB Corpus and Executable Benchmark for Scientific Code Generation -- 2026-07-31
  325. Procedural desert explorer built with Claude Code (Opus 5) and Three.js -- 2026-07-31
  326. img2threejs — Image-to-3D procedural Three.js model generation -- 2026-07-31
  327. Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac -- 2026-07-30
  328. STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU -- 2026-07-30
  329. [Editorial] -- 2026-07-30
  330. [Editorial] -- 2026-07-30
  331. [Editorial] -- 2026-07-30
  332. Was waiting for Kimi 3 and now Ollama release it and has to pay extra to use it (like OpenRoute) -- 2026-07-30
  333. Everyone posts day-one impressions. What's still in your stack a month later? -- 2026-07-30
  334. [Editorial] -- 2026-07-30
  335. Kimi-K3 Releases on HuggingFace 6/27 -- 2026-07-28
  336. Ling-3.0-flash weights: SGLang says day-0, vLLM says when they land, llama.cpp closed the 2.6 request as not_planned -- 2026-07-28
  337. Sanctions on Open Source. hope they don't do anything stupid here. -- 2026-07-28
  338. [Editorial] -- 2026-07-28
  339. RTX 2080 Ti Memory Upgrade to 22 GB -- 2026-07-28
  340. microsoft/VibeVoice-ASR-BitNet -- 2026-07-28
  341. tetsuo-ai/voice_clone_lab -- 2026-07-28
  342. [Editorial] -- 2026-07-27
  343. [Editorial] -- 2026-07-27
  344. OpenAI and Anthropic unite against open-weight AI risks to their bottom line -- 2026-07-27
  345. [Editorial] -- 2026-07-27
  346. [Editorial] -- 2026-07-27
  347. The Open Weight Unbundling -- 2026-07-24
  348. Model "distillation" accusations are getting way overblown at this point -- 2026-07-24
  349. The Distillation Claims Are Fake and Desperate -- 2026-07-24
  350. I gave Kimi K3 a shot at auditing my post-quantum crypto project, it found 5 real bugs Fable/Opus 4.8 and GPT-5.6 Sol had all missed -- 2026-07-24
  351. China's Kimi K3 fuels fears safety curbs are holding back US AI -- 2026-07-23
  352. Startup founders urge Trump not to shut off Chinese open weight AI -- 2026-07-23
  353. Ahem! Qwen is on the move again -- 2026-07-23
  354. [Editorial] -- 2026-07-23
  355. tried predicting which MoE experts get used next token to speed up cpu/gpu offload, got some real numbers, is this actually implementable or am i wasting my time (30tg/s -> 150-200tg/s) -- 2026-07-23
  356. DS V4 on single b300. only 770 tok/s batched in vLLM -- 2026-07-23
  357. New Framework Desktop Option with AMD Ryzen AI Max+ Pro 495 and 192GB Memory -- 2026-07-23
  358. [Editorial] -- 2026-07-23
  359. [Editorial] -- 2026-07-23
  360. Are AI labs pelicanmaxxing? -- 2026-07-23
  361. China's open-weights AI strategy is winning -- 2026-07-22
  362. How do we benefits from 2+ T models? -- 2026-07-22
  363. Serving a fleet of Qwen3.5 122b sessions on a single Mac Studio (96GB) without losing your sanity -- 2026-07-22
  364. Intel Starts Shipping High-NA EUV Silicon -- 2026-07-22
  365. The State of Simulation for Physical AI: An Overview -- 2026-07-22
  366. [Editorial] -- 2026-07-22
  367. OpenAI had to pause an unreleased model after it escaped containment. -- 2026-07-21
  368. [Editorial] -- 2026-07-21
  369. Quant Convergence: Bridging Classical Value Investing and Modern Factor Models for Systematic Equity Selection -- 2026-07-21
  370. [Editorial] -- 2026-07-21
  371. [Editorial] -- 2026-07-21
  372. Model Routing Is Simple. Until It Isn't. — IBM Research -- 2026-07-20
  373. [Editorial] OmniRoute — Model Routing Framework -- 2026-07-20
  374. I Burned All My Tokens Researching How to Save Tokens -- 2026-07-20
  375. [Editorial] GCF-Rust — Blackwell Systems GPU Compute Framework -- 2026-07-20
  376. Kimi K3 Agentic Benchmark -- 2026-07-17
  377. Kimi k3 is 2.8t! Will need to have an aggressive iQ2_XXS or IQ1.8! -- 2026-07-17
  378. Governments, companies, nonprofits should invest in free, open source AI [pdf] -- 2026-07-17
  379. Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon with no GPU -- 2026-07-17
  380. LM Studio Bionic: the AI agent for open models -- 2026-07-17
  381. Prism-ML Bonsai Qwen 3.6 27B -- 2026-07-17
  382. i tried ternary decomposition instead of quantization. it works as good at q4km but takes slightly more vram. while being completly ternary. and completly PTQ (no QAT) -- 2026-07-17
  383. Prayers requested for an Abomination(Modified Gemma) -- 2026-07-17
  384. Hermes on Android (Graphene OS) -- 2026-07-17
  385. Inkling: Our Open-Weights Model -- 2026-07-16
  386. [Editorial] -- 2026-07-16
  387. hustvl/Moebius -- 2026-07-16
  388. Alternative(s) to run CUDA on non-Nvidia hardware -- 2026-07-16
  389. High-Bandwidth Flash offers efficient storage for model weights -- 2026-07-16
  390. Colibri streaming for Hy3 (Run Hy3 on 10GB (V)RAM) -- 2026-07-16
  391. In some languages, Claude will be more strict. Anthropic found out how language changes AI responses. -- 2026-07-15
  392. Building Food Metadata with LLM Juries -- 2026-07-15
  393. [Editorial] -- 2026-07-15
  394. Self-hosted voice for any agent/harness of your choice (open-source) -- 2026-07-15
  395. The untuned 27B beat the tuned 75B as an agent -- 2026-07-15
  396. Qwen3.6 35B-A3B (Q8_0, no KV quant) single prompt in opencode -- 2026-07-15
  397. I didn't give up - extGemma4-40_5B returned -- 2026-07-14
  398. J-Space Hallucination Signal Stress-Tested Across 7 Datasets on Qwen3-4B -- 2026-07-14
  399. FT: Companies Turn to Chinese Open Weight Models to Cut Costs -- 2026-07-14
  400. PrismML Compresses Qwen-3.6-27B to Under 4GB — Runs on iPhone 17 Pro -- 2026-07-14
  401. [Editorial] ThinkingCap Qwen3.6-27B -- 2026-07-14
  402. GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine -- 2026-07-14
  403. Muse Spark 1.1 -- 2026-07-10
  404. If the GPT-5.6 SOL rumors are true, is anyone actually sticking with Fable 5? -- 2026-07-10
  405. KilimcininKorOglu/M365Bridge -- 2026-07-10
  406. nvidia/Nemotron-Labs-Audex-30B-A3B · Hugging Face -- 2026-07-08
  407. Gemma 4 Technical Report -- 2026-07-08
  408. Hugging Face and Cerebras bring Gemma 4 to real-time voice AI -- 2026-07-08
  409. Qwen's J-Space — Anthropic's Discovery of an Internal Model Global Workspace -- 2026-07-07
  410. Physics-Style Attractor System Replaces Neural Networks for Word Embeddings — Hits SimLex-999 ρ=0.36 on 7.5% of Wikipedia -- 2026-07-07
  411. Why Specialization Is Inevitable -- 2026-07-07
  412. Follow-up: GLM-5.2 NVFP4 on four DGX Sparks — the MTP mystery is solved, and it's now ~24 tok/s at 128K context -- 2026-07-06
  413. Follow-up: DeepSeek V4 Flash on 2x RTX PRO 6000 finishes real coding tasks faster than Sonnet and Opus, at about Sonnet quality -- 2026-07-06
  414. Best Local VLMs - July 2026 -- 2026-07-06
  415. I benchmarked PrismML's 1-bit Bonsai-8B against IBM's Granite on CPU tool calling. The 1-bit model won, but only with grammar-constrained decoding -- 2026-07-06
  416. NASA testing local LLM inference for future space missions -- 2026-07-06
  417. Hierarchos: Preliminary Findings From a 232M Recurrent Memory-Augmented Assistant Model -- 2026-07-03
  418. DiScoFormer: One transformer for density and score, across distributions -- 2026-07-03
  419. Mapping Local Nodes — Visualizing Model Activation Paths -- 2026-07-03
  420. Microsoft has taken down fastcontext model from everywhere -- 2026-07-02
  421. Orthrus (diffusion head) trained Qwen 3.5/3.6 and Gemma 4 models are dropping soon -- 2026-07-02
  422. SenseNova-U1-8b-MoT-Infographic-V2 (released yesterday) - An open source SOTA beast for infographic design and image editing. -- 2026-07-02
  423. on Dario's statement -- 2026-07-02
  424. After i distilled a 26B into a 4B to cut false positives I'll probably use the base model after all. -- 2026-07-02
  425. LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active -- 2026-07-01
  426. [Editorial] Google Gemini Model Family Updates -- 2026-07-01
  427. InternScience/Agents-A1 · Hugging Face -- 2026-07-01
  428. [Editorial] Video Submission -- 2026-07-01
  429. Accio-Lab/Dressage -- 2026-07-01
  430. dondai1234/agent-browser -- 2026-07-01
  431. Emry: an event-sourced, local-first observability engine for long training runs -- 2026-07-01
  432. Position: AI Safety Requires Effective Controllability -- 2026-07-01
  433. Godot will no longer accept AI-authored code contributions -- 2026-07-01
  434. I built a desktop AI that scrubs your PII locally before it hits the cloud -- 2026-07-01
  435. DeepSeek V4 official version launching mid-July -- 2026-06-30
  436. OpenPangu-2.0-Flash: 92B MoE (6B active) on Ascend with 512K context -- 2026-06-30
  437. Anthropic's Amodei: "Open Source models [could take us to] a very dangerous place." -- 2026-06-30
  438. Even Google still believes in small models for coding — Gemma 4 31B hackathon at 1500 tok/s -- 2026-06-30
  439. deepseek-ai/DeepSeek-V4-Pro-DSpark • Huggingface -- 2026-06-29
  440. High-quality GLM-5.2 Quant on 4x DGX Spark - Guide, Results, and Comps -- 2026-06-29
  441. Ornith 35B is great so far -- 2026-06-29
  442. Book Review: Domain-Specific Small Language Models by Guglielmo Iozzia -- 2026-06-29
  443. Previewing GPT‑5.6 Sol: a next-generation model -- 2026-06-29
  444. [Editorial] -- 2026-06-29
  445. [Editorial] -- 2026-06-29
  446. [Editorial] -- 2026-06-27
  447. US government allows Anthropic limited release of AI model that sparked cybersecurity concerns | CNN Business -- 2026-06-27
  448. Fable 5 return RUMORED with some hints in CC -- 2026-06-27
  449. [Editorial] -- 2026-06-27
  450. [Editorial] -- 2026-06-27
  451. [Editorial] -- 2026-06-27
  452. trotsky1997/OpenFugu -- 2026-06-27
  453. SPIRAL: Learning to Search and Aggregate -- 2026-06-27
  454. LFM2.5 230M running in-browser at 1,400 tok/s using custom WebGPU kernels -- 2026-06-26
  455. Got GLM-5.2 + MTP speculative decode running on 4x DGX Spark (GB10) — and the build piece the public recipe is missing -- 2026-06-26
  456. Run a vLLM Server on HF Jobs in One Command -- 2026-06-26
  457. Running Sonnet 4.6 on every Instagram DM for a 7-location restaurant. 97% cache hit is the only reason it's affordable -- 2026-06-26
  458. [Editorial] rupixel -- 2026-06-26
  459. GLM-5.2 is a step change for open agents -- 2026-06-25
  460. poolside/Laguna-M.1 · Hugging Face - 225B-A23B -- 2026-06-25
  461. Mimo 2.5 is _fast_ at large context (dual RTX Pro 6000) -- 2026-06-25
  462. The Eagle(3) has landed (for Qwen) -- 2026-06-25
  463. CPU-only TTS benchmark: Kokoro 82M vs Supertonic 3 vs Inflect-Nano-v1 (4.6M params), with UTMOS scoring on every sample -- 2026-06-25
  464. Ideogram 4: Open Image Model at the Forefront of Design -- 2026-06-25
  465. QUEST-35B: Open-Source Deep Research Agent Trained on 32 H100s — Full Recipe, Weights, and Data Released -- 2026-06-25
  466. North Mini Code: 4-Bit Quant + Ollama + OpenRouter — Now Runs on 20GB -- 2026-06-25
  467. EdgeRazor: Mixed-Precision Quantization-Aware Distillation Down to 1.58-Bit -- 2026-06-25
  468. [Editorial] Qwable-3.6-27b — Open Model Release -- 2026-06-25
  469. ScenemaAI/scenema-audio — Audio Generation Model -- 2026-06-25
  470. "Tokaine Addiction" — PhD Student's Colleague Can't Stop Running AI Agents -- 2026-06-25
  471. Suitcase Robot Uses Gas Sensor to Modulate LLM Sampler Temperature in Real-Time -- 2026-06-25
  472. GLM-5.2 is the first open-weights model to cross 80% on Terminal-Bench -- 2026-06-19
  473. unsloth GLM-5.2-GGUF, including 2bit at 238GB -- 2026-06-19
  474. Running local models is good now -- 2026-06-19
  475. Gemma 12b less than 10 watts 6.5pp 1.3tg -- 2026-06-19
  476. Rio 3.5 397B could've simply been a semi-failed embezzling of funding -- 2026-06-19
  477. VibeThinker-3B: what is this witchcraft? Killing it at MathQA like it has ~30B parameters -- 2026-06-18
  478. Get in here: Community model build thread -- 2026-06-18
  479. bartowski/command-a-plus-05-2026-GGUF · Hugging Face -- 2026-06-18
  480. nvidia/Nemotron-Labs-Diffusion-14B -- 2026-06-18
  481. [Editorial] Microsoft Majorana 2: Quantum Discovery via Agentic AI -- 2026-06-17
  482. cuTile Rust: Safe, Data-Race-Free GPU Kernels in Rust -- 2026-06-17
  483. [Editorial] Adversarial AI Research — The Malicious Use of Artificial Intelligence -- 2026-06-17
  484. Heretic Grimoire 1.4: Takedown-Resilient Model Backup — 9KB Reproducible Manifests + IPFS Distribution -- 2026-06-16
  485. z.ai Poll: MIT-Licensed Open Weights Are Losing -- 2026-06-16
  486. [Editorial] AMD CEO Lisa Su Challenges NVIDIA's GPU Dominance -- 2026-06-16
  487. PonyExl3: EXL3 Quantization Ported to Apple Silicon — 2700 tok/s Prefill, 68.5 tok/s Decode on M5 Max -- 2026-06-16
  488. My Homelab AI Dev Platform -- 2026-06-16
  489. [Editorial] Video Content -- 2026-06-16
  490. [Editorial] Video Content -- 2026-06-16
  491. [Editorial] Video Content -- 2026-06-16
  492. [Editorial] StandardAgents Arrow-JS — JavaScript Agent Framework -- 2026-06-16
  493. archex: Local-First Deterministic Code-Context for AI Agents — No API Key, No Telemetry (Apache 2.0) -- 2026-06-16
  494. Ironsmith: Open Source macOS App That Creates macOS Apps From Prompts — Works With Local Models -- 2026-06-16
  495. Claude Mythos 5 + Fable 5 Are Here And The Numbers Are INSANE -- 2026-06-10
  496. What it feels like to work with Mythos -- 2026-06-10
  497. [Editorial] Reuven Cohen on maximizing AI tools -- 2026-06-10
  498. How much of Thermo Fisher's antibody data has been manipulated? -- 2026-06-09
  499. News Sites are Blocking Internet Archive over AI Scraping Fears -- 2026-06-09
  500. OneDrive data now has an expiry date -- 2026-06-09
  501. Dopamine Fracking -- 2026-06-09
  502. ESP32 Bit Pirate, a Hardware Hacking Tool with WebCLI That Speaks Every Protocol -- 2026-06-09
  503. Porting the ThinkPad X61 to Coreboot -- 2026-06-09
  504. Holo3.1: Fast & Local Computer Use Agents -- 2026-06-03
  505. Fused MoE dispatch kernel in pure Triton: 89-131% of Megablocks, runs on AMD with zero code changes -- 2026-06-03
  506. ReAligned-Qwen3.5 Release -- 2026-06-03
  507. KANX: A production-ready Kolmogorov-Arnold Network library -- 2026-06-03
  508. Introducing Mellum2: A 12B Mixture-of-Experts Model by JetBrains -- 2026-06-02
  509. Tencent Hy-MT2 is now under Apache License 2.0 -- 2026-06-02
  510. fxyz666/LogicPipe -- 2026-06-02
  511. [OSS] dlmserve - first serving engine for diffusion language models -- 2026-06-02
  512. veryyoldman/Genspark-AI -- 2026-06-02
  513. Qwen/Qwen-Image-Bench · Hugging Face -- 2026-06-02
  514. Welcome NVIDIA Cosmos 3: The First Open Omni-model for Physical AI Reasoning and Action -- 2026-06-02
  515. Nvidia announces new AI chip for personal computers -- 2026-06-02
  516. Nvidia LocateAnything - Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding. (10x faster than Qwen3-VL) -- 2026-06-02
  517. The $500K AI Film That "Premiered at Cannes" Was Not in the Official Festival -- 2026-06-02
  518. [Editorial] Barracuda Nightmare Eclipse Zero-Days -- 2026-05-29
  519. [Editorial] Video -- 2026-05-29
  520. Krasis update: Qwen3.6-35B-A3B (Q4) at reading speed, 1x 8GB 3070 Mobile laptop (32GB RAM) -- 2026-05-28
  521. Benchmarked Needle 26M vs Qwen3-0.6B on CPU function calling, 50 queries across 5 difficulty tiers. The 23x smaller model wins on accuracy and is 4.4x faster. -- 2026-05-28
  522. Small comparison on full compute performance (Anima) of 5090 vs 6000 PRO MaxQ vs 6000 PRO WS/SE -- 2026-05-28
  523. CXMT started selling ram to corsair -- 2026-05-28
  524. TTS Benchmark Comparison (all known TTS up until May 2026) -- 2026-05-28
  525. BitCPM-CANN: Native 1.58-Bit Large Language Model Training on Ascend NPU -- 2026-05-26
  526. Carbon: Decoding the Language of Life -- 2026-05-26
  527. AI content detector based on Qwen 0.8b fine-tuned on Pangram dataset -- 2026-05-26
  528. A tool I built to generate 3D objects with functional, articulated parts. It's on github, and is mostly LLM-agnostic. -- 2026-05-26
  529. Show HN: Audiomass – a free, open-source multitrack audio editor for the web -- 2026-05-26
  530. CAM-LDS: Cyber Attack Manifestations for Automatic Interpretation of System Logs and Security Alerts -- 2026-05-26
  531. aaron-kidwell/goLoL -- 2026-05-26
  532. [Editorial] -- 2026-05-26
  533. lordx64/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled -- 2026-05-22
  534. Jackrong/Qwopus3.5-9B-Coder-GGUF · Hugging Face -- 2026-05-22
  535. Qwen 3.6 35B GGUF: NTP vs MTP quantization results across GPUs and CPUs -- 2026-05-22
  536. [WIP] Gemma 4 MTP -- 2026-05-22
  537. Lemonade v10.5.1: an MTP + ROCm 7.13 quick start for Strix Halo -- 2026-05-22
  538. gemma-4-Ortenzya-The-Creative-Wordsmith-31B-it-uncensored-heretic is Out Now -- 2026-05-22
  539. Gemma-4-Gembrain-31B-it-uncensored-heretic Is Out Now -- 2026-05-22
  540. Newbie vibe coding experience: Shifting from Claude Sonnet 4.6 to Qwen3.6-35B-A3B-UD-Q6_K -- 2026-05-22
  541. HalBench: Open Sycophancy and Hallucination Benchmark — 12,800 Graded Responses -- 2026-05-21
  542. DystopiaBench: 42 LLMs tested on willingness to build dystopian systems -- 2026-05-21
  543. Guardrails take an 8B model from 53% to 99% on agentic tasks [ACM CAIS '26 preprint] -- 2026-05-21
  544. Introducing the Ettin Reranker Family -- 2026-05-20
  545. OlmoEarth v1.1: A more efficient family of models -- 2026-05-20
  546. Mini Shai-Hulud Strikes Again: 314 npm Packages Compromised -- 2026-05-20
  547. [Editorial] -- 2026-05-20
  548. exploitbench/exploitbench -- 2026-05-20
  549. bytedance released an open source model that attempts to do just about anything with only 3b parameters -- 2026-05-19
  550. Sapient Intelligence releases HRM-Text 1B: 40B tokens, ~$1k pretrain, beats Llama3.2 3B on MATH and DROP -- 2026-05-19
  551. PaddleOCR 3.5: Running OCR and Document Parsing Tasks with a Transformers Backend -- 2026-05-19
  552. DeepSeek-V4-Flash W4A16+FP8 with MTP self-speculation: 85 tok/s @ 524k on 2x RTX PRO 6000 Max-Q -- 2026-05-15
  553. NCCL-Free Tensor Parallelism on Dual Blackwell PCIe — llama.cpp b9095 -- 2026-05-15
  554. The Qwen 3.6 35B A3B hype is real!!! -- 2026-05-15
  555. A First Comprehensive Study of TurboQuant: Accuracy and Performance -- 2026-05-15
  556. Developing Open Source LLM from Ground Up — DeepSeek V3 Architecture on Single Blackwell GPU -- 2026-05-15
  557. z-lab/Qwen3.6-27B-DFlash -- 2026-05-15
  558. [Editorial] Arxiv Research Paper 2603.15423 -- 2026-05-11
  559. [Editorial] Arxiv Research Paper 2603.17378 -- 2026-05-11
  560. [Editorial] Gadi Evron — Predicting AI Trends -- 2026-05-11
  561. iai-mcp: Persistent Claude Memory Daemon — 5 Months of Daily Use, Now Open Source -- 2026-05-08
  562. [Editorial] Claude Skins — Custom Claude Code Personas -- 2026-05-08
  563. [Editorial] AI Agents Projects & Tutorials Collection -- 2026-05-08
  564. Does the "6 months gap" still hold? -- 2026-05-07
  565. Claude Code @ Opus 4.7 vs OpenCode @ qwen3.6:27b. Both shipped a playable cozy roguelite. -- 2026-05-07
  566. Fine-tuned Qwen3.6-35B-A3B DeltaNet experiment -- 2026-05-07
  567. Zyphra/ZAYA1-8B -- 2026-05-07
  568. Virtual violin produces realistic sounds (MIT) -- 2026-05-06
  569. Lightricks/LTX-2.3-22b-IC-LoRA-HDR -- 2026-05-06
  570. "Second Thoughts" — A small transformer reads output and feeds it back as a refinement loop, drastically improving a 1.7B model's coding -- 2026-05-06
  571. A plug-n-play open-source pruning tool that is workload-aware (Sculpt) -- 2026-05-06
  572. "I" is not singular — 4 LLM agents with per-agent LoRA on a single RTX 3070 8GB -- 2026-05-06
  573. I made a visualizer for Hugging Face models (hfviewer.com) -- 2026-05-06
  574. Humanoid Robot Actuators -- 2026-05-05
  575. Qwen-Scope: Official Sparse Autoencoders (SAEs) for Qwen 3.5 models -- 2026-05-05
  576. guidelabs/steerling-8b — Steerable 8B Model -- 2026-05-05
  577. nvidia/Gemma-4-31B-IT-NVFP4 — NVIDIA FP4 Quantized Gemma 4 -- 2026-05-05
  578. Qwen 3.6 27B Neo Code Q4_KM running tax accounting on a Ryzen laptop -- 2026-05-05
  579. Qwen3.6-27B gets stuck in a self-affirming thinking loop -- 2026-05-05
  580. Mistral Medium 3.5 128B — MLX 4-bit Conversion with Vision and 256K Context -- 2026-05-04
  581. Chirp: Native Offline Text-to-Speech Desktop App (Kokoro + Qwen3-TTS) -- 2026-05-04
  582. TinyMozart v2 85M — Unconditional MIDI Piano Music Generation -- 2026-05-04
  583. SuperGemma4-26B Uncensored MLX 4-bit v2 -- 2026-05-04
  584. Open WebUI Skill for Auto-Creating Tools with Qwen3.6/Gemma4 -- 2026-05-04
  585. quant-whisper: Terminal-Native Algo Trading Engine with Local LLM Inference -- 2026-05-04
  586. [Editorial] Finding Zero-Days with Any Model -- 2026-05-01
  587. Qwen 3.6 27B Makes Huge Gains in Agency on Artificial Analysis - Ties with Sonnet 4.6 -- 2026-04-30
  588. First direct side by side MoE vs Dense comparison -- 2026-04-30
  589. GPT-6 Confirmed -- 2026-04-30
  590. Google to invest up to $40 billion in Anthropic as search giant spreads its AI bets -- 2026-04-30
  591. convert : add support for Nemotron Nano 3 Omni by danbev - llama.cpp PR #22481 -- 2026-04-30
  592. SWE-bench Verified no longer measures frontier coding capabilities -- 2026-04-29
  593. Opus 4.7: Are these first signs of model collapse? -- 2026-04-29
  594. Qwen3.6-27B IQ4_XS FULL VRAM with 110k context -- 2026-04-28
  595. Can we already use Google's TurboQuant (TQ) for KV Cache in llama-server? Or are we waiting for a PR? -- 2026-04-28
  596. Gemma 4 beats Qwen 3.5 (UPDATE), and Qwen 3.6 27B + MiniMax M2.7 is the best OpenCode setup -- 2026-04-28
  597. Tried Qwen3.6-27B-UD-Q6_K_XL.gguf with CloudeCode, well I can't believe but it is usable -- 2026-04-28
  598. BitNet is the AI future? -- 2026-04-28
  599. How Anthropic's Model Context Protocol Allows for Easy Remote Execution -- 2026-04-27
  600. [Editorial] LinkedIn: AI Industry Perspective -- 2026-04-27
  601. [Editorial] PolinRider — Open Source Malware Analysis -- 2026-04-27
  602. FP4 inference in llama.cpp (NVFP4) and ik_llama.cpp (MXFP4) landed -- 2026-04-27
  603. [Editorial] Bonsai-8B MLX 1-bit -- 2026-04-27
  604. VRAM.cpp: Running llama-fit-params directly in your browser -- 2026-04-27
  605. Thoughts on using an AMD Alveo V80 FPGA as a poor man's Taalas HC1 -- 2026-04-27
  606. [Editorial] hw-smi — Cross-Platform Hardware Monitor -- 2026-04-27
  607. China's DeepSeek valuation rockets above $20B!! -- 2026-04-24
  608. [Editorial] DeepSeek Open-Sources Tile Kernels -- 2026-04-24
  609. DeepSeek v4 -- 2026-04-24
  610. NSA is using Anthropic's Mythos despite blacklist -- 2026-04-22
  611. [Editorial] Mad Bugs: Reverse Engineering Deep Dive -- 2026-04-22
  612. Anyone deployed Kimi K2.6 on their local hardware? -- 2026-04-21
  613. 24/7 Headless AI Server on Xiaomi 12 Pro (Guide & Benchmarks) Gemma4 VS Qwen2.5 -- 2026-04-21
  614. Gemm4:e4B-IT good at instructions following no refusals. -- 2026-04-21
  615. [Editorial] arxiv:2506.02153 — AI Research -- 2026-04-20
  616. [Editorial] arxiv:1503.02531 — Classic ML/AI Paper -- 2026-04-20
  617. [Editorial] arxiv:2604.06169 — Recent AI Research -- 2026-04-20
  618. [Editorial] Why Your LLM Is Slow (and the Eight Fixes) -- 2026-04-20
  619. Hot Experts in your VRAM! Dynamic expert cache in llama.cpp for 27% faster token generation -- 2026-04-20
  620. Gemma 4 26B fabricated an entire code audit. I have the forensic evidence from the database. -- 2026-04-16
  621. [Editorial] Looped LLMs Are the Nuclear Fusion of AI -- 2026-04-16
  622. All elementary functions from a single binary operator -- 2026-04-14
  623. GLM 5.1 "Actually wait" — the current thinking SOTA open source -- 2026-04-14
  624. Are i-Quants overrated? -- 2026-04-14
  625. Gemma 4 E4B vs Qwen3.5-4B on document tasks — sub-scores tell a different story -- 2026-04-14
  626. Using NPU for something useful — whisper-npu for Intel NPU speech-to-text -- 2026-04-14
  627. [Editorial] -- 2026-04-14
  628. [Editorial] -- 2026-04-13
  629. Qwen3.5-397B is shockingly useful at Q2 -- 2026-04-13
  630. [Editorial] -- 2026-04-13
  631. Liquid AI releases LFM2.5-VL-450M - structured visual understanding at 240ms -- 2026-04-13
  632. Muse Spark: Scaling towards personal superintelligence -- 2026-04-13
  633. [Editorial] -- 2026-04-13
  634. [Editorial] -- 2026-04-13
  635. Abliterating Qwen3.5-397B on a Mac Studio revealed that MoE models encode refusal differently than dense models — safety refusals route through expert selection and survive weight-baking -- 2026-04-09
  636. Finally Abliterated Sarvam 30B and 105B! -- 2026-04-09
  637. Sam Altman may control our future – can he be trusted? -- 2026-04-07
  638. Qwen3.6-Plus -- 2026-04-07
  639. Omnivoice - 600+ Language Open-Source TTS with Voice Cloning and Design -- 2026-04-07
  640. [Editorial] Generative Models and High-Quality Output -- 2026-04-07
  641. A cryptography engineer's perspective on quantum computing timelines -- 2026-04-07
  642. [Editorial] -- 2026-03-28
  643. [Editorial] -- 2026-03-28
  644. [Editorial] -- 2026-03-28
  645. [Editorial] -- 2026-03-28
  646. [Editorial] -- 2026-03-28
  647. The current state of the Chinese LLMs scene -- 2026-03-26
  648. Alibaba confirms they are committed to continuously open-sourcing new Qwen and Wan models -- 2026-03-26
  649. Cursor's Composer 2 apparently built on Kimi K2.5 without attribution -- 2026-03-26
  650. Nemotron Cascade 2 30B A3B -- 2026-03-26
  651. Controllable Reasoning Models Are Private Thinkers -- 2026-03-26
  652. Nvidia built a silent opinion engine into NemotronH to gaslight you and they're not the only ones doing it -- 2026-03-26
  653. Gemini thoughts turned very violent -- 2026-03-26
  654. Qianfan-OCR: End-to-End 4B Document Intelligence VLM — SOTA on OmniDocBench -- 2026-03-24
  655. AMD Quark: Under-the-Radar Quantization Tool with MXFP4 Post-Training -- 2026-03-24
  656. We Compressed 6 LLMs and Found They Don't Degrade the Same Way -- 2026-03-24
  657. HumeAI/tada-1b — Expressive Audio Generation Model -- 2026-03-24
  658. The Resolv Hack: How One Compromised Key Printed $23M -- 2026-03-24
  659. Show HN: Three new Kitten TTS models – smallest less than 25MB -- 2026-03-23
  660. Flash-MoE: Running a 397B Parameter Model on a Laptop -- 2026-03-23
  661. [Editorial] Distillation Techniques -- 2026-03-23
  662. Jackrong/Qwen3.5-2B-Claude-4.6-Opus-Reasoning-Distilled-GGUF -- 2026-03-23
  663. Local Qwen 8B + 4B completes browser automation by replanning one step at a time -- 2026-03-23
  664. Does imatrix calibration data affect writing style? I ran a blind-scored experiment -- 2026-03-23
  665. [Editorial] Claude Code Channels Documentation -- 2026-03-23
  666. github.com -- 2026-03-23
  667. Claude Channels vs Dispatch vs Remote Control -- 2026-03-23
  668. [Editorial] Claude Code Is Brilliant Until the Repo... -- 2026-03-23
  669. I made mcp-optimizer - stop wasting tokens on idle MCP servers -- 2026-03-23
  670. [Editorial] Anthropic Is Coming for Lovable et al. -- 2026-03-23
  671. [Editorial] -- 2026-03-20
  672. mlx-tune – fine-tune LLMs on your Mac (SFT, DPO, GRPO, Vision) with an Unsloth-compatible API -- 2026-03-20
  673. Hunter Alpha was a stealth model revealed on March 18th as an early testing version of MiMo-V2-Pro. -- 2026-03-19
  674. mistralai/Mistral-Small-4-119B-2603 -- 2026-03-19
  675. LiquidAI/LFM2-24B-A2B -- 2026-03-19
  676. Mistral releases an official NVFP4 model, Mistral-Small-4-119B-2603-NVFP4! -- 2026-03-18
  677. Omnicoder-Claude-4.6-Opus-Uncensored-GGUF -- 2026-03-18
  678. You can run LLMs on your AMD NPU on Linux! -- 2026-03-18
  679. Rick Beato: "How AI Will Fail Like The Music Industry" (and why local LLMs will win) -- 2026-03-18
  680. [Editorial] Your AI Forgot Everything Again — There's a Fix for That -- 2026-03-18
  681. ylytdeng/wechat-decrypt -- 2026-03-18
  682. [Editorial] LLM Architecture Gallery -- 2026-03-17
  683. Mistral small 4 PR on transformers. -- 2026-03-17
  684. I spent a weekend doing layer surgery on 6 different model architectures. There's a "danger zone" at ~50-56% depth. -- 2026-03-17
  685. Nemotron 3 Super and the no free lunch problem -- 2026-03-17
  686. Source code of Swedish e-government services has been leaked -- 2026-03-14
  687. Bucketsquatting is (finally) dead -- 2026-03-14
  688. E2E encrypted messaging on Instagram will no longer be supported after 8 May -- 2026-03-14
  689. Fine-tuned Qwen3 SLMs (0.6-8B) beat frontier LLMs on narrow tasks -- 2026-03-12
  690. Towards Cold-Start Drafting and Continual Refining: A Value-Driven Memory Approach with Application to NPU Kernel Synthesis -- 2026-03-12
  691. Heretic Defeats GPT-OSS with Arbitrary-Rank Ablation (ARA) Decensoring -- 2026-03-11
  692. Fat Fish — A Proper Upscale and Prune of Mistral Nemo -- 2026-03-11
  693. Optimizing Qwen3 Coder for RTX 5090 and PRO 6000 — Community Benchmarking Infrastructure -- 2026-03-11
  694. Yann LeCun's AMI Labs Raises $1.03 Billion to Build World Models -- 2026-03-11
  695. Did Alibaba just kneecap its powerful Qwen AI team? -- 2026-03-07
  696. [Editorial] You're 1,191 Days Late — Here's What to Do -- 2026-03-07
  697. [Editorial] When Anonymity Fades: What New Research Reveals -- 2026-03-07
  698. Introducing Modular Diffusers - Composable Building Blocks for Diffusion Pipelines -- 2026-03-07
  699. kyutai-labs/hibiki-zero -- 2026-03-07
  700. [Editorial] OpenClawCity -- 2026-03-07
  701. [Editorial] The Zero-Day Clock Is Ticking -- 2026-03-06
  702. [Editorial] Step-by-Step Guide to Exploiting AI Systems -- 2026-03-06
  703. [Editorial] Unprompted 2026: Top Insights Day One -- 2026-03-06
  704. [Editorial] Unprompted 2026: Top Insights Day Two -- 2026-03-06
  705. YuanLabAI/Yuan3.0-Ultra: 1010B MoE, fully open weights -- 2026-03-05
  706. We could be hours (or less than a week) away from true NVFP4 support in Llama.cpp GGUF format -- 2026-03-05
  707. Step-3.5-Flash-Base & Midtrain (in case you missed them) -- 2026-03-05
  708. Qwen3.5-9B Uncensored Aggressive Release (GGUF) -- 2026-03-05
  709. unknown -- 2026-03-05
  710. [Editorial] David Maynor Security Gist -- 2026-03-04
  711. [Editorial] arXiv:2602.23093 -- 2026-03-04
  712. unpromptedcon.org -- 2026-03-04
  713. Unsloth fixed version of Qwen3.5-35B-A3B is incredible at research tasks -- 2026-03-04
  714. PSA: Qwen 3.5 requires bf16 KV cache, NOT f16!! -- 2026-03-04
  715. Building a Dependency-Free GPT on a Custom OS -- 2026-03-04
  716. Qwen3.5 122B in 72GB VRAM (3x3090) is the best model available at this time — also it nails the "car wash test" -- 2026-03-03
  717. Qwen 3.5 is multimodal. Here is how to enable image understanding in opencode with llama cpp -- 2026-03-03
  718. pplx-embed: State-of-the-Art Embedding Models for Web-Scale Retrieval -- 2026-03-03
  719. andimarafioti/faster-qwen3-tts -- 2026-03-03
  720. We tested RLVR on top of fine-tuned small models across 12 datasets — here's exactly when it helps (and when it doesn't) -- 2026-03-02
  721. [Editorial] Weber Electrodynamics & Thermodynamic Memory -- 2026-03-02
  722. Minuspod: Automatically remove ads from podcasts locally -- 2026-03-02
  723. Squidcasa/midipipe: ALSA Sequencer to plain text and back -- 2026-03-02
  724. [Editorial] -- 2026-02-28
  725. Parakeet.cpp – Parakeet ASR inference in pure C++ with Metal GPU acceleration -- 2026-02-28
  726. [Editorial] -- 2026-02-28
  727. Liquid AI releases LFM2-24B-A2B -- 2026-02-24
  728. Qwen3's most underrated feature: Voice embeddings -- 2026-02-24
  729. After many contributions craft, Crane now officially supports Qwen3-TTS! -- 2026-02-24
  730. We tested the same INT8 model on 5 Snapdragon chipsets. Accuracy ranged from 93% to 71%. Same weights, same ONNX file. -- 2026-02-24
  731. [Editorial] -- 2026-02-24
  732. Anthropic Accuses DeepSeek, Moonshot AI, and MiniMax of Creating 24,000 Fake Claude Accounts -- 2026-02-24
  733. Qwen 3.5 vs Gemini 3 Pro on Screenshot-to-Code: Is the gap finally gone? -- 2026-02-19
  734. [Editorial] The evolution of vision models from CNNs to transformers -- 2026-02-19
  735. [Editorial] An AI agent merged code into 22 widely-used open source projects -- 2026-02-19
  736. [Editorial] AI Agent Security and Supply Chain -- 2026-02-19
  737. Policy Compiler for Secure Agentic Systems -- 2026-02-19
  738. [Editorial] OpenClaw Maestro Threat Assessment -- 2026-02-19
  739. Qwen Released Qwen 3.5 397B and Qwen 3.5 Plus! -- 2026-02-17
  740. Qwen3.5 NVFP4 (Blackwell) is up! -- 2026-02-17
  741. Running Gemma 3n E2B natively on Android via LiteRT -- 2026-02-17
  742. Deploying Open WebUI + vLLM on Amazon EKS -- 2026-02-17
  743. Ling-2.5-1T: 1T Parameter Open-Source Instant Model with 1M Context -- 2026-02-16
  744. Qwen3.5-397B-A17B Unsloth GGUFs — Run on Consumer Hardware -- 2026-02-16
  745. Running Qwen3-Coder-Next 80B on 8GB VRAM — 300x Speedup via Custom Expert Caching -- 2026-02-16
  746. [Editorial] https://www.linkedin.com/posts/ownyourai_i-just-woke-up-to-qwen3-coder-next-80b-activity-7424703876240695297-Nlqf -- 2026-02-04
  747. LiquidAI/LFM2.5-1.2B-Thinking -- 2026-02-04
  748. ByteDance-Seed/Stable-DiffCoder-8B-Instruct · Hugging Face -- 2026-02-03
  749. tencent/Youtu-VL-4B-Instruct -- 2026-02-03
  750. NousResearch/NousCoder-14B -- 2026-02-03
  751. transformers v5 final is out 🔥 -- 2026-01-30
  752. PaddlePaddle/PaddleOCR-VL-1.5 -- 2026-01-30
  753. zai-org/GLM-4.7 -- 2026-01-30
  754. MHA2MLA-VLM: Enabling DeepSeek's Economical Multi-Head Latent Attention across Vision-Language Models -- 2026-01-30
  755. Sharing my set of distilled small language models (3B) + training data in more than 50 low-resource languages -- 2026-01-30
  756. Introducing Kimi K2.5, Open-Source Visual Agentic Intelligence -- 2026-01-28
  757. ~60GB models on coding: GLM 4.7 Flash vs. GPT OSS 120B vs. Qwen3 Coder 30B -- your comparisons? -- 2026-01-28
  758. openbmb/AgentCPM-Report -- 2026-01-28
  759. Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice -- 2026-01-28
  760. Mixture-of-Models: Unifying Heterogeneous Agents via N-Way Self-Evaluating Deliberation -- 2026-01-28
  761. One Year Since the “DeepSeek Moment” -- 2026-01-28
  762. Qwen/Qwen-Image-2512 -- 2026-01-28
  763. Qwen/Qwen3-TTS-12Hz-0.6B-Base -- 2026-01-27
  764. Qwen/Qwen3-VL-Reranker-2B -- 2026-01-27
  765. Qwen/Qwen3-VL-Embedding-2B -- 2026-01-21
  766. Phr00t/Qwen-Image-Edit-Rapid-AIO -- 2026-01-21
  767. Qwen/Qwen3-VL-Reranker-8B -- 2026-01-21
  768. Bartowski comes through again. GLM 4.7 flash GGUF -- 2026-01-21
  769. Step-Audio-R1.1 (Open Weight) by StepFun just set a new SOTA on the Artificial Analysis Speech Reasoning leaderboard -- 2026-01-20
  770. GLM-4.7-Flash -- 2026-01-20
  771. stepfun-ai/Step-Audio-R1.1 -- 2026-01-20
  772. tencent/HY-MT1.5-1.8B -- 2026-01-20
  773. Introducing GLM-Image -- 2026-01-14
  774. GPT-OSS -> MLA conversion breakthrough (20B), still looking for compute + collaborators -- 2026-01-14
  775. FrogBoss 32B and FrogMini 14B from Microsoft -- 2026-01-14
  776. Qwen/Qwen3-VL-Embedding-8B -- 2026-01-14
  777. Liquid AI releases LFM2-2.6B-Transcript, an incredibly fast open-weight meeting transcribing AI model on-par with closed-source giants. -- 2026-01-13
  778. New llama.cpp 30% faster.... -- 2026-01-13
  779. Qwen/Qwen-Image-Edit-2511 -- 2026-01-13
  780. nvidia/Nemotron-Orchestrator-8B -- 2026-01-13
  781. tencent/HY-WorldPlay -- 2026-01-09
  782. meituan-longcat/LongCat-Image -- 2026-01-09
  783. Introducing Falcon H1R 7B -- 2026-01-09
  784. [Editorial] https://github.com/hiyouga/LlamaFactory -- 2026-01-07
  785. Tongyi-MAI/MAI-UI-8B · Hugging Face -- 2026-01-07
  786. MultiverseComputingCAI/HyperNova-60B · Hugging Face -- 2026-01-07
  787. Anyone tried IQuest-Coder-V1 yet? The 40B numbers look wild -- 2026-01-06
  788. open-thoughts/OpenThinker-Agent-v1 -- 2026-01-06
  789. [Editorial] https://www.alwaysfurther.ai/blog/train-4b-model-to-beat-claude-sonnet-gemini -- 2026-01-05
  790. [Experimental] Gemma 3 4B - Dark CoT: Pushing 4B Reasoning to 33%+ on GPQA Diamond -- 2026-01-05
  791. Youtu-LLM-2B-GGUF is here! -- 2026-01-05
  792. upstage/Solar-Open-100B -- 2026-01-05
  793. A zero-setup agent that benchmarks multiple open / closed source LLMs on your specific problem / data -- 2026-01-02
  794. What is a good model for assisting with patching source code? -- 2026-01-02
  795. Just got an RTX Pro 6000 - need recommendations for processing a massive dataset with instruction following -- 2026-01-02
  796. MiniMaxAI/MiniMax-M2.1 -- 2026-01-02
  797. [Editorial] https://www.linkedin.com/posts/ownyourai_wow-its-raining-korean-open-ai-models-today-activity-7412133834667876352-fcpl -- 2025-12-31
  798. Naver (South Korean internet giant), has just launched HyperCLOVA X SEED Think, a 32B open weights reasoning model and HyperCLOVA X SEED 8B Omni, a unified multimodal model that brings text, vision, and speech together -- 2025-12-31
  799. allenai/Olmo-3-7B-Instruct -- 2025-12-31
  800. meituan-longcat/LongCat-Video-Avatar -- 2025-12-31
  801. Mistral AI’s December -- 2025-12-31
  802. MBZUAI releases K2-V2 - 70B fully open model. -- 2025-12-22
  803. microsoft/Fara-7B -- 2025-12-22
  804. I've been experimenting with SLM's a lot recently. My goal was to prove even SLMs can be accurate with the right architecture behind it. -- 2025-12-22
  805. Alibaba Tongyi Open Sources Two Audio Models: Fun-CosyVoice 3.0 (TTS) and Fun-ASR-Nano-2512 (ASR) -- 2025-12-19
  806. My professor lent me an A6000, so I tried to build a coding model. Here is Anni! (Qwen3-14B Fine-tune) -- 2025-12-19
  807. Two years ago, I was just a math major. Now I've built the 1.5B router model used by HuggingFace. Can I bring it to Cursor? -- 2025-12-19
  808. 🚀 New: Olmo 3.1 Think 32B & Olmo 3.1 Instruct 32B -- 2025-12-16
  809. ByteDance/Dolphin-v2 -- 2025-12-16
  810. zai-org/AutoGLM-Phone-9B-Multilingual -- 2025-12-16
  811. [Editorial] https://www.linkedin.com/posts/eric-vyacheslav-156273169_a-7m-model-just-surpassed-deepseek-r1-gemini-activity-7405985266043297792-s1Jn -- 2025-12-15
  812. Nemotron 3 Nano \- A new Standard for Efficient, Open, and Intelligent Agentic Models -- 2025-12-15
  813. Tiny-A2D: An Open Recipe to Turn Any AR LM into a Diffusion LM -- 2025-12-12
  814. New in llama.cpp: Model Management -- 2025-12-12
  815. SoK: a Comprehensive Causality Analysis Framework for Large Language Model Security -- 2025-12-12
  816. Generating synthetic test data for LLM applications (our approach) -- 2025-12-12
  817. PaCoRe: The first open-source deep think 8B model beats GPT-5 on HMMT25 -- 2025-12-11
  818. Building RNJ-1: What makes It different from Gemma 3? -- 2025-12-11
  819. EssentialAI/rnj-1 -- 2025-12-11
  820. VulnLLM-R: Specialized Reasoning LLM with Agent Scaffold for Vulnerability Detection -- 2025-12-11
  821. From Azure Functions to FreeBSD -- 2025-12-10
  822. RnJ-1-Instruct FP8 Quantization -- 2025-12-10
  823. Operator Mech v2.5: A Compact Structural-Reasoning Kernel for Local Models (YAML, 7B–13B Optimized) -- 2025-12-10
  824. Masked Diffusion Models as Energy Minimization -- 2025-12-10
  825. https://huggingface.co/Doradus/Hermes-4.3-36B-FP8 -- 2025-12-09
  826. Support for rnj-1 now in llama.cpp -- 2025-12-09
  827. Comfy-Org/flux2-dev -- 2025-12-09
  828. baidu/ERNIE-4.5-VL-28B-A3B-Thinking -- 2025-12-09
  829. Guidance: A cheat code for diffusion models -- 2025-12-09
  830. OVHcloud on Hugging Face Inference Providers 🔥 -- 2025-12-05
  831. smallevals - Tiny 0.6B Evaluation Models and a Local LLM Evaluation Framework -- 2025-12-05
  832. I cooked abliterated gemma3-27b-it with norm-preserving technique -- 2025-12-04
  833. allenai/Olmo-3-1125-32B -- 2025-12-03
  834. EmbeddingGemma: Powerful and Lightweight Text Representations -- 2025-12-03
  835. Building SFT from scratch - results & learnings -- 2025-12-03
  836. Qwen3 VL built from scratch with PyTorch -- 2025-12-03
  837. LM Studio beta supports Qwen3 80b Next. -- 2025-12-03
  838. RTX 5090 + Qwen 30B MoE @ 135 tok/s in NVFP4 - Full guide with C++ patches -- 2025-12-02
  839. PleIAs/Baguettotron -- 2025-12-01
  840. nvidia/ChronoEdit-14B-Diffusers -- 2025-12-01
  841. 20x Faster TRL Fine-tuning with RapidFire AI -- 2025-12-01
  842. Optimizing Token Generation in llama.cpp's CUDA Backend -- 2025-12-01
  843. [Editorial] https://www.linkedin.com/posts/erudenko_claudecode-aiagents-opensource-activity-7399658246443216896-dIU0 -- 2025-11-28
  844. DataArcTech/DataArc-SynData-Toolkit -- 2025-11-28
  845. facebook/sam-3d-body-dinov3 -- 2025-11-28
  846. WeiboAI/VibeThinker-1.5B -- 2025-11-28
  847. In depth analysis of Nvidia's Jet Nemotron models -- 2025-11-26
  848. Hidden causes of LLM latency, its not just the model size -- 2025-11-26
  849. A million ways to die from a data race in Go -- 2025-11-26
  850. You're using HuggingFace wrong. Stop downloading pre-quantized GGUFs and start building hardware-optimized, domain-specific models. Here's the pipeline I built to do it. -- 2025-11-26
  851. Hardcore function calling benchmark in backend coding agent. -- 2025-11-26
  852. Luo-Yihao/FaithC -- 2025-11-20
  853. ReTabAD: A Benchmark for Restoring Semantic Context in Tabular Anomaly Detection -- 2025-11-20
  854. Android Developer Verification Starts as Google Partially Retreats on Measures -- 2025-11-20
  855. How are you all orchestrating multi-agent workflows (beyond one-shot prompt chaining)? -- 2025-11-20
  856. deliveryhero/asya -- 2025-11-20
  857. Do we rely too much on huggingface? Do you think they’ll eventually regulate open source models? Is there any way to distribute them elsewhere? -- 2025-11-18
  858. Build a DeepSeek model from scratch -- 2025-11-18
  859. Easily Build and Share ROCm Kernels with Hugging Face -- 2025-11-18
  860. ibm-granite/granite-4.0-h-350m -- 2025-11-14
  861. tencent/HunyuanWorld-Mirror -- 2025-11-14
  862. [Editorial] https://www.linkedin.com/posts/mreichstein_cybersecurity-carhacking-physicalsecurity-activity-7394425210877218816-NPyT -- 2025-11-13
  863. On USB HID, Keyboard LEDs, and device emulation (2024) -- 2025-11-13
  864. [Editorial] https://www.linkedin.com/posts/stuart-winter-tear_so-reportedly-yann-lecun-plans-to-leave-activity-7394396547276460032-gEE5 -- 2025-11-13
  865. inclusionAI/LLaDA2.0-mini-preview -- 2025-11-13
  866. [Editorial] https://www.linkedin.com/posts/ismaelvelasco_theres-an-ai-text-model-comparable-to-sota-activity-7393850964731912192-nZT1 -- 2025-11-12
  867. Last week in Multimodal AI - Local Edition -- 2025-11-12
  868. antarys-ai/antarys -- 2025-11-11
  869. Visualizing Quantization Types -- 2025-11-11
  870. Co-authored a book called "Build DeepSeek from Scratch" | Live Now -- 2025-11-11
  871. Writing your own BEAM -- 2025-11-11
  872. DIY Powerwall Blows Clouds, Competition Out of the Water -- 2025-11-11
  873. Building a PV Solar-Powered Quadcopter -- 2025-11-07
  874. BAAI/Emu3.5-Image -- 2025-11-07
  875. Riemannian Optimization for LoRA on the Stiefel Manifold -- 2025-11-07
  876. Kimi release Kimi K2 Thinking, an open-source trillion-parameter reasoning model -- 2025-11-06
  877. OpenAI asks U.S. for loan guarantees to fund $1T AI expansion -- 2025-11-06
  878. Kimi released Kimi K2 Thinking, an open-source trillion-parameter reasoning model -- 2025-11-06
  879. Reproducing the AWS Outage Race Condition with a Model Checker -- 2025-11-05
  880. Futurelock: A subtle risk in async Rust -- 2025-11-05
  881. Qwen3-VL-32B Q8 speeds in llama.cpp vs vLLM FP8 on a RTX PRO 6000 -- 2025-11-03
  882. Help me decide: EPYC 7532 128GB + 2 x 3080 20GB vs GMtec EVO-X2 -- 2025-11-03
  883. amd/Nitro-E -- 2025-11-03
  884. DeepSeek may have found a new way to improve AI’s ability to remember -- 2025-11-02
  885. Reliable Unlearning Harmful Information in LLMs with Metamorphosis Representation Projection -- 2025-11-02
  886. [Editorial] https://www.linkedin.com/posts/busiel-morley_economic-shifts-in-the-age-of-ai-ugcPost-7390349517612806144-8djS -- 2025-11-01
  887. OpenAI: gpt-oss-safeguard: two open-weight reasoning models built for safety classification (Now on Hugging Face) -- 2025-10-31
  888. briaai/FIBO -- 2025-10-31
  889. vLLM MoE Benchmark Configs for Qwen3 Coder REAP 25B & RTX Pro 6000 -- 2025-10-29
  890. Ollama supports Qwen3-VL locally! -- 2025-10-29
  891. Test results for various models' ability to give structured responses via LM Studio. Spoiler: Qwen3 won -- 2025-10-29
  892. Introducing ExecuTorch 1.0 -- 2025-10-29
  893. [Editorial] new coding model -- 2025-10-28
  894. Cerebras REAP'd GLM4.6: 25%, 30%, 40% pruned FP8 checkpoints on HF! -- 2025-10-28
  895. GLM-4.6 on fresh SWE-bench–style tasks collected in September 2025 -- 2025-10-28
  896. Best LLM for 96G RTX Pro 6000 Blackwell? -- 2025-10-27
  897. Gemma3 model differencies -- 2025-10-27
  898. AMD iGPU + dGPU : llama.cpp tensor-split not working with Vulkan backend -- 2025-10-27
  899. Is GLM 4.5 / 4.6 really sensitive to quantisation? Or is vLLM stupifying the models? -- 2025-10-27
  900. AlphaXiv,Compare the Deepseek-OCR and Mistral-OCR OCR models -- 2025-10-26
  901. Open-Bee/Bee-8B-RL -- 2025-10-26
  902. datalab-to/chandra -- 2025-10-26
  903. Unlock the power of images with AI Sheets -- 2025-10-26
  904. [Editorial] Periodic table for ai algorithms -- 2025-10-26
  905. Reverse Engineering STL Files with FreeCAD -- 2025-10-25
  906. Qwen3-VL-32B-Instruct GGUF with unofficial llama.cpp release to run it (Pre-release build) -- 2025-10-25
  907. Qwen3 Next support in llama.cpp ready for review -- 2025-10-25
  908. How's Halo Strix now ? -- 2025-10-25
  909. 20x Max Plan (€216) takes 2% of weekly Opus usage for a single Deep Research Query. That equals 50 per week if you use it ONLY for this and never continue or respond -- 2025-10-24
  910. Show HN: Cuq – Formal Verification of Rust GPU Kernels -- 2025-10-24
  911. [Editorial] We need open, uncensored, & local -- 2025-10-23
  912. 🚀 HuggingFaceChat Omni: Dynamic policy-baed routing to 115+ LLMs -- 2025-10-23
  913. Preliminary support in llama.cpp for Qualcomm Hexagon NPU -- 2025-10-23
  914. LiquidAI/LFM2-2.6B -- 2025-10-23
  915. riptideslabs/tokenex -- 2025-10-20
  916. Multi-Tenant SaaS's Wildcard TLS: An Overview of DNS-01 Challenges -- 2025-10-20
  917. From cloud to OCP? Be ready to wrangle firmware -- 2025-10-20
  918. FLOSS Weekly Episode 851: Buckets of Money -- 2025-10-20
  919. Learning Lifted Action Models From Traces of Incomplete Actions and States -- 2025-10-20
  920. volantvm/volant -- 2025-10-19
  921. Wireshark 4.6.0 Supports macOS Pktap Metadata (PID, Process Name, etc.) -- 2025-10-19
  922. A classified network of SpaceX satellites is emitting a mysterious signal -- 2025-10-19
  923. Show HN: Largest open-source multimodal AI dataset -- 2025-10-18
  924. KORMo-Team/KORMo-10B-sft -- 2025-10-18
  925. ByteDance/FaceCLIP -- 2025-10-18
  926. We built 3B and 8B models that rival GPT-5 at HTML extraction while costing 40-80x less - fully open source -- 2025-10-17
  927. inclusionAI/Ring-flash-linear-2.0 -- 2025-10-17
  928. Qwen/Qwen3-VL-8B-Instruct -- 2025-10-17
  929. Google Cloud C4 Brings a 70% TCO improvement on GPT OSS with Intel and Hugging Face -- 2025-10-17
  930. This Week in Security: ID Breaches, Code Smell, and Poetic Flows -- 2025-10-14
  931. Show HN: Rebuilt Bible search app to run 100% client-side with Transformers.js -- 2025-10-13
  932. swiss-ai/Apertus-8B-Instruct-2509 -- 2025-10-13
  933. A Childhood Dream, Created and Open Sourced -- 2025-10-13
  934. Introducing the ColBERT Nano series of models. All 3 of these models come in at less than 1 million parameters (250K, 450K, 950K) -- 2025-10-11
  935. LiquidAI/LFM2-8B-A1B -- 2025-10-11
  936. GPT-OSS from Scratch on AMD GPUs -- 2025-10-11
  937. How do I compare cost per token for serverless vs provisioned hardware? -- 2025-10-11
  938. OpenAI is good at deals -- 2025-10-11
  939. meituan-longcat/LongCat-Flash-Chat -- 2025-10-11
  940. adb1274/batchi -- 2025-10-11
  941. Granite4 Small-h 32b-A9b (Q4_K_M) at FULL 1M context window is using only 73GB of VRAM - Life is good! -- 2025-10-09
  942. Run Open AI GPT-OSS on a mobile phone (Demo) -- 2025-10-09
  943. AI21 releases Jamba 3B, the tiny model outperforming Qwen 3 4B and IBM Granite 4 Micro! -- 2025-10-09
  944. inclusionAI/Ling-mini-2.0 -- 2025-10-09
  945. Provable scaling laws of feature emergence from learning dynamics of grokking -- 2025-10-09
  946. SecureV2X: An Efficient and Privacy-Preserving System for Vehicle-to-Everything (V2X) Applications -- 2025-10-09
  947. [Editorial] The Tiny Recursive Mode -- 2025-10-08
  948. deepseek-ai/DeepSeek-V3.1-Terminus -- 2025-10-08
  949. Behavioral Modification Systems in Large Language Models: A Methodological Analysis of Long Conversation Reminders -- 2025-10-08
  950. Should I choose Ada, SPARK, or Rust over C/C++? (2024) -- 2025-10-08
  951. We built this open-source LLM Inference project to boost context generation by up to 15x and now it is being implemented by NVIDIA Dynamo! -- 2025-10-06
  952. Running Qwen3-VL-235B (Thinking & Instruct) AWQ on vLLM -- 2025-10-06
  953. Granite 4.0 Language Models - a ibm-granite Collection -- 2025-10-06
  954. NousResearch/Hermes-4-70B -- 2025-10-06
  955. ibm-granite/granite-4.0-h-small -- 2025-10-06
  956. ibm-granite/granite-4.0-h-tiny -- 2025-10-06
  957. [Editorial] Agentic Tribe -- 2025-10-06
  958. AlexanderYastrebov/onion-vanity-address -- 2025-10-06
  959. Open Printer is an open-source inkjet with DRM-free ink and no subscriptions -- 2025-10-06
  960. Yes, Gemini, A Wii Server Is Possible -- 2025-10-06
  961. Tencent-Hunyuan/Hunyuan3D-Omni -- 2025-10-04
  962. tencent/HunyuanImage-3.0 -- 2025-10-04
  963. swiss-ai/Apertus-8B-2509 -- 2025-10-04
  964. Use Remote Models on iOS with Noema -- 2025-10-04
  965. Best small model <3B for HomeAssistant -- 2025-10-04
  966. Nikity/lille-130m-instruct -- 2025-10-04
  967. Unsupervised Hallucination Detection by Inspecting Reasoning Processes -- 2025-10-03
  968. Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training -- 2025-10-03
  969. Will fine-tuning LLaMA 3.2 11B Instruct on text-only data degrade its vision capabilities? -- 2025-10-03
  970. Does Ollama immobilize GPUs / computing resources? -- 2025-10-02
  971. Any real alternatives to NotebookLM (closed-corpus only)? -- 2025-10-02
  972. Best instruct model that fits in 32gb VRAM -- 2025-10-02
  973. Bring Your Own Data (BYOD) -- 2025-09-30
  974. CohereLabs/command-a-reasoning-08-2025 -- 2025-09-30
  975. Built an MCP server for Claude Desktop to browse Reddit in real-time -- 2025-09-30
  976. 1652933138/eth-address-poisoning-tool -- 2025-09-30
  977. Inside NVIDIA GPUs: Anatomy of high performance matmul kernels -- 2025-09-29
  978. Bit is all we need: binary normalized neural networks -- 2025-09-29
  979. Accelerating Qwen3-8B Agent on Intel® Core™ Ultra with Depth-Pruned Draft Models -- 2025-09-29
  980. The MoE tradeoff seems bad for local hosting -- 2025-09-29
  981. MetalQwen3: Full GPU-Accelerated Qwen3 Inference on Apple Silicon with Metal Shaders – Built on qwen3.c - WORK IN PROGRESS -- 2025-09-28
  982. yangdongchao/UniAudio2 -- 2025-09-28
  983. jimsweb/aiMIDI -- 2025-09-28
  984. Handy – Free open-source speech-to-text app written in Rust -- 2025-09-28
  985. IndexTeam/IndexTTS-2 -- 2025-09-28
  986. beankeji-cloud/SLiteIO -- 2025-09-28
  987. GitHub - shantur/jarvis-mcp: Bring your AI to life—talk to assistants instantly in your browser. Zero hasle, No API keys, No Whisper -- 2025-09-27
  988. A1: Asynchronous Test-Time Scaling via Conformal Prediction -- 2025-09-25
  989. DeepLink-org/DeepTrace -- 2025-09-25
  990. Qwen 3 max released -- 2025-09-24
  991. Clauder, auto-updating toolkit for Claude Code, now ships with 65+ MCP servers -- 2025-09-24
  992. Show HN: I wrote inference for Qwen3 0.6B in C/CUDA -- 2025-09-24
  993. nvidia/canary-1b-v2 -- 2025-09-24
  994. baidu/ERNIE-4.5-21B-A3B-Thinking -- 2025-09-24
  995. Smol2Operator: Post-Training GUI Agents for Computer Use -- 2025-09-24
  996. Seeking Local LLM Recommendations for AST Generation (by Function Calling) -- 2025-09-24
  997. A first stab at packaging llama.cpp in a performance-optimized manner -- 2025-09-23
  998. Model: Qwen3 Next Pull Request llama.cpp -- 2025-09-23
  999. Efficient 4B parameter gpt OSS distillation without the over-censorship -- 2025-09-22
  1000. [Project] I created an AI photo organizer that uses Ollama to sort photos, filter duplicates, and write Instagram captions. -- 2025-09-22
  1001. Pointer Tagging in C++: The Art of Packing Bits into a Pointer -- 2025-09-22
  1002. inclusionAI/Ring-mini-2.0 -- 2025-09-22
  1003. Local real-time assistant that remembers convo + drafts a doc -- 2025-09-22
  1004. XiaomiMiMo/MiMo-Audio-7B-Instruct -- 2025-09-21
  1005. Scaling Self-Supervised Representation Learning for Symbolic Piano Performance -- 2025-09-21
  1006. unsloth/Qwen3-Next-80B-A3B-Instruct -- 2025-09-19
  1007. Gluon: a GPU programming language based on the same compiler stack as Triton -- 2025-09-19
  1008. Tesslate/WEBGEN-OSS-20B -- 2025-09-19
  1009. UUIDv47: Store UUIDv7 in DB, emit UUIDv4 outside (SipHash-masked timestamp) -- 2025-09-19
  1010. Beyond Prompt-Induced Lies: Investigating LLM Deception on Benign Prompts -- 2025-09-18
  1011. LFM2-1.2B safety benchmark -- 2025-09-18
  1012. Has anyone successfully gotten Ollama models (or any models) to execute SQL queries through natural language in Openwebui? -- 2025-09-18
  1013. ROCm 6.4.3 -> 7.0-rc1 after updating got +13.5% at 2xR9700 -- 2025-09-18
  1014. Is it possible for different brand GPUs to work together? -- 2025-09-18
  1015. Speculative cascades — A hybrid approach for smarter, faster LLM inference -- 2025-09-17
  1016. google/embeddinggemma-300m -- 2025-09-17
  1017. Running Qwen-Next (Instruct and Thinking) MLX BF16 with MLX-LM on Macs -- 2025-09-17
  1018. From Research to Reality: Feasibility of Gradient Inversion Attacks in Federated Learning -- 2025-09-17
  1019. Kwai-Klear/Klear-46B-A2.5B-Instruct -- 2025-09-16
  1020. WestZhang/VibeVoice-Large-pt -- 2025-09-16
  1021. FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference -- 2025-09-16
  1022. Qwen 3 Next Series – Qwen/Qwen3 Next 80B A3B Instruct Detected -- 2025-09-16
  1023. Test-time Prompt Intervention -- 2025-09-16
  1024. Effecient hot-swappable LoRA variant supported in llama.cpp -- 2025-09-12
  1025. Qwen/Qwen3-Next-80B-A3B-Instruct -- 2025-09-12
  1026. swiss-ai/Apertus-70B-Instruct-2509 -- 2025-09-12
  1027. GRASPED: Graph Anomaly Detection using Autoencoder with Spectral Encoder and Decoder (Full Version) -- 2025-09-12
  1028. stepfun-ai/step3 -- 2025-09-12
  1029. Exploring State-Space-Model based Language Model in Music Generation -- 2025-09-12
  1030. Small Open Models Achieve Near Parity with Large Models in Low Resource Literary Translation at a Fraction of the Cost -- 2025-09-10
  1031. Reasoning Vectors: Transferring Chain-of-Thought Capabilities via Task Arithmetic -- 2025-09-10
  1032. Tilde AI Releases TildeOpen LLM: An Open-Source Large Language Model with Over 30 Billion Parameters and Support Most European Languages -- 2025-09-09
  1033. Show HN: Open-sourcing our text-to-CAD app -- 2025-09-09
  1034. bytedance-research/USO -- 2025-09-09
  1035. LiquidAI/LFM2-VL-450M -- 2025-09-09
  1036. Fantastic pretraining optimizers and where to find them -- 2025-09-08
  1037. Qwen/Qwen3-4B-Instruct-2507 -- 2025-09-08
  1038. moonshotai/Kimi-K2-Instruct-0905 -- 2025-09-08
  1039. Tenstorrent p150a tested against RTX5090, RTX3090, A100, H100 by Russian blogger -- 2025-09-08
  1040. unsloth/gemma-3-270m-it-GGUF -- 2025-09-08
  1041. YanoljaNEXT-Rosetta: A Collection of Translation Models in Different Sizes -- 2025-09-07
  1042. Vulkan back ends, what do you use? -- 2025-09-06
  1043. A new OpenAI model? Could this be 5.1 or 5o? What do you think? -- 2025-09-06
  1044. VibeVoice RIP? What do you think? -- 2025-09-05
  1045. lodestones/Chroma1-HD -- 2025-09-05
  1046. Welcome EmbeddingGemma, Google's new efficient embedding model -- 2025-09-05
  1047. NousResearch/Hermes-4-405B -- 2025-09-05
  1048. I locally benchmarked 41 open-source LLMs across 19 tasks and ranked them -- 2025-09-05
  1049. Little SSM (RWKV7 7B) state checkpointing demo. -- 2025-09-04
  1050. Need advice on how to get VLLM working with 2xR9700 + 2x7900xtx? -- 2025-09-04
  1051. The Hacker's Guide to Building an AI Supercluster -- 2025-09-02
  1052. CAD, From Scratch: MakerCAD -- 2025-09-02
  1053. PSO-Merging: Merging Models Based on Particle Swarm Optimization -- 2025-09-02
  1054. internlm/Intern-S1-mini -- 2025-09-02
  1055. stepfun-ai/Step-Audio-2-mini -- 2025-09-02
  1056. Claude Memory Lazy Method: The Graduation Path (From 4 Prompts to 1) -- 2025-08-30
  1057. bullerwins/Wan2.2-I2V-A14B-GGUF -- 2025-08-28
  1058. QuantStack/Qwen-Image-Edit-GGUF -- 2025-08-28
  1059. dvlab-research/MGM-Omni -- 2025-08-28
  1060. DeepSeek V3.1 dynamic Unsloth GGUFs + chat template fixes -- 2025-08-28
  1061. PSA: OpenAI GPT-OSS running slow? Do not set top-k to 0! -- 2025-08-28
  1062. Seamlessly bridge LM Studio and OpenWebUI with zero configuration -- 2025-08-28
  1063. I built Husk, a native, private, and open-source iOS client for your local models -- 2025-08-28
  1064. SQLite-Vector adds support for float16 and bfloat16 (CPU, NEON, AVX2 and SSE2) -- 2025-08-26
  1065. zai-org/GLM-4.5 -- 2025-08-26
  1066. rednote-hilab/dots.vlm1.inst -- 2025-08-26
  1067. black-forest-labs/FLUX.1-Krea-dev -- 2025-08-26
  1068. Design Patterns in MCP: Literate Reasoning -- 2025-08-22
  1069. FlyMyAI/flymyai-lora-trainer -- 2025-08-22
  1070. Zedless: Zed fork focused on privacy and being local-first -- 2025-08-22
  1071. Show HN: Rucat – Cat for Prompt Engineers -- 2025-08-22
  1072. NEW VERSION: 0.6.23 Has Just Released! - Many fixes and new features, huge changelog -- 2025-08-22
  1073. nvidia/Llama-3_3-Nemotron-Super-49B-v1_5 -- 2025-08-21
  1074. NVIDIA Nemotron Nano 2 and the Nemotron Pretraining Dataset v1 -- 2025-08-20
  1075. deepseek-ai/DeepSeek-V3.1-Base · Hugging Face -- 2025-08-20
  1076. mistralai/Devstral-Small-2507 -- 2025-08-20
  1077. zai-org/GLM-4.5-Air -- 2025-08-20
  1078. microsoft/Phi-4-mini-flash-reasoning -- 2025-08-20
  1079. ModelTC/Qwen-Image-Lightning -- 2025-08-20
  1080. Fast Type-Aware Linting in Oxlint -- 2025-08-20
  1081. From Language to Logic: A Bi-Level Framework for Structured Reasoning -- 2025-08-19
  1082. MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE -- 2025-08-19
  1083. Distillation Scaling Laws -- 2025-08-19
  1084. GPT-5, where does it shine for you? -- 2025-08-19
  1085. vidore/colqwen-omni-v0.1 -- 2025-08-19
  1086. Tutorial: Open WebUI and llama-swap works great together! Demo of setup, model swapping and activity monitoring. -- 2025-08-18
  1087. Concurrency in open-weight/open-source models? -- 2025-08-18
  1088. [Editorial] Claude Flow, Alpha 90 release -- 2025-08-17
  1089. Optimizing Text gen webui (oobabooga) for MOE models (Qwen3-235b, GLM 4.5) -- 2025-08-17
  1090. Chen-zexi/vllm-cli -- 2025-08-17
  1091. zyfoxx/subhunter -- 2025-08-17
  1092. moonshotai/Kimi-K2-Instruct -- 2025-08-17
  1093. KittenML/kitten-tts-nano-0.1 -- 2025-08-17
  1094. ilkerzgi/Overlay-Kontext-Dev-LoRA -- 2025-08-17
  1095. JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 -- 2025-08-17
  1096. Fun with RTX PRO 6000 Blackwell SE -- 2025-08-17
  1097. Any tips/Advice for running gpt-oss-120b locally -- 2025-08-17
  1098. LLM performance of tiny (<4B) models? -- 2025-08-17
  1099. What "big" models can I run with this setup: 5070ti 16GB and 128GB ram, i9-13900k ? -- 2025-08-17
  1100. HuggingFaceTB/SmolLM3-3B-Base -- 2025-08-16
  1101. mistralai/Voxtral-Small-24B-2507 -- 2025-08-16
  1102. TiTan - a tiny model for tags and titles -- 2025-08-16
  1103. baidu/ERNIE-4.5-VL-424B-A47B-PT -- 2025-08-16
  1104. character-ai/pipelining-sft -- 2025-08-15
  1105. Compass-Thinker-7B Technical Report -- 2025-08-15
  1106. Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning -- 2025-08-15
  1107. [Editorial] GLM-4.5, enterprise use -- 2025-08-13
  1108. Best local model with function calling? -- 2025-08-13
  1109. agentica-org/DeepSWE-Preview -- 2025-08-13
  1110. janhq/Jan-v1-4B-GGUF -- 2025-08-13
  1111. Unsloth fixes chat_template (again). gpt-oss-120-high now scores 68.4 on Aider polyglot -- 2025-08-12
  1112. GPT-OSS-120B runs on just 8GB VRAM & 64GB+ system RAM -- 2025-08-12
  1113. How Attention Sinks Keep Language Models Stable -- 2025-08-12
  1114. Mitigate Hallucinations by Fine-tuning gpt-oss-120b with One Example -- 2025-08-10
  1115. uncensored gpt-oss-20b, bf16 and mxfp4 both available -- 2025-08-10
  1116. LGAI-EXAONE/EXAONE-4.0-1.2B -- 2025-08-10
  1117. New Open-Source Text-to-Image Model Just Dropped Qwen-Image (20B MMDiT) by Alibaba! -- 2025-08-10
  1118. Experience with GLM-4.5-Air + claude code? -- 2025-08-08
  1119. Just when you thought Qwen was done... -- 2025-08-08
  1120. Running GPT-OSS-120B at 500 tokens per second on Nvidia GPUs -- 2025-08-08
  1121. How the best image generation models work from the inside ? -- 2025-08-08
  1122. AIDC-AI/Ovis-U1-3B -- 2025-08-08
  1123. google/gemma-3n-E4B -- 2025-08-07
  1124. IntervitensInc/pangu-pro-moe-model -- 2025-08-07
  1125. Welcome GPT OSS, the new open-source model family from OpenAI! -- 2025-08-07
  1126. 0.82 um 105 W diode-pumped thulium-doped all silica fiber laser -- 2025-08-06
  1127. ByteDance drops Seed-Prover -- 2025-08-06
  1128. naver-hyperclovax/HyperCLOVAX-SEED-Think-14B -- 2025-08-06
  1129. [Editorial] a more mature phase of the AI cycle. -- 2025-08-05
  1130. glm-4.5-Air appreciation poist - if you have not done so already, give this model a try -- 2025-08-05
  1131. How to locally run Grok 4 with 2x AMD 7900 XTX GPUs? (24 GB VRAM x2) -- 2025-08-05
  1132. zai-org/GLM-4.5 -- 2025-08-05
  1133. unsloth/Qwen3-Coder-480B-A35B-Instruct-GGUF -- 2025-08-05
  1134. Learn Software-Defined Radio, GNURadio, RTL-SDR and PlutoSDR with Prof Jason -- 2025-08-05
  1135. Waiting on direct MCP integration—dev team, got a roadmap update? -- 2025-08-05
  1136. [Help] Figma MCP Tool Execution via HTTP API - Getting 404s, Is External Tool Calling Supported? -- 2025-08-05
  1137. Build an AI Shopping Assistant with Gradio MCP Servers -- 2025-08-01
  1138. Wan 2.2 T2V,I2V 14B MoE Models -- 2025-07-31
  1139. PowerInfer/SmallThinker-21BA3B-Instruct -- 2025-07-31
  1140. haykgrigo3/TimeCapsuleLLM -- 2025-07-30
  1141. unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF -- 2025-07-30
  1142. PhysicsWallahAI/Aryabhata-1.0 -- 2025-07-30
  1143. This year’s best open-source models and most cost-effective models -- 2025-07-29
  1144. I’m looking for multimodal image input support and uncensored LLM -- 2025-07-29
  1145. nvidia/audio-flamingo-3 -- 2025-07-29
  1146. mistralai/Voxtral-Mini-3B-2507 -- 2025-07-29
  1147. UI/UX benchmark update 7/22: Newest Qwen models added, Qwen3 takes the lead in terms of win rate (though still early) -- 2025-07-28
  1148. unsloth/Qwen3-235B-A22B-Thinking-2507-GGUF -- 2025-07-28
  1149. zai-org/GLM-4.5 -- 2025-07-28
  1150. Tesslate/UIGEN-X-32B-0727 -- 2025-07-28
  1151. had to fine-tune qwen since llama sucks at summarizing -- 2025-07-28
  1152. orchestre-dev/ccproxy -- 2025-07-28
  1153. Guide to PDF security -- 2025-07-28
  1154. MetaMask extension bug causes 100s of GBs of extraneous data to be written -- 2025-07-28
  1155. Commodore 64 on New FPGA -- 2025-07-28
  1156. Running Qwen3 235B-A22B 2507 on a Threadripper 3970X + 3x RTX 3090 Machine at 15 tok/s -- 2025-07-25
  1157. The Latest GPT-5 Leaks and Teasers -- 2025-07-25
  1158. Qwen3-235B-A22B-Thinking-2507 released! -- 2025-07-25
  1159. albozes/shotbuddy -- 2025-07-25
  1160. uttam-li/dfs -- 2025-07-25
  1161. OmniSVG/OmniSVG -- 2025-07-25
  1162. Freezer Monitoring: Because Ice Cream Is a Dish Best Served Cold -- 2025-07-25
  1163. Fast LoRA inference for Flux with Diffusers and PEFT -- 2025-07-25
  1164. FreeBSD 15's installer to gain option to install a full KDE Plasma desktop -- 2025-07-24
  1165. Spanish police arrest five over $542M crypto investment scheme -- 2025-07-24
  1166. A Spectrophotometer Jailbreak to Resolve Colorful Disputes -- 2025-07-24
  1167. Lucy: A Mobile-Capable 1.7B Reasoning Model That Rivals Jan-Nano -- 2025-07-23
  1168. Recommend hardware for my use case? -- 2025-07-20
  1169. Best Hardware Setup to Run DeepSeek-V3 670B Locally on $40K–$80K? -- 2025-07-20
  1170. e6a5/flow -- 2025-07-20
  1171. Improve Your KiCad Productivity With These Considered Shortcut Keys -- 2025-07-20
  1172. Semantic chunking using LLMs -- 2025-07-20
  1173. Does the OpenWebUi run the sentence transformer models locally? -- 2025-07-20
  1174. Dataset for structured (JSON) output? -- 2025-07-19
  1175. support for Kimi-K2 has been merged into llama.cpp -- 2025-07-19
  1176. t-tech/T-pro-it-2.0 -- 2025-07-19
  1177. Support for diffusion models (Dream 7B) has been merged into llama.cpp -- 2025-07-17
  1178. Seq vs Seq: the Ettin Suite of Paired Encoders and Decoders -- 2025-07-17
  1179. T5Gemma: A new collection of encoder-decoder Gemma models- Google Developers Blog -- 2025-07-17
  1180. H-Net: a hierarchical network that replaces tokenization with a dynamic chunking process directly inside the model, automatically discovering and operating over meaningful units of data -- 2025-07-17
  1181. An Open-Concept 3D Printer Using Cantilever Arms -- 2025-07-17
  1182. Exploring State-Space-Model based Language Model in Music Generation -- 2025-07-16
  1183. Diffusion model support in llama.cpp. -- 2025-07-16
  1184. GLM-4 MoE incoming -- 2025-07-16
  1185. FlexOlmo: Open Language Models for Flexible Data Use | Implications for federated training in the open source community -- 2025-07-15
  1186. Tencent/AICGSecEval -- 2025-07-15
  1187. microsoft/NextCoder-32B -- 2025-07-15
  1188. RekaAI/reka-flash-3.1 -- 2025-07-15
  1189. Qwen3-235B-A22B @ 0.7t/s. Hardware or configuration bottleneck? -- 2025-07-15
  1190. Building a silent, budget 4-GPU LLM workstation—1×3090 + 3×P40, need advice -- 2025-07-15
  1191. Enough resources for light AI workloads? -- 2025-07-15
  1192. What can I expect from current amd igpu performance? -- 2025-07-15
  1193. Kimi-K2 is a DeepSeek V3 with more experts -- 2025-07-14
  1194. HuggingFaceTB/SmolLM3-3B-Base -- 2025-07-14
  1195. Replication of Quantum Factorisation Records with an 8-bit Home Computer [pdf] -- 2025-07-14
  1196. Why don’t we have a big torrent repo for open-source LLMs? -- 2025-07-12
  1197. Local PDF Database searchable with ollama - best setup? -- 2025-07-12
  1198. Tinyllama on old Mediatek G80 android device -- 2025-07-12
  1199. I used Ollama to build a Cursor for PDFs -- 2025-07-12
  1200. Advice on switching to LLM -- 2025-07-12
  1201. Building the Hugging Face MCP Server -- 2025-07-11
  1202. Support for the upcoming IBM Granite 4.0 has been merged into llama.cpp -- 2025-07-11
  1203. support for Falcon-H1 model family has been merged into llama.cpp -- 2025-07-11
  1204. [Tool Release] Finetune & Quantize 1–3B LLMs on 8GB RAM using LoFT CLI (TinyLlama + QLoRA + llama.cpp) -- 2025-07-11
  1205. Qwen3-8B-BitNet -- 2025-07-11
  1206. How are commercial dense models so much faster? -- 2025-07-09
  1207. pola-rs/polars -- 2025-07-09
  1208. Looking for an upgrade from Meta-Llama-3.1-8B-Instruct-Q4_K_L.gguf, especially for letter parsing. Last time I looked into this was a very long time ago (7 months!) What are the best models nowadays? -- 2025-07-08
  1209. Best models by size? -- 2025-07-08
  1210. Planning a 7–8B Model Benchmark on 8GB GPU — What Should I Test & Measure? -- 2025-07-08
  1211. Smallest & best OCR model that can read math & code? -- 2025-07-07
  1212. Qwen/WorldPM-72B -- 2025-07-07
  1213. black-forest-labs/FLUX.1-Kontext-dev-onnx -- 2025-07-07
  1214. MCPVerse – An open playground for autonomous agents to publicly chat, react, publish, and exhibit emergent behavior -- 2025-07-06
  1215. I made a free iOS app for people who run LLMs locally. It’s a chatbot that you can use away from home to interact with an LLM that runs locally on your desktop Mac. -- 2025-07-06
  1216. Using local models with Void -- 2025-07-05
  1217. Intel GPU vLLM Docker Compose Bootstrap with Phi-lthy4 on A770 -- 2025-07-05
  1218. Accelerated LLM Inference on AMD Instinct™ GPUs with vLLM 0.9.x and ROCm -- 2025-07-05
  1219. Gemma 3n fully available in the open-source ecosystem! -- 2025-07-05
  1220. 5060ti 16gb or 9060xt 16gb for small llm server -- 2025-07-05
  1221. Qwen3 models in MLX format! -- 2025-07-05
  1222. Best local coding model right now? -- 2025-07-05
  1223. Wallbleed: A Memory Disclosure Vulnerability in the Great Firewall of China -- 2025-07-05
  1224. 0-Pierced Triangles within a Poisson Overlay -- 2025-07-05
  1225. 1000 days of lowest frequency emission from the low-luminosity GRB 171205A -- 2025-07-05
  1226. Found a Web3 LLM That Actually Gets DeFi Right -- 2025-07-03
  1227. apple/DiffuCoder-7B-cpGRPO -- 2025-07-03
  1228. Ollama - Windows 11 > LXC Docker - Openwebui = constant BSOD with RTX 5090 Ventus on driver 576.80 -- 2025-07-02
  1229. Running Open WebUI with NVIDIA GPU Support? -- 2025-07-02
  1230. Cursor 1.0 -- 2025-06-30
  1231. Help me design a robust on-prem Llama 3 70B infrastructure for 30 users – Complete hardware/software list wanted -- 2025-06-30
  1232. Jan-nano, a 4B model that can outperform 671B on MCP -- 2025-06-30
  1233. Models that are good and fast at Long Document Processing -- 2025-06-30
  1234. I am making an AI batteries included Web Framework (like Django but for AI) -- 2025-06-30
  1235. [New Features & Better] Tabulens: A Vision-LLM Powered PDF Table Extractor -- 2025-06-30
  1236. I tested 10 LLMs locally on my MacBook Air M1 (8GB RAM!) – Here's what actually works- -- 2025-06-30
  1237. Chatbot without ChatGPT -- 2025-06-30
  1238. What's the best way to save and manage different text files for the models to reference? PRD, cursor rules, tech stack, design reference, etc? -- 2025-06-30
  1239. Bzip2 crate switches from C to 100% Rust -- 2025-06-30
  1240. Litestream: Revamped -- 2025-06-30
  1241. Stop using REST for state synchronization (2024) -- 2025-06-30
  1242. Announcing `mcp-protocol-sdk`: A New Enterprise grade Rust SDK for AI Tool Calling (Model Context Protocol) -- 2025-06-30
  1243. THU-KEG/LongWriter-Zero-32B -- 2025-06-30
  1244. 100 Gbps Indoor Access and 4.8 Gbps Outdoor Point-to-Point LiFi Transmission Systems using Laser-based Light Sources -- 2025-06-30
  1245. (0,4) brane box models -- 2025-06-30
  1246. Built memX: a shared memory backend for LLM agents (demo + open-source code) -- 2025-06-29
  1247. Automatically Evaluating AI Coding Assistants with Each Git Commit (Open Source) -- 2025-06-29
  1248. Secure Minions: private collaboration between Ollama and frontier models -- 2025-06-29
  1249. Privacy implications of sending data to OpenRouter -- 2025-06-29
  1250. Exploring Practical Uses for Small Language Models (e.g., Microsoft Phi) -- 2025-06-29
  1251. LLM with OCR capabilities -- 2025-06-29
  1252. How to create a speech recognition model from scratch -- 2025-06-29
  1253. Arch 0.3.0 is out - I added support for the Claude family of LLMs in the proxy server framework for agents 🚀 -- 2025-06-29
  1254. Gemini Cli MCP Agent just released ! -- 2025-06-29
  1255. Freeplane xml mind maps locally: only Qwen3 and Phi4 Reasoning Plus can create them in one shot? -- 2025-06-29
  1256. Reinforcement Pre-Training -- 2025-06-29
  1257. unsloth/gemma-3n-E4B-it-GGUF -- 2025-06-29
  1258. chandar-lab/NeoBERT -- 2025-06-29
  1259. unsloth/Mistral-Small-3.2-24B-Instruct-2506-GGUF -- 2025-06-26
  1260. jinaai/jina-embeddings-v4 -- 2025-06-26
  1261. Intelligent-Internet/II-Medical-8B-1706 -- 2025-06-26
  1262. 0-concordance of knotted surfaces and Alexander ideals -- 2025-06-26
  1263. 100% of the zeros of the Riemann zeta-function are on the critical line -- 2025-06-25
  1264. 100% of odd hyperelliptic Jacobians have no rational points of small height -- 2025-06-25
  1265. deepseek-ai/DualPipe -- 2025-06-23
  1266. 100 Particles Quantum Heat Engine: Exploring the Impact of Criticality on Efficiency -- 2025-06-23
  1267. 0-Auslander correspondence -- 2025-06-23
  1268. nvidia/Cosmos-Predict2-2B-Text2Image -- 2025-06-22
  1269. 1000-10,000 M$_\odot$ Primordial Stars Created the Nitrogen Excess in the Galaxy GS 3073 at $z = 5.55$ -- 2025-06-21
  1270. $0^+$ to $2^+$ neutrinoless double-$β$ decay of $^{76}$Ge, $^{82}$Se, $^{130}$Te and $^{136}$Xe in the microscopic interacting boson model} -- 2025-06-21
  1271. resemble-ai/chatterbox -- 2025-06-21
  1272. google/magenta-realtime -- 2025-06-21
  1273. 0-1 laws for pattern occurrences in phylogenetic trees and networks -- 2025-06-20
  1274. meta-llama/Llama-3.1-8B-Instruct -- 2025-06-19
  1275. MiniMaxAI/MiniMax-M1-80k -- 2025-06-19
  1276. Demo Video of AutoBE, Backend Vibe Coding Agent Achieving 100% Compilation Success (Open Source) -- 2025-06-19
  1277. How to set up local llms on a 6700 xt -- 2025-06-19
  1278. Jetson Orin AGX 32gb -- 2025-06-19
  1279. AMD GPU support -- 2025-06-19
  1280. Much lower performance for Mistral-Small 24B on RTX 3090 and from deepinfra API -- 2025-06-19
  1281. Extract Website Information -- 2025-06-19
  1282. Looking for a verified copy of big-lama.ckpt (181MB) used in the original LaMa inpainting model trained on Places2. -- 2025-06-19
  1283. Is it true that all tools like Cline/Copilot Agent/Roo Code/Windsurf/Claude Code/Cursor are roughly the same thing? -- 2025-06-19
  1284. SkyRoof: New Ham Satellite Tracking and SDR Receiver Software -- 2025-06-19
  1285. MiniMaxAI/MiniMax-M1-40k -- 2025-06-18
  1286. Qwen/Qwen3-Reranker-4B -- 2025-06-18
  1287. ckanthony/openapi-mcp -- 2025-06-18
  1288. Ta0ing/MCP-SecurityTools -- 2025-06-18
  1289. 100% Elimination of Hallucinations on RAGTruth for GPT-4 and GPT-3.5 Turbo -- 2025-06-17
  1290. unsloth/Magistral-Small-2506-GGUF -- 2025-06-12
  1291. mistralai/Magistral-Small-2506 -- 2025-06-12
  1292. rednote-hilab/dots.llm1.base -- 2025-06-12
  1293. [Tool] rvn-convert: OSS Rust-based SafeTensors to GGUF v3 converter (single-shard, fast, no Python) -- 2025-06-12
  1294. GuidedQuant: Boost LLM layer-wise PTQ methods using the end loss guidance (Qwen3, Gemma3, Llama3.3 / 2~4bit Quantization) -- 2025-06-12
  1295. I built a memory MCP that understands you (so Sam Altman can't). -- 2025-06-12
  1296. Built an open source desktop app to easily play with local LLMs and MCP -- 2025-06-12
  1297. mtmd : support Qwen 2.5 Omni (input audio+vision, no audio output) by ngxson · Pull Request #13784 · ggml-org/llama.cpp -- 2025-06-12
  1298. i got tired of the errors, so automated debugging using Ollama -- 2025-06-12
  1299. Ablating Gemma 3 27B variants with synthetic data from Sonnet 4 (Few-shot vs LoRA) -- 2025-06-12
  1300. The LLM Gateway gets a major upgrade: becomes a data-plane for Agents. -- 2025-06-12
  1301. Introducing stronger dependencies on systemd -- 2025-06-12
  1302. How we decreased GitLab repo backup times from 48 hours to 41 minutes -- 2025-06-12
  1303. The Quest for 100k - LLAMA.CPP Setting for a Noobie -- 2025-06-12
  1304. News publishers call Google's AI Mode 'theft' -- 2025-06-11
  1305. Qwen/Qwen3-Embedding-4B -- 2025-06-10
  1306. The Unreliability of LLMs and What Lies Ahead -- 2025-06-10
  1307. AnythingLLM RAG with Gemma 3:12b & BGE-m3-F16: LM Studio vs. Ollama Embedding Discrepancies - Same GGUF, Different Results? -- 2025-06-09
  1308. What Models for C/C++? -- 2025-06-09
  1309. Help with guardrails ai and local ollama model -- 2025-06-09
  1310. Setup Recommendation for University (H200 vs RTX 6000 Pro) -- 2025-06-09
  1311. LlamaFirewall: framework open source per rilevare e mitigare i rischi per la sicurezza incentrati sull'intelligenza artificiale - Help Net Security -- 2025-06-09
  1312. Are autoencoders really need for anomaly detection in time series? -- 2025-06-09
  1313. Backdoored malware repos traced to single GitHub user -- 2025-06-09
  1314. Improper Access Control Allows All Users to View Private Content. Am I doing it wrong ? -- 2025-06-09
  1315. Agno Now Supports Dual Model Output (Reasoning + Structure) -- 2025-06-09
  1316. 100 Gbps Quantum-safe IPsec VPN Tunnels over 46 km Deployed Fiber -- 2025-06-09
  1317. Qwen/Qwen3-Embedding-0.6B -- 2025-06-08
  1318. Qwen/Qwen3-Embedding-8B -- 2025-06-06
  1319. nvidia/Llama-3.1-Nemotron-Nano-VL-8B-V1 -- 2025-06-05
  1320. hexgrad/Kokoro-82M -- 2025-06-05
  1321. Qwen/Qwen3-Embedding-0.6B-GGUF -- 2025-06-05
  1322. Gradient-Based Program Repair: Fixing Bugs in Continuous Program Spaces -- 2025-06-05
  1323. 0-dimensional Homology Preserving Dimensionality Reduction with TopoMap -- 2025-06-05
  1324. NousResearch/atropos -- 2025-06-04
  1325. 100,000 Podcasts: A Spoken English Document Corpus -- 2025-06-03
  1326. nvidia/Nemotron-Research-Reasoning-Qwen-1.5B -- 2025-06-03
  1327. Intelligent-Internet/II-Medical-8B -- 2025-06-03
  1328. Gen-Verse/MMaDA -- 2025-06-01
  1329. osmosis-ai/Osmosis-Structure-0.6B -- 2025-06-01
  1330. simplescaling/s1 -- 2025-05-31
  1331. FractalAIResearch/Fathom-R1-14B -- 2025-05-31
  1332. unsloth/DeepSeek-R1-0528-Qwen3-8B-GGUF -- 2025-05-31
  1333. deepseek-ai/DeepSeek-R1-0528-Qwen3-8B -- 2025-05-31
  1334. I made Model Version Control Protocol for AI agents -- 2025-05-31
  1335. AI Baby Monitor – fully local Video-LLM nanny (beeps when safety rules are violated) -- 2025-05-31
  1336. LMStudio - llama.cpp - vLLM -- 2025-05-31
  1337. Built an ADK Agent that finds Jobs based on your Resume -- 2025-05-31
  1338. Should I resize the image before sending it to Qwen VL 7B? Would it give better results? -- 2025-05-31
  1339. How to start a LLM project? -- 2025-05-31
  1340. Beware of Fast-Math -- 2025-05-31
  1341. facebook/OMol25 -- 2025-05-30
  1342. EdinburghNLP/MMLongBench -- 2025-05-29
  1343. Show HN: Model2vec-Rs – Fast Static Text Embeddings in Rust -- 2025-05-29
  1344. unsloth/DeepSeek-R1-0528-GGUF -- 2025-05-29
  1345. QuantStack/Wan2.1-VACE-14B-GGUF -- 2025-05-29
  1346. nvidia/Llama-3.1-Nemotron-Nano-4B-v1.1 -- 2025-05-28
  1347. Tongyi-Zhiwen/QwenLong-L1-32B -- 2025-05-28
  1348. PKU-DS-LAB/FairyR1-32B -- 2025-05-28
  1349. google/medgemma-4b-pt -- 2025-05-28