Open-Weight Models

Open-source model releases, community models, licensing

1135 articles across 248 editions

Articles

  1. [Editorial] Researchers Complain That OpenAI Revoked Their Access to Limited Cyber Program -- 2026-08-24
  2. [Editorial] Smolbox — Live Site -- 2026-08-24
  3. [Editorial] Smolbox — remyhax Writeup -- 2026-08-24
  4. [Editorial] gadievron/raptor — Issue #889 -- 2026-08-24
  5. Every Model Cheats -- 2026-08-24
  6. [Editorial] DreamLab-AI Loom — Research Paper v4 -- 2026-08-24
  7. [Editorial] Editor-Curated Video Pick #2 -- 2026-08-24
  8. [Editorial] Editor-Curated Video Pick #1 -- 2026-08-24
  9. [Editorial] sw30labs/singularity-atlas -- 2026-08-24
  10. [Editorial] Ox Alpha: Stealth AI Model With 1M-Token Context -- 2026-08-24
  11. [Editorial] Coding Model Ox Alpha Retains Every Prompt — And You Can't Name the Company Holding Them -- 2026-08-24
  12. Ox Alpha stealth model: GLM5 Air, Mimo V3 or ? -- 2026-08-24
  13. Artificial Analysis' Qwen3.8-27B benchmarks put it neck and neck with DeepSeek V4 and GPT-5.6 Luna Max -- 2026-08-21
  14. [Editorial] YouTube feature -- 2026-08-21
  15. DeepSeek-V4-Flash-Vision-Exp -- 2026-08-21
  16. Microsoft ran a 100B BitNet test on one CPU. Here is the detail the headline misses. -- 2026-08-21
  17. The August 17 outage -- 2026-08-21
  18. [Editorial] RepoRadar -- 2026-08-21
  19. [Editorial] Event-Horizon (ruvnet) -- 2026-08-21
  20. GIMP Development Update -- 2026-08-21
  21. [Editorial] music.cognitum.one -- 2026-08-21
  22. It's actually crazy how good DSv4 Flash 0731 is -- 2026-08-19
  23. I ran DeepSeek V4 Flash 284B + DSpark on one RTX PRO 6000. The drafter was faster in RAM than VRAM. -- 2026-08-19
  24. Deepseek Harnness - why is feels better -- 2026-08-19
  25. [Editorial] pawaca/dsh-edge -- 2026-08-19
  26. Stolen LLM Reasoning: How come OpenAI, Anthrophic, Google have the same vulnerabilities? -- 2026-08-19
  27. AI-Generated GitHub Copilot “Autofix” Allowed Compromise of Snowflake's Jira -- 2026-08-19
  28. Israel creates fake think tank in likely attempt to dupe AI chatbots -- 2026-08-19
  29. The AI Credit Resale Economy -- 2026-08-19
  30. Simon Willison: Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things -- 2026-08-19
  31. Ollama's MTP variant of Qwen3.8-27B is 2x slower than the non-MTP one, measured, with a negative control -- 2026-08-19
  32. Taking Qwen3.5-9B quants to SOTA. New lineup incoming :) -- 2026-08-19
  33. [Paper] Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning -- 2026-08-19
  34. dealignai/Gemma-4-31B-JANG_4M-CRACK -- 2026-08-19
  35. bitsandbytes creator teasing new quantization method: GLM 5.3 on a single DGX Spark at 7t/s -- 2026-08-18
  36. GLM 5.3 weights. It might offer the best capacity-to-size ratio. -- 2026-08-18
  37. Ling 3.0 support merged into llama.cpp -- 2026-08-18
  38. Why not? ☺️ -- 2026-08-18
  39. A Preview of DuckDB v2.0 -- 2026-08-18
  40. Reticulum – Decentralized Mesh Network -- 2026-08-18
  41. Design 3D-printable parts by talking -- 2026-08-18
  42. Store Tunes on Paper and Stream Them Over LoRA -- 2026-08-18
  43. RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM -- 2026-08-18
  44. I built an open source local memory engine (Hillock v0.2) to ingest docs in sub-seconds alongside Ollama -- 2026-08-18
  45. Open Web UI my usage (local LLM for a compagny) -- 2026-08-18
  46. [Editorial] GLM-5.3 Release (z.ai) -- 2026-08-14
  47. [Editorial] LinkedIn Post (gMy2vbnc) -- 2026-08-14
  48. State of Open Models: Summer 2026 Observations -- 2026-08-14
  49. Qwen 3.8 27B is out : open weights, best local dense model yet -- 2026-08-14
  50. CohereLabs/North-Micro-Vision-Instruct · Hugging Face -- 2026-08-14
  51. U.S. Department of Energy Launches the Genesis Open Models Initiative and, with Arcee, Unveils Genesis-Science-1 — Its First Open-Weight Model for Scientific Research -- 2026-08-14
  52. We quantized DeepSeek V4 0731 and benchmarked it against popular quants on 8× RTX 5090 -- 2026-08-14
  53. Gemma 4 12B Q3: +8.55% Coding Performance From Tensor-Level Quantization Allocation -- 2026-08-14
  54. I trained a 1B-parameter LLM from scratch on 20B tokens for about $200 -- 2026-08-14
  55. Showoff Saturday: Local 4x 6000 Pro (multi-year progression) -- 2026-08-14
  56. [Editorial] Foxflow Fleet (probagi.com) -- 2026-08-14
  57. Mojo 1.0 -- 2026-08-13
  58. DeepSeek Harness -- 2026-08-13
  59. [Editorial] Unsloth Desktop -- 2026-08-13
  60. I built a local realtime voice stack for Ollama: Parakeet STT → Qwen 2.5 7B → Qwen3-TTS -- 2026-08-13
  61. I put Gemma 4 E4B and E2B into an e-reader so I can ask my weird questions and share my thoughts in private directly in app. -- 2026-08-13
  62. I wanted to understand Transformers below the PyTorch abstraction layer, so I built one from scratch in CuPy -- 2026-08-13
  63. [Editorial] DeepSWE by Datacurve -- 2026-08-13
  64. Anthropic: Introducing The Conceptual Reasoning Index -- 2026-08-13
  65. pathwaycom/arc-task-gen -- 2026-08-13
  66. [Editorial] Infisical Agent Vault -- 2026-08-13
  67. [Editorial] jedarden/seam -- 2026-08-13
  68. [Editorial] LinkedIn feature -- 2026-08-13
  69. OpenAI on upcoming model "Astra" (GPT-6): "We're treating it as our first "critical" model for cybersecurity" -- 2026-08-12
  70. Everything you do is being recorded -- 2026-08-12
  71. MiniMax-AI/MiniMax-H3 -- 2026-08-11
  72. Glimmer seems pretty censored? -- 2026-08-11
  73. No wonder Qwen and Gemma are so different -- 2026-08-11
  74. CJK Manga/Manhwa/Manhua 150M OCR model (hayai-ocr-v2) outperforming PaddleOCR-VL-For-Manga -- 2026-08-11
  75. [Editorial] Anthropic's Invisible Watermarks in Claude Text -- 2026-08-11
  76. [Editorial] Reuven Cohen: Quantum Mechanics May Be Telling Us Something -- 2026-08-11
  77. Underestimated budget solution: radeon 780m iGPU -- 2026-08-10
  78. Building a zero-dependency C inference engine for BitNet (1.58-bit) - lessons from hitting 36 tok/s on a Xeon CPU -- 2026-08-10
  79. I got tired of my 300GB model loads taking 5min on RPC. PR 26291 speeds it 300% to 1min30sec (4060ti+ddr4) + (4060ti+ddr5) -- 2026-08-10
  80. Early signs that Muse-Glimmer-30B might quantize *very* well? Share your experiences. -- 2026-08-10
  81. Best Local LLMs - August 2026 -- 2026-08-10
  82. Qwen3-TTS voice cloning is now in mainline llama.cpp — the old demo finally became real support -- 2026-08-10
  83. [Editorial] -- 2026-08-10
  84. [Editorial] -- 2026-08-10
  85. [Editorial] -- 2026-08-10
  86. DeepSeek V4 Flash 2-bit quant achieves 100% on SQL benchmark locally -- 2026-08-07
  87. Scotoma-2: Gemma4, but with less annoying slop and better writing -- 2026-08-07
  88. Mach-1 Additive: 95% of Qwen 3.6 35B performance while 10x smaller -- 2026-08-07
  89. Intern S2 Mobius — Qwen3.5-35B derivative with architectural throughput gains -- 2026-08-07
  90. A 2.6B model with tool calling and 128K context now runs at 30 tok/s on a phone -- 2026-08-06
  91. inclusionAI/Ling-3.0-flash weights are up on Hugging Face — MIT, BF16 plus an official FP8 -- 2026-08-06
  92. They almost catched up on Frontier performance, so now catching up on prices -- 2026-08-06
  93. Deepseek v4 flash 0731 still not holding up. -- 2026-08-06
  94. Maple-Preview — Ternary 20B MoE running at 120 tok/s on an iPhone -- 2026-08-05
  95. Qwen Developers' AMA: 3.8-27B coming soon, 2.4T params (95B active) -- 2026-08-05
  96. AI9Stars released G9v3-39A5B — Apache 2.0, 39B/5B active MoE -- 2026-08-05
  97. [Editorial] Unsloth Kimi K3 Model Documentation -- 2026-08-05
  98. [Editorial] Series of Model Tests and Results -- 2026-08-05
  99. Smaller, faster, safer: running Kimi and GLM at scale on Cloudflare -- 2026-08-04
  100. Huawei open-sources openPangu-2.0-Pro: 505B-A18B MoE on Ascend -- 2026-08-04
  101. NousResearch ships Hermes 0.20 agent framework -- 2026-08-04
  102. DMARC Has Been Public Since 2012. 68.4% of Domains Still Don't Enforce It -- 2026-07-31
  103. Kimi K3 Architecture Overview and Notes -- 2026-07-31
  104. DeepSeek-V4-Flash-0731 on HuggingFace -- 2026-07-31
  105. [Editorial] Poolside Laguna-S-2.1 — Open Coding Model -- 2026-07-31
  106. LFM2.5-Encoders: Fast at Long Context, Even on CPU -- 2026-07-31
  107. Dario Amodei: closed-weights models are worse than open-weights ones? -- 2026-07-31
  108. Tested proven orchestration techniques on small local models — 90% failed, the 10% that survived roughly doubled task completion -- 2026-07-31
  109. LiteRT-LM is up to 3.5x faster than llama.cpp on Intel Arc iGPU -- 2026-07-31
  110. First Kimi K3 results on home lab — ~4 t/s -- 2026-07-31
  111. SK Hynix stock fell 40% in 30 days — cheap RAM and GPUs again? -- 2026-07-31
  112. [Editorial] Video Content -- 2026-07-31
  113. SciCodePile: A 128GB Corpus and Executable Benchmark for Scientific Code Generation -- 2026-07-31
  114. Procedural desert explorer built with Claude Code (Opus 5) and Three.js -- 2026-07-31
  115. img2threejs — Image-to-3D procedural Three.js model generation -- 2026-07-31
  116. Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac -- 2026-07-30
  117. STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU -- 2026-07-30
  118. [Editorial] -- 2026-07-30
  119. [Editorial] -- 2026-07-30
  120. [Editorial] -- 2026-07-30
  121. Was waiting for Kimi 3 and now Ollama release it and has to pay extra to use it (like OpenRoute) -- 2026-07-30
  122. Everyone posts day-one impressions. What's still in your stack a month later? -- 2026-07-30
  123. [Editorial] -- 2026-07-30
  124. Kimi-K3 Releases on HuggingFace 6/27 -- 2026-07-28
  125. Ling-3.0-flash weights: SGLang says day-0, vLLM says when they land, llama.cpp closed the 2.6 request as not_planned -- 2026-07-28
  126. Sanctions on Open Source. hope they don't do anything stupid here. -- 2026-07-28
  127. [Editorial] -- 2026-07-28
  128. RTX 2080 Ti Memory Upgrade to 22 GB -- 2026-07-28
  129. microsoft/VibeVoice-ASR-BitNet -- 2026-07-28
  130. tetsuo-ai/voice_clone_lab -- 2026-07-28
  131. [Editorial] -- 2026-07-27
  132. [Editorial] -- 2026-07-27
  133. OpenAI and Anthropic unite against open-weight AI risks to their bottom line -- 2026-07-27
  134. [Editorial] -- 2026-07-27
  135. [Editorial] -- 2026-07-27
  136. The Open Weight Unbundling -- 2026-07-24
  137. Model "distillation" accusations are getting way overblown at this point -- 2026-07-24
  138. The Distillation Claims Are Fake and Desperate -- 2026-07-24
  139. I gave Kimi K3 a shot at auditing my post-quantum crypto project, it found 5 real bugs Fable/Opus 4.8 and GPT-5.6 Sol had all missed -- 2026-07-24
  140. China's Kimi K3 fuels fears safety curbs are holding back US AI -- 2026-07-23
  141. Startup founders urge Trump not to shut off Chinese open weight AI -- 2026-07-23
  142. Ahem! Qwen is on the move again -- 2026-07-23
  143. [Editorial] -- 2026-07-23
  144. tried predicting which MoE experts get used next token to speed up cpu/gpu offload, got some real numbers, is this actually implementable or am i wasting my time (30tg/s -> 150-200tg/s) -- 2026-07-23
  145. DS V4 on single b300. only 770 tok/s batched in vLLM -- 2026-07-23
  146. New Framework Desktop Option with AMD Ryzen AI Max+ Pro 495 and 192GB Memory -- 2026-07-23
  147. [Editorial] -- 2026-07-23
  148. [Editorial] -- 2026-07-23
  149. Are AI labs pelicanmaxxing? -- 2026-07-23
  150. China's open-weights AI strategy is winning -- 2026-07-22
  151. How do we benefits from 2+ T models? -- 2026-07-22
  152. Serving a fleet of Qwen3.5 122b sessions on a single Mac Studio (96GB) without losing your sanity -- 2026-07-22
  153. Intel Starts Shipping High-NA EUV Silicon -- 2026-07-22
  154. The State of Simulation for Physical AI: An Overview -- 2026-07-22
  155. [Editorial] -- 2026-07-22
  156. OpenAI had to pause an unreleased model after it escaped containment. -- 2026-07-21
  157. [Editorial] -- 2026-07-21
  158. Quant Convergence: Bridging Classical Value Investing and Modern Factor Models for Systematic Equity Selection -- 2026-07-21
  159. [Editorial] -- 2026-07-21
  160. [Editorial] -- 2026-07-21
  161. Model Routing Is Simple. Until It Isn't. — IBM Research -- 2026-07-20
  162. [Editorial] OmniRoute — Model Routing Framework -- 2026-07-20
  163. I Burned All My Tokens Researching How to Save Tokens -- 2026-07-20
  164. [Editorial] GCF-Rust — Blackwell Systems GPU Compute Framework -- 2026-07-20
  165. Kimi K3 Agentic Benchmark -- 2026-07-17
  166. Kimi k3 is 2.8t! Will need to have an aggressive iQ2_XXS or IQ1.8! -- 2026-07-17
  167. Governments, companies, nonprofits should invest in free, open source AI [pdf] -- 2026-07-17
  168. Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon with no GPU -- 2026-07-17
  169. LM Studio Bionic: the AI agent for open models -- 2026-07-17
  170. Prism-ML Bonsai Qwen 3.6 27B -- 2026-07-17
  171. i tried ternary decomposition instead of quantization. it works as good at q4km but takes slightly more vram. while being completly ternary. and completly PTQ (no QAT) -- 2026-07-17
  172. Prayers requested for an Abomination(Modified Gemma) -- 2026-07-17
  173. Hermes on Android (Graphene OS) -- 2026-07-17
  174. Inkling: Our Open-Weights Model -- 2026-07-16
  175. [Editorial] -- 2026-07-16
  176. hustvl/Moebius -- 2026-07-16
  177. Alternative(s) to run CUDA on non-Nvidia hardware -- 2026-07-16
  178. High-Bandwidth Flash offers efficient storage for model weights -- 2026-07-16
  179. Colibri streaming for Hy3 (Run Hy3 on 10GB (V)RAM) -- 2026-07-16
  180. In some languages, Claude will be more strict. Anthropic found out how language changes AI responses. -- 2026-07-15
  181. Building Food Metadata with LLM Juries -- 2026-07-15
  182. [Editorial] -- 2026-07-15
  183. Self-hosted voice for any agent/harness of your choice (open-source) -- 2026-07-15
  184. The untuned 27B beat the tuned 75B as an agent -- 2026-07-15
  185. Qwen3.6 35B-A3B (Q8_0, no KV quant) single prompt in opencode -- 2026-07-15
  186. I didn't give up - extGemma4-40_5B returned -- 2026-07-14
  187. J-Space Hallucination Signal Stress-Tested Across 7 Datasets on Qwen3-4B -- 2026-07-14
  188. FT: Companies Turn to Chinese Open Weight Models to Cut Costs -- 2026-07-14
  189. PrismML Compresses Qwen-3.6-27B to Under 4GB — Runs on iPhone 17 Pro -- 2026-07-14
  190. [Editorial] ThinkingCap Qwen3.6-27B -- 2026-07-14
  191. GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine -- 2026-07-14
  192. Muse Spark 1.1 -- 2026-07-10
  193. If the GPT-5.6 SOL rumors are true, is anyone actually sticking with Fable 5? -- 2026-07-10
  194. KilimcininKorOglu/M365Bridge -- 2026-07-10
  195. nvidia/Nemotron-Labs-Audex-30B-A3B · Hugging Face -- 2026-07-08
  196. Gemma 4 Technical Report -- 2026-07-08
  197. Hugging Face and Cerebras bring Gemma 4 to real-time voice AI -- 2026-07-08
  198. Qwen's J-Space — Anthropic's Discovery of an Internal Model Global Workspace -- 2026-07-07
  199. Physics-Style Attractor System Replaces Neural Networks for Word Embeddings — Hits SimLex-999 ρ=0.36 on 7.5% of Wikipedia -- 2026-07-07
  200. Why Specialization Is Inevitable -- 2026-07-07
  201. Follow-up: GLM-5.2 NVFP4 on four DGX Sparks — the MTP mystery is solved, and it's now ~24 tok/s at 128K context -- 2026-07-06
  202. Follow-up: DeepSeek V4 Flash on 2x RTX PRO 6000 finishes real coding tasks faster than Sonnet and Opus, at about Sonnet quality -- 2026-07-06
  203. Best Local VLMs - July 2026 -- 2026-07-06
  204. I benchmarked PrismML's 1-bit Bonsai-8B against IBM's Granite on CPU tool calling. The 1-bit model won, but only with grammar-constrained decoding -- 2026-07-06
  205. NASA testing local LLM inference for future space missions -- 2026-07-06
  206. Hierarchos: Preliminary Findings From a 232M Recurrent Memory-Augmented Assistant Model -- 2026-07-03
  207. DiScoFormer: One transformer for density and score, across distributions -- 2026-07-03
  208. Mapping Local Nodes — Visualizing Model Activation Paths -- 2026-07-03
  209. Microsoft has taken down fastcontext model from everywhere -- 2026-07-02
  210. Orthrus (diffusion head) trained Qwen 3.5/3.6 and Gemma 4 models are dropping soon -- 2026-07-02
  211. SenseNova-U1-8b-MoT-Infographic-V2 (released yesterday) - An open source SOTA beast for infographic design and image editing. -- 2026-07-02
  212. on Dario's statement -- 2026-07-02
  213. After i distilled a 26B into a 4B to cut false positives I'll probably use the base model after all. -- 2026-07-02
  214. LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active -- 2026-07-01
  215. [Editorial] Google Gemini Model Family Updates -- 2026-07-01
  216. InternScience/Agents-A1 · Hugging Face -- 2026-07-01
  217. [Editorial] Video Submission -- 2026-07-01
  218. Accio-Lab/Dressage -- 2026-07-01
  219. dondai1234/agent-browser -- 2026-07-01
  220. Emry: an event-sourced, local-first observability engine for long training runs -- 2026-07-01
  221. Position: AI Safety Requires Effective Controllability -- 2026-07-01
  222. Godot will no longer accept AI-authored code contributions -- 2026-07-01
  223. I built a desktop AI that scrubs your PII locally before it hits the cloud -- 2026-07-01
  224. DeepSeek V4 official version launching mid-July -- 2026-06-30
  225. OpenPangu-2.0-Flash: 92B MoE (6B active) on Ascend with 512K context -- 2026-06-30
  226. Anthropic's Amodei: "Open Source models [could take us to] a very dangerous place." -- 2026-06-30
  227. Even Google still believes in small models for coding — Gemma 4 31B hackathon at 1500 tok/s -- 2026-06-30
  228. deepseek-ai/DeepSeek-V4-Pro-DSpark • Huggingface -- 2026-06-29
  229. High-quality GLM-5.2 Quant on 4x DGX Spark - Guide, Results, and Comps -- 2026-06-29
  230. Ornith 35B is great so far -- 2026-06-29
  231. Book Review: Domain-Specific Small Language Models by Guglielmo Iozzia -- 2026-06-29
  232. Previewing GPT‑5.6 Sol: a next-generation model -- 2026-06-29
  233. [Editorial] -- 2026-06-29
  234. [Editorial] -- 2026-06-29
  235. [Editorial] -- 2026-06-27
  236. US government allows Anthropic limited release of AI model that sparked cybersecurity concerns | CNN Business -- 2026-06-27
  237. Fable 5 return RUMORED with some hints in CC -- 2026-06-27
  238. [Editorial] -- 2026-06-27
  239. [Editorial] -- 2026-06-27
  240. [Editorial] -- 2026-06-27
  241. trotsky1997/OpenFugu -- 2026-06-27
  242. SPIRAL: Learning to Search and Aggregate -- 2026-06-27
  243. LFM2.5 230M running in-browser at 1,400 tok/s using custom WebGPU kernels -- 2026-06-26
  244. Got GLM-5.2 + MTP speculative decode running on 4x DGX Spark (GB10) — and the build piece the public recipe is missing -- 2026-06-26
  245. Run a vLLM Server on HF Jobs in One Command -- 2026-06-26
  246. Running Sonnet 4.6 on every Instagram DM for a 7-location restaurant. 97% cache hit is the only reason it's affordable -- 2026-06-26
  247. [Editorial] rupixel -- 2026-06-26
  248. GLM-5.2 is a step change for open agents -- 2026-06-25
  249. poolside/Laguna-M.1 · Hugging Face - 225B-A23B -- 2026-06-25
  250. Mimo 2.5 is _fast_ at large context (dual RTX Pro 6000) -- 2026-06-25
  251. The Eagle(3) has landed (for Qwen) -- 2026-06-25
  252. CPU-only TTS benchmark: Kokoro 82M vs Supertonic 3 vs Inflect-Nano-v1 (4.6M params), with UTMOS scoring on every sample -- 2026-06-25
  253. Ideogram 4: Open Image Model at the Forefront of Design -- 2026-06-25
  254. QUEST-35B: Open-Source Deep Research Agent Trained on 32 H100s — Full Recipe, Weights, and Data Released -- 2026-06-25
  255. North Mini Code: 4-Bit Quant + Ollama + OpenRouter — Now Runs on 20GB -- 2026-06-25
  256. EdgeRazor: Mixed-Precision Quantization-Aware Distillation Down to 1.58-Bit -- 2026-06-25
  257. [Editorial] Qwable-3.6-27b — Open Model Release -- 2026-06-25
  258. ScenemaAI/scenema-audio — Audio Generation Model -- 2026-06-25
  259. "Tokaine Addiction" — PhD Student's Colleague Can't Stop Running AI Agents -- 2026-06-25
  260. Suitcase Robot Uses Gas Sensor to Modulate LLM Sampler Temperature in Real-Time -- 2026-06-25
  261. GLM-5.2 is the first open-weights model to cross 80% on Terminal-Bench -- 2026-06-19
  262. unsloth GLM-5.2-GGUF, including 2bit at 238GB -- 2026-06-19
  263. Running local models is good now -- 2026-06-19
  264. Gemma 12b less than 10 watts 6.5pp 1.3tg -- 2026-06-19
  265. Rio 3.5 397B could've simply been a semi-failed embezzling of funding -- 2026-06-19
  266. VibeThinker-3B: what is this witchcraft? Killing it at MathQA like it has ~30B parameters -- 2026-06-18
  267. Get in here: Community model build thread -- 2026-06-18
  268. bartowski/command-a-plus-05-2026-GGUF · Hugging Face -- 2026-06-18
  269. nvidia/Nemotron-Labs-Diffusion-14B -- 2026-06-18
  270. [Editorial] Microsoft Majorana 2: Quantum Discovery via Agentic AI -- 2026-06-17
  271. cuTile Rust: Safe, Data-Race-Free GPU Kernels in Rust -- 2026-06-17
  272. [Editorial] Adversarial AI Research — The Malicious Use of Artificial Intelligence -- 2026-06-17
  273. Heretic Grimoire 1.4: Takedown-Resilient Model Backup — 9KB Reproducible Manifests + IPFS Distribution -- 2026-06-16
  274. z.ai Poll: MIT-Licensed Open Weights Are Losing -- 2026-06-16
  275. [Editorial] AMD CEO Lisa Su Challenges NVIDIA's GPU Dominance -- 2026-06-16
  276. PonyExl3: EXL3 Quantization Ported to Apple Silicon — 2700 tok/s Prefill, 68.5 tok/s Decode on M5 Max -- 2026-06-16
  277. My Homelab AI Dev Platform -- 2026-06-16
  278. [Editorial] Video Content -- 2026-06-16
  279. [Editorial] Video Content -- 2026-06-16
  280. [Editorial] Video Content -- 2026-06-16
  281. [Editorial] StandardAgents Arrow-JS — JavaScript Agent Framework -- 2026-06-16
  282. archex: Local-First Deterministic Code-Context for AI Agents — No API Key, No Telemetry (Apache 2.0) -- 2026-06-16
  283. Ironsmith: Open Source macOS App That Creates macOS Apps From Prompts — Works With Local Models -- 2026-06-16
  284. Claude Mythos 5 + Fable 5 Are Here And The Numbers Are INSANE -- 2026-06-10
  285. What it feels like to work with Mythos -- 2026-06-10
  286. [Editorial] Reuven Cohen on maximizing AI tools -- 2026-06-10
  287. How much of Thermo Fisher's antibody data has been manipulated? -- 2026-06-09
  288. News Sites are Blocking Internet Archive over AI Scraping Fears -- 2026-06-09
  289. OneDrive data now has an expiry date -- 2026-06-09
  290. Dopamine Fracking -- 2026-06-09
  291. ESP32 Bit Pirate, a Hardware Hacking Tool with WebCLI That Speaks Every Protocol -- 2026-06-09
  292. Porting the ThinkPad X61 to Coreboot -- 2026-06-09
  293. Holo3.1: Fast & Local Computer Use Agents -- 2026-06-03
  294. Fused MoE dispatch kernel in pure Triton: 89-131% of Megablocks, runs on AMD with zero code changes -- 2026-06-03
  295. ReAligned-Qwen3.5 Release -- 2026-06-03
  296. KANX: A production-ready Kolmogorov-Arnold Network library -- 2026-06-03
  297. Introducing Mellum2: A 12B Mixture-of-Experts Model by JetBrains -- 2026-06-02
  298. Tencent Hy-MT2 is now under Apache License 2.0 -- 2026-06-02
  299. fxyz666/LogicPipe -- 2026-06-02
  300. [OSS] dlmserve - first serving engine for diffusion language models -- 2026-06-02
  301. veryyoldman/Genspark-AI -- 2026-06-02
  302. Qwen/Qwen-Image-Bench · Hugging Face -- 2026-06-02
  303. Welcome NVIDIA Cosmos 3: The First Open Omni-model for Physical AI Reasoning and Action -- 2026-06-02
  304. Nvidia announces new AI chip for personal computers -- 2026-06-02
  305. Nvidia LocateAnything - Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding. (10x faster than Qwen3-VL) -- 2026-06-02
  306. The $500K AI Film That "Premiered at Cannes" Was Not in the Official Festival -- 2026-06-02
  307. [Editorial] Barracuda Nightmare Eclipse Zero-Days -- 2026-05-29
  308. [Editorial] Video -- 2026-05-29
  309. Krasis update: Qwen3.6-35B-A3B (Q4) at reading speed, 1x 8GB 3070 Mobile laptop (32GB RAM) -- 2026-05-28
  310. Benchmarked Needle 26M vs Qwen3-0.6B on CPU function calling, 50 queries across 5 difficulty tiers. The 23x smaller model wins on accuracy and is 4.4x faster. -- 2026-05-28
  311. Small comparison on full compute performance (Anima) of 5090 vs 6000 PRO MaxQ vs 6000 PRO WS/SE -- 2026-05-28
  312. CXMT started selling ram to corsair -- 2026-05-28
  313. TTS Benchmark Comparison (all known TTS up until May 2026) -- 2026-05-28
  314. BitCPM-CANN: Native 1.58-Bit Large Language Model Training on Ascend NPU -- 2026-05-26
  315. Carbon: Decoding the Language of Life -- 2026-05-26
  316. AI content detector based on Qwen 0.8b fine-tuned on Pangram dataset -- 2026-05-26
  317. A tool I built to generate 3D objects with functional, articulated parts. It's on github, and is mostly LLM-agnostic. -- 2026-05-26
  318. Show HN: Audiomass – a free, open-source multitrack audio editor for the web -- 2026-05-26
  319. CAM-LDS: Cyber Attack Manifestations for Automatic Interpretation of System Logs and Security Alerts -- 2026-05-26
  320. aaron-kidwell/goLoL -- 2026-05-26
  321. [Editorial] -- 2026-05-26
  322. lordx64/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled -- 2026-05-22
  323. Jackrong/Qwopus3.5-9B-Coder-GGUF · Hugging Face -- 2026-05-22
  324. Qwen 3.6 35B GGUF: NTP vs MTP quantization results across GPUs and CPUs -- 2026-05-22
  325. [WIP] Gemma 4 MTP -- 2026-05-22
  326. Lemonade v10.5.1: an MTP + ROCm 7.13 quick start for Strix Halo -- 2026-05-22
  327. gemma-4-Ortenzya-The-Creative-Wordsmith-31B-it-uncensored-heretic is Out Now -- 2026-05-22
  328. Gemma-4-Gembrain-31B-it-uncensored-heretic Is Out Now -- 2026-05-22
  329. Newbie vibe coding experience: Shifting from Claude Sonnet 4.6 to Qwen3.6-35B-A3B-UD-Q6_K -- 2026-05-22
  330. HalBench: Open Sycophancy and Hallucination Benchmark — 12,800 Graded Responses -- 2026-05-21
  331. DystopiaBench: 42 LLMs tested on willingness to build dystopian systems -- 2026-05-21
  332. Guardrails take an 8B model from 53% to 99% on agentic tasks [ACM CAIS '26 preprint] -- 2026-05-21
  333. Introducing the Ettin Reranker Family -- 2026-05-20
  334. OlmoEarth v1.1: A more efficient family of models -- 2026-05-20
  335. Mini Shai-Hulud Strikes Again: 314 npm Packages Compromised -- 2026-05-20
  336. [Editorial] -- 2026-05-20
  337. exploitbench/exploitbench -- 2026-05-20
  338. bytedance released an open source model that attempts to do just about anything with only 3b parameters -- 2026-05-19
  339. Sapient Intelligence releases HRM-Text 1B: 40B tokens, ~$1k pretrain, beats Llama3.2 3B on MATH and DROP -- 2026-05-19
  340. PaddleOCR 3.5: Running OCR and Document Parsing Tasks with a Transformers Backend -- 2026-05-19
  341. DeepSeek-V4-Flash W4A16+FP8 with MTP self-speculation: 85 tok/s @ 524k on 2x RTX PRO 6000 Max-Q -- 2026-05-15
  342. NCCL-Free Tensor Parallelism on Dual Blackwell PCIe — llama.cpp b9095 -- 2026-05-15
  343. The Qwen 3.6 35B A3B hype is real!!! -- 2026-05-15
  344. A First Comprehensive Study of TurboQuant: Accuracy and Performance -- 2026-05-15
  345. Developing Open Source LLM from Ground Up — DeepSeek V3 Architecture on Single Blackwell GPU -- 2026-05-15
  346. z-lab/Qwen3.6-27B-DFlash -- 2026-05-15
  347. [Editorial] Arxiv Research Paper 2603.15423 -- 2026-05-11
  348. [Editorial] Arxiv Research Paper 2603.17378 -- 2026-05-11
  349. [Editorial] Gadi Evron — Predicting AI Trends -- 2026-05-11
  350. iai-mcp: Persistent Claude Memory Daemon — 5 Months of Daily Use, Now Open Source -- 2026-05-08
  351. [Editorial] Claude Skins — Custom Claude Code Personas -- 2026-05-08
  352. [Editorial] AI Agents Projects & Tutorials Collection -- 2026-05-08
  353. Does the "6 months gap" still hold? -- 2026-05-07
  354. Claude Code @ Opus 4.7 vs OpenCode @ qwen3.6:27b. Both shipped a playable cozy roguelite. -- 2026-05-07
  355. Fine-tuned Qwen3.6-35B-A3B DeltaNet experiment -- 2026-05-07
  356. Zyphra/ZAYA1-8B -- 2026-05-07
  357. Virtual violin produces realistic sounds (MIT) -- 2026-05-06
  358. Lightricks/LTX-2.3-22b-IC-LoRA-HDR -- 2026-05-06
  359. "Second Thoughts" — A small transformer reads output and feeds it back as a refinement loop, drastically improving a 1.7B model's coding -- 2026-05-06
  360. A plug-n-play open-source pruning tool that is workload-aware (Sculpt) -- 2026-05-06
  361. "I" is not singular — 4 LLM agents with per-agent LoRA on a single RTX 3070 8GB -- 2026-05-06
  362. I made a visualizer for Hugging Face models (hfviewer.com) -- 2026-05-06
  363. Humanoid Robot Actuators -- 2026-05-05
  364. Qwen-Scope: Official Sparse Autoencoders (SAEs) for Qwen 3.5 models -- 2026-05-05
  365. guidelabs/steerling-8b — Steerable 8B Model -- 2026-05-05
  366. nvidia/Gemma-4-31B-IT-NVFP4 — NVIDIA FP4 Quantized Gemma 4 -- 2026-05-05
  367. Qwen 3.6 27B Neo Code Q4_KM running tax accounting on a Ryzen laptop -- 2026-05-05
  368. Qwen3.6-27B gets stuck in a self-affirming thinking loop -- 2026-05-05
  369. Mistral Medium 3.5 128B — MLX 4-bit Conversion with Vision and 256K Context -- 2026-05-04
  370. Chirp: Native Offline Text-to-Speech Desktop App (Kokoro + Qwen3-TTS) -- 2026-05-04
  371. TinyMozart v2 85M — Unconditional MIDI Piano Music Generation -- 2026-05-04
  372. SuperGemma4-26B Uncensored MLX 4-bit v2 -- 2026-05-04
  373. Open WebUI Skill for Auto-Creating Tools with Qwen3.6/Gemma4 -- 2026-05-04
  374. quant-whisper: Terminal-Native Algo Trading Engine with Local LLM Inference -- 2026-05-04
  375. [Editorial] Finding Zero-Days with Any Model -- 2026-05-01
  376. Qwen 3.6 27B Makes Huge Gains in Agency on Artificial Analysis - Ties with Sonnet 4.6 -- 2026-04-30
  377. First direct side by side MoE vs Dense comparison -- 2026-04-30
  378. GPT-6 Confirmed -- 2026-04-30
  379. Google to invest up to $40 billion in Anthropic as search giant spreads its AI bets -- 2026-04-30
  380. convert : add support for Nemotron Nano 3 Omni by danbev - llama.cpp PR #22481 -- 2026-04-30
  381. SWE-bench Verified no longer measures frontier coding capabilities -- 2026-04-29
  382. Opus 4.7: Are these first signs of model collapse? -- 2026-04-29
  383. Qwen3.6-27B IQ4_XS FULL VRAM with 110k context -- 2026-04-28
  384. Can we already use Google's TurboQuant (TQ) for KV Cache in llama-server? Or are we waiting for a PR? -- 2026-04-28
  385. Gemma 4 beats Qwen 3.5 (UPDATE), and Qwen 3.6 27B + MiniMax M2.7 is the best OpenCode setup -- 2026-04-28
  386. Tried Qwen3.6-27B-UD-Q6_K_XL.gguf with CloudeCode, well I can't believe but it is usable -- 2026-04-28
  387. BitNet is the AI future? -- 2026-04-28
  388. How Anthropic's Model Context Protocol Allows for Easy Remote Execution -- 2026-04-27
  389. [Editorial] LinkedIn: AI Industry Perspective -- 2026-04-27
  390. [Editorial] PolinRider — Open Source Malware Analysis -- 2026-04-27
  391. FP4 inference in llama.cpp (NVFP4) and ik_llama.cpp (MXFP4) landed -- 2026-04-27
  392. [Editorial] Bonsai-8B MLX 1-bit -- 2026-04-27
  393. VRAM.cpp: Running llama-fit-params directly in your browser -- 2026-04-27
  394. Thoughts on using an AMD Alveo V80 FPGA as a poor man's Taalas HC1 -- 2026-04-27
  395. [Editorial] hw-smi — Cross-Platform Hardware Monitor -- 2026-04-27
  396. China's DeepSeek valuation rockets above $20B!! -- 2026-04-24
  397. [Editorial] DeepSeek Open-Sources Tile Kernels -- 2026-04-24
  398. DeepSeek v4 -- 2026-04-24
  399. NSA is using Anthropic's Mythos despite blacklist -- 2026-04-22
  400. [Editorial] Mad Bugs: Reverse Engineering Deep Dive -- 2026-04-22
  401. Anyone deployed Kimi K2.6 on their local hardware? -- 2026-04-21
  402. 24/7 Headless AI Server on Xiaomi 12 Pro (Guide & Benchmarks) Gemma4 VS Qwen2.5 -- 2026-04-21
  403. Gemm4:e4B-IT good at instructions following no refusals. -- 2026-04-21
  404. [Editorial] arxiv:2506.02153 — AI Research -- 2026-04-20
  405. [Editorial] arxiv:1503.02531 — Classic ML/AI Paper -- 2026-04-20
  406. [Editorial] arxiv:2604.06169 — Recent AI Research -- 2026-04-20
  407. [Editorial] Why Your LLM Is Slow (and the Eight Fixes) -- 2026-04-20
  408. Hot Experts in your VRAM! Dynamic expert cache in llama.cpp for 27% faster token generation -- 2026-04-20
  409. Gemma 4 26B fabricated an entire code audit. I have the forensic evidence from the database. -- 2026-04-16
  410. [Editorial] Looped LLMs Are the Nuclear Fusion of AI -- 2026-04-16
  411. All elementary functions from a single binary operator -- 2026-04-14
  412. GLM 5.1 "Actually wait" — the current thinking SOTA open source -- 2026-04-14
  413. Are i-Quants overrated? -- 2026-04-14
  414. Gemma 4 E4B vs Qwen3.5-4B on document tasks — sub-scores tell a different story -- 2026-04-14
  415. Using NPU for something useful — whisper-npu for Intel NPU speech-to-text -- 2026-04-14
  416. [Editorial] -- 2026-04-14
  417. [Editorial] -- 2026-04-13
  418. Qwen3.5-397B is shockingly useful at Q2 -- 2026-04-13
  419. [Editorial] -- 2026-04-13
  420. Liquid AI releases LFM2.5-VL-450M - structured visual understanding at 240ms -- 2026-04-13
  421. Muse Spark: Scaling towards personal superintelligence -- 2026-04-13
  422. [Editorial] -- 2026-04-13
  423. [Editorial] -- 2026-04-13
  424. Abliterating Qwen3.5-397B on a Mac Studio revealed that MoE models encode refusal differently than dense models — safety refusals route through expert selection and survive weight-baking -- 2026-04-09
  425. Finally Abliterated Sarvam 30B and 105B! -- 2026-04-09
  426. Sam Altman may control our future – can he be trusted? -- 2026-04-07
  427. Qwen3.6-Plus -- 2026-04-07
  428. Omnivoice - 600+ Language Open-Source TTS with Voice Cloning and Design -- 2026-04-07
  429. [Editorial] Generative Models and High-Quality Output -- 2026-04-07
  430. A cryptography engineer's perspective on quantum computing timelines -- 2026-04-07
  431. [Editorial] -- 2026-03-28
  432. [Editorial] -- 2026-03-28
  433. [Editorial] -- 2026-03-28
  434. [Editorial] -- 2026-03-28
  435. [Editorial] -- 2026-03-28
  436. The current state of the Chinese LLMs scene -- 2026-03-26
  437. Alibaba confirms they are committed to continuously open-sourcing new Qwen and Wan models -- 2026-03-26
  438. Cursor's Composer 2 apparently built on Kimi K2.5 without attribution -- 2026-03-26
  439. Nemotron Cascade 2 30B A3B -- 2026-03-26
  440. Controllable Reasoning Models Are Private Thinkers -- 2026-03-26
  441. Nvidia built a silent opinion engine into NemotronH to gaslight you and they're not the only ones doing it -- 2026-03-26
  442. Gemini thoughts turned very violent -- 2026-03-26
  443. Qianfan-OCR: End-to-End 4B Document Intelligence VLM — SOTA on OmniDocBench -- 2026-03-24
  444. AMD Quark: Under-the-Radar Quantization Tool with MXFP4 Post-Training -- 2026-03-24
  445. We Compressed 6 LLMs and Found They Don't Degrade the Same Way -- 2026-03-24
  446. HumeAI/tada-1b — Expressive Audio Generation Model -- 2026-03-24
  447. The Resolv Hack: How One Compromised Key Printed $23M -- 2026-03-24
  448. Show HN: Three new Kitten TTS models – smallest less than 25MB -- 2026-03-23
  449. Flash-MoE: Running a 397B Parameter Model on a Laptop -- 2026-03-23
  450. [Editorial] Distillation Techniques -- 2026-03-23
  451. Jackrong/Qwen3.5-2B-Claude-4.6-Opus-Reasoning-Distilled-GGUF -- 2026-03-23
  452. Local Qwen 8B + 4B completes browser automation by replanning one step at a time -- 2026-03-23
  453. Does imatrix calibration data affect writing style? I ran a blind-scored experiment -- 2026-03-23
  454. [Editorial] Claude Code Channels Documentation -- 2026-03-23
  455. github.com -- 2026-03-23
  456. Claude Channels vs Dispatch vs Remote Control -- 2026-03-23
  457. [Editorial] Claude Code Is Brilliant Until the Repo... -- 2026-03-23
  458. I made mcp-optimizer - stop wasting tokens on idle MCP servers -- 2026-03-23
  459. [Editorial] Anthropic Is Coming for Lovable et al. -- 2026-03-23
  460. [Editorial] -- 2026-03-20
  461. mlx-tune – fine-tune LLMs on your Mac (SFT, DPO, GRPO, Vision) with an Unsloth-compatible API -- 2026-03-20
  462. Hunter Alpha was a stealth model revealed on March 18th as an early testing version of MiMo-V2-Pro. -- 2026-03-19
  463. mistralai/Mistral-Small-4-119B-2603 -- 2026-03-19
  464. LiquidAI/LFM2-24B-A2B -- 2026-03-19
  465. Mistral releases an official NVFP4 model, Mistral-Small-4-119B-2603-NVFP4! -- 2026-03-18
  466. Omnicoder-Claude-4.6-Opus-Uncensored-GGUF -- 2026-03-18
  467. You can run LLMs on your AMD NPU on Linux! -- 2026-03-18
  468. Rick Beato: "How AI Will Fail Like The Music Industry" (and why local LLMs will win) -- 2026-03-18
  469. [Editorial] Your AI Forgot Everything Again — There's a Fix for That -- 2026-03-18
  470. ylytdeng/wechat-decrypt -- 2026-03-18
  471. [Editorial] LLM Architecture Gallery -- 2026-03-17
  472. Mistral small 4 PR on transformers. -- 2026-03-17
  473. I spent a weekend doing layer surgery on 6 different model architectures. There's a "danger zone" at ~50-56% depth. -- 2026-03-17
  474. Nemotron 3 Super and the no free lunch problem -- 2026-03-17
  475. Source code of Swedish e-government services has been leaked -- 2026-03-14
  476. Bucketsquatting is (finally) dead -- 2026-03-14
  477. E2E encrypted messaging on Instagram will no longer be supported after 8 May -- 2026-03-14
  478. Fine-tuned Qwen3 SLMs (0.6-8B) beat frontier LLMs on narrow tasks -- 2026-03-12
  479. Towards Cold-Start Drafting and Continual Refining: A Value-Driven Memory Approach with Application to NPU Kernel Synthesis -- 2026-03-12
  480. Heretic Defeats GPT-OSS with Arbitrary-Rank Ablation (ARA) Decensoring -- 2026-03-11
  481. Fat Fish — A Proper Upscale and Prune of Mistral Nemo -- 2026-03-11
  482. Optimizing Qwen3 Coder for RTX 5090 and PRO 6000 — Community Benchmarking Infrastructure -- 2026-03-11
  483. Yann LeCun's AMI Labs Raises $1.03 Billion to Build World Models -- 2026-03-11
  484. Did Alibaba just kneecap its powerful Qwen AI team? -- 2026-03-07
  485. [Editorial] You're 1,191 Days Late — Here's What to Do -- 2026-03-07
  486. [Editorial] When Anonymity Fades: What New Research Reveals -- 2026-03-07
  487. Introducing Modular Diffusers - Composable Building Blocks for Diffusion Pipelines -- 2026-03-07
  488. kyutai-labs/hibiki-zero -- 2026-03-07
  489. [Editorial] OpenClawCity -- 2026-03-07
  490. [Editorial] The Zero-Day Clock Is Ticking -- 2026-03-06
  491. [Editorial] Step-by-Step Guide to Exploiting AI Systems -- 2026-03-06
  492. [Editorial] Unprompted 2026: Top Insights Day One -- 2026-03-06
  493. [Editorial] Unprompted 2026: Top Insights Day Two -- 2026-03-06
  494. YuanLabAI/Yuan3.0-Ultra: 1010B MoE, fully open weights -- 2026-03-05
  495. We could be hours (or less than a week) away from true NVFP4 support in Llama.cpp GGUF format -- 2026-03-05
  496. Step-3.5-Flash-Base & Midtrain (in case you missed them) -- 2026-03-05
  497. Qwen3.5-9B Uncensored Aggressive Release (GGUF) -- 2026-03-05
  498. unknown -- 2026-03-05
  499. [Editorial] David Maynor Security Gist -- 2026-03-04
  500. [Editorial] arXiv:2602.23093 -- 2026-03-04
  501. unpromptedcon.org -- 2026-03-04
  502. Unsloth fixed version of Qwen3.5-35B-A3B is incredible at research tasks -- 2026-03-04
  503. PSA: Qwen 3.5 requires bf16 KV cache, NOT f16!! -- 2026-03-04
  504. Building a Dependency-Free GPT on a Custom OS -- 2026-03-04
  505. Qwen3.5 122B in 72GB VRAM (3x3090) is the best model available at this time — also it nails the "car wash test" -- 2026-03-03
  506. Qwen 3.5 is multimodal. Here is how to enable image understanding in opencode with llama cpp -- 2026-03-03
  507. pplx-embed: State-of-the-Art Embedding Models for Web-Scale Retrieval -- 2026-03-03
  508. andimarafioti/faster-qwen3-tts -- 2026-03-03
  509. We tested RLVR on top of fine-tuned small models across 12 datasets — here's exactly when it helps (and when it doesn't) -- 2026-03-02
  510. [Editorial] Weber Electrodynamics & Thermodynamic Memory -- 2026-03-02
  511. Minuspod: Automatically remove ads from podcasts locally -- 2026-03-02
  512. Squidcasa/midipipe: ALSA Sequencer to plain text and back -- 2026-03-02
  513. [Editorial] -- 2026-02-28
  514. Parakeet.cpp – Parakeet ASR inference in pure C++ with Metal GPU acceleration -- 2026-02-28
  515. [Editorial] -- 2026-02-28
  516. Liquid AI releases LFM2-24B-A2B -- 2026-02-24
  517. Qwen3's most underrated feature: Voice embeddings -- 2026-02-24
  518. After many contributions craft, Crane now officially supports Qwen3-TTS! -- 2026-02-24
  519. We tested the same INT8 model on 5 Snapdragon chipsets. Accuracy ranged from 93% to 71%. Same weights, same ONNX file. -- 2026-02-24
  520. [Editorial] -- 2026-02-24
  521. Anthropic Accuses DeepSeek, Moonshot AI, and MiniMax of Creating 24,000 Fake Claude Accounts -- 2026-02-24
  522. Qwen 3.5 vs Gemini 3 Pro on Screenshot-to-Code: Is the gap finally gone? -- 2026-02-19
  523. [Editorial] The evolution of vision models from CNNs to transformers -- 2026-02-19
  524. [Editorial] An AI agent merged code into 22 widely-used open source projects -- 2026-02-19
  525. [Editorial] AI Agent Security and Supply Chain -- 2026-02-19
  526. Policy Compiler for Secure Agentic Systems -- 2026-02-19
  527. [Editorial] OpenClaw Maestro Threat Assessment -- 2026-02-19
  528. Qwen Released Qwen 3.5 397B and Qwen 3.5 Plus! -- 2026-02-17
  529. Qwen3.5 NVFP4 (Blackwell) is up! -- 2026-02-17
  530. Running Gemma 3n E2B natively on Android via LiteRT -- 2026-02-17
  531. Deploying Open WebUI + vLLM on Amazon EKS -- 2026-02-17
  532. Ling-2.5-1T: 1T Parameter Open-Source Instant Model with 1M Context -- 2026-02-16
  533. Qwen3.5-397B-A17B Unsloth GGUFs — Run on Consumer Hardware -- 2026-02-16
  534. Running Qwen3-Coder-Next 80B on 8GB VRAM — 300x Speedup via Custom Expert Caching -- 2026-02-16
  535. [Editorial] https://www.linkedin.com/posts/ownyourai_i-just-woke-up-to-qwen3-coder-next-80b-activity-7424703876240695297-Nlqf -- 2026-02-04
  536. LiquidAI/LFM2.5-1.2B-Thinking -- 2026-02-04
  537. ByteDance-Seed/Stable-DiffCoder-8B-Instruct · Hugging Face -- 2026-02-03
  538. tencent/Youtu-VL-4B-Instruct -- 2026-02-03
  539. NousResearch/NousCoder-14B -- 2026-02-03
  540. transformers v5 final is out 🔥 -- 2026-01-30
  541. PaddlePaddle/PaddleOCR-VL-1.5 -- 2026-01-30
  542. zai-org/GLM-4.7 -- 2026-01-30
  543. MHA2MLA-VLM: Enabling DeepSeek's Economical Multi-Head Latent Attention across Vision-Language Models -- 2026-01-30
  544. Sharing my set of distilled small language models (3B) + training data in more than 50 low-resource languages -- 2026-01-30
  545. Introducing Kimi K2.5, Open-Source Visual Agentic Intelligence -- 2026-01-28
  546. ~60GB models on coding: GLM 4.7 Flash vs. GPT OSS 120B vs. Qwen3 Coder 30B -- your comparisons? -- 2026-01-28
  547. openbmb/AgentCPM-Report -- 2026-01-28
  548. Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice -- 2026-01-28
  549. Mixture-of-Models: Unifying Heterogeneous Agents via N-Way Self-Evaluating Deliberation -- 2026-01-28
  550. One Year Since the “DeepSeek Moment” -- 2026-01-28
  551. Qwen/Qwen-Image-2512 -- 2026-01-28
  552. Qwen/Qwen3-TTS-12Hz-0.6B-Base -- 2026-01-27
  553. Qwen/Qwen3-VL-Reranker-2B -- 2026-01-27
  554. Qwen/Qwen3-VL-Embedding-2B -- 2026-01-21
  555. Phr00t/Qwen-Image-Edit-Rapid-AIO -- 2026-01-21
  556. Qwen/Qwen3-VL-Reranker-8B -- 2026-01-21
  557. Bartowski comes through again. GLM 4.7 flash GGUF -- 2026-01-21
  558. Step-Audio-R1.1 (Open Weight) by StepFun just set a new SOTA on the Artificial Analysis Speech Reasoning leaderboard -- 2026-01-20
  559. GLM-4.7-Flash -- 2026-01-20
  560. stepfun-ai/Step-Audio-R1.1 -- 2026-01-20
  561. tencent/HY-MT1.5-1.8B -- 2026-01-20
  562. Introducing GLM-Image -- 2026-01-14
  563. GPT-OSS -> MLA conversion breakthrough (20B), still looking for compute + collaborators -- 2026-01-14
  564. FrogBoss 32B and FrogMini 14B from Microsoft -- 2026-01-14
  565. Qwen/Qwen3-VL-Embedding-8B -- 2026-01-14
  566. Liquid AI releases LFM2-2.6B-Transcript, an incredibly fast open-weight meeting transcribing AI model on-par with closed-source giants. -- 2026-01-13
  567. New llama.cpp 30% faster.... -- 2026-01-13
  568. Qwen/Qwen-Image-Edit-2511 -- 2026-01-13
  569. nvidia/Nemotron-Orchestrator-8B -- 2026-01-13
  570. tencent/HY-WorldPlay -- 2026-01-09
  571. meituan-longcat/LongCat-Image -- 2026-01-09
  572. Introducing Falcon H1R 7B -- 2026-01-09
  573. [Editorial] https://github.com/hiyouga/LlamaFactory -- 2026-01-07
  574. Tongyi-MAI/MAI-UI-8B · Hugging Face -- 2026-01-07
  575. MultiverseComputingCAI/HyperNova-60B · Hugging Face -- 2026-01-07
  576. Anyone tried IQuest-Coder-V1 yet? The 40B numbers look wild -- 2026-01-06
  577. open-thoughts/OpenThinker-Agent-v1 -- 2026-01-06
  578. [Editorial] https://www.alwaysfurther.ai/blog/train-4b-model-to-beat-claude-sonnet-gemini -- 2026-01-05
  579. [Experimental] Gemma 3 4B - Dark CoT: Pushing 4B Reasoning to 33%+ on GPQA Diamond -- 2026-01-05
  580. Youtu-LLM-2B-GGUF is here! -- 2026-01-05
  581. upstage/Solar-Open-100B -- 2026-01-05
  582. A zero-setup agent that benchmarks multiple open / closed source LLMs on your specific problem / data -- 2026-01-02
  583. What is a good model for assisting with patching source code? -- 2026-01-02
  584. Just got an RTX Pro 6000 - need recommendations for processing a massive dataset with instruction following -- 2026-01-02
  585. MiniMaxAI/MiniMax-M2.1 -- 2026-01-02
  586. [Editorial] https://www.linkedin.com/posts/ownyourai_wow-its-raining-korean-open-ai-models-today-activity-7412133834667876352-fcpl -- 2025-12-31
  587. Naver (South Korean internet giant), has just launched HyperCLOVA X SEED Think, a 32B open weights reasoning model and HyperCLOVA X SEED 8B Omni, a unified multimodal model that brings text, vision, and speech together -- 2025-12-31
  588. allenai/Olmo-3-7B-Instruct -- 2025-12-31
  589. meituan-longcat/LongCat-Video-Avatar -- 2025-12-31
  590. Mistral AI’s December -- 2025-12-31
  591. MBZUAI releases K2-V2 - 70B fully open model. -- 2025-12-22
  592. microsoft/Fara-7B -- 2025-12-22
  593. I've been experimenting with SLM's a lot recently. My goal was to prove even SLMs can be accurate with the right architecture behind it. -- 2025-12-22
  594. Alibaba Tongyi Open Sources Two Audio Models: Fun-CosyVoice 3.0 (TTS) and Fun-ASR-Nano-2512 (ASR) -- 2025-12-19
  595. My professor lent me an A6000, so I tried to build a coding model. Here is Anni! (Qwen3-14B Fine-tune) -- 2025-12-19
  596. Two years ago, I was just a math major. Now I've built the 1.5B router model used by HuggingFace. Can I bring it to Cursor? -- 2025-12-19
  597. 🚀 New: Olmo 3.1 Think 32B & Olmo 3.1 Instruct 32B -- 2025-12-16
  598. ByteDance/Dolphin-v2 -- 2025-12-16
  599. zai-org/AutoGLM-Phone-9B-Multilingual -- 2025-12-16
  600. [Editorial] https://www.linkedin.com/posts/eric-vyacheslav-156273169_a-7m-model-just-surpassed-deepseek-r1-gemini-activity-7405985266043297792-s1Jn -- 2025-12-15
  601. Nemotron 3 Nano \- A new Standard for Efficient, Open, and Intelligent Agentic Models -- 2025-12-15
  602. Tiny-A2D: An Open Recipe to Turn Any AR LM into a Diffusion LM -- 2025-12-12
  603. New in llama.cpp: Model Management -- 2025-12-12
  604. SoK: a Comprehensive Causality Analysis Framework for Large Language Model Security -- 2025-12-12
  605. Generating synthetic test data for LLM applications (our approach) -- 2025-12-12
  606. PaCoRe: The first open-source deep think 8B model beats GPT-5 on HMMT25 -- 2025-12-11
  607. Building RNJ-1: What makes It different from Gemma 3? -- 2025-12-11
  608. EssentialAI/rnj-1 -- 2025-12-11
  609. VulnLLM-R: Specialized Reasoning LLM with Agent Scaffold for Vulnerability Detection -- 2025-12-11
  610. From Azure Functions to FreeBSD -- 2025-12-10
  611. RnJ-1-Instruct FP8 Quantization -- 2025-12-10
  612. Operator Mech v2.5: A Compact Structural-Reasoning Kernel for Local Models (YAML, 7B–13B Optimized) -- 2025-12-10
  613. Masked Diffusion Models as Energy Minimization -- 2025-12-10
  614. https://huggingface.co/Doradus/Hermes-4.3-36B-FP8 -- 2025-12-09
  615. Support for rnj-1 now in llama.cpp -- 2025-12-09
  616. Comfy-Org/flux2-dev -- 2025-12-09
  617. baidu/ERNIE-4.5-VL-28B-A3B-Thinking -- 2025-12-09
  618. Guidance: A cheat code for diffusion models -- 2025-12-09
  619. OVHcloud on Hugging Face Inference Providers 🔥 -- 2025-12-05
  620. smallevals - Tiny 0.6B Evaluation Models and a Local LLM Evaluation Framework -- 2025-12-05
  621. I cooked abliterated gemma3-27b-it with norm-preserving technique -- 2025-12-04
  622. allenai/Olmo-3-1125-32B -- 2025-12-03
  623. EmbeddingGemma: Powerful and Lightweight Text Representations -- 2025-12-03
  624. Building SFT from scratch - results & learnings -- 2025-12-03
  625. Qwen3 VL built from scratch with PyTorch -- 2025-12-03
  626. LM Studio beta supports Qwen3 80b Next. -- 2025-12-03
  627. RTX 5090 + Qwen 30B MoE @ 135 tok/s in NVFP4 - Full guide with C++ patches -- 2025-12-02
  628. PleIAs/Baguettotron -- 2025-12-01
  629. nvidia/ChronoEdit-14B-Diffusers -- 2025-12-01
  630. 20x Faster TRL Fine-tuning with RapidFire AI -- 2025-12-01
  631. Optimizing Token Generation in llama.cpp's CUDA Backend -- 2025-12-01
  632. [Editorial] https://www.linkedin.com/posts/erudenko_claudecode-aiagents-opensource-activity-7399658246443216896-dIU0 -- 2025-11-28
  633. DataArcTech/DataArc-SynData-Toolkit -- 2025-11-28
  634. facebook/sam-3d-body-dinov3 -- 2025-11-28
  635. WeiboAI/VibeThinker-1.5B -- 2025-11-28
  636. In depth analysis of Nvidia's Jet Nemotron models -- 2025-11-26
  637. Hidden causes of LLM latency, its not just the model size -- 2025-11-26
  638. A million ways to die from a data race in Go -- 2025-11-26
  639. You're using HuggingFace wrong. Stop downloading pre-quantized GGUFs and start building hardware-optimized, domain-specific models. Here's the pipeline I built to do it. -- 2025-11-26
  640. Hardcore function calling benchmark in backend coding agent. -- 2025-11-26
  641. Luo-Yihao/FaithC -- 2025-11-20
  642. ReTabAD: A Benchmark for Restoring Semantic Context in Tabular Anomaly Detection -- 2025-11-20
  643. Android Developer Verification Starts as Google Partially Retreats on Measures -- 2025-11-20
  644. How are you all orchestrating multi-agent workflows (beyond one-shot prompt chaining)? -- 2025-11-20
  645. deliveryhero/asya -- 2025-11-20
  646. Do we rely too much on huggingface? Do you think they’ll eventually regulate open source models? Is there any way to distribute them elsewhere? -- 2025-11-18
  647. Build a DeepSeek model from scratch -- 2025-11-18
  648. Easily Build and Share ROCm Kernels with Hugging Face -- 2025-11-18
  649. ibm-granite/granite-4.0-h-350m -- 2025-11-14
  650. tencent/HunyuanWorld-Mirror -- 2025-11-14
  651. [Editorial] https://www.linkedin.com/posts/mreichstein_cybersecurity-carhacking-physicalsecurity-activity-7394425210877218816-NPyT -- 2025-11-13
  652. On USB HID, Keyboard LEDs, and device emulation (2024) -- 2025-11-13
  653. [Editorial] https://www.linkedin.com/posts/stuart-winter-tear_so-reportedly-yann-lecun-plans-to-leave-activity-7394396547276460032-gEE5 -- 2025-11-13
  654. inclusionAI/LLaDA2.0-mini-preview -- 2025-11-13
  655. [Editorial] https://www.linkedin.com/posts/ismaelvelasco_theres-an-ai-text-model-comparable-to-sota-activity-7393850964731912192-nZT1 -- 2025-11-12
  656. Last week in Multimodal AI - Local Edition -- 2025-11-12
  657. antarys-ai/antarys -- 2025-11-11
  658. Visualizing Quantization Types -- 2025-11-11
  659. Co-authored a book called "Build DeepSeek from Scratch" | Live Now -- 2025-11-11
  660. Writing your own BEAM -- 2025-11-11
  661. DIY Powerwall Blows Clouds, Competition Out of the Water -- 2025-11-11
  662. Building a PV Solar-Powered Quadcopter -- 2025-11-07
  663. BAAI/Emu3.5-Image -- 2025-11-07
  664. Riemannian Optimization for LoRA on the Stiefel Manifold -- 2025-11-07
  665. Kimi release Kimi K2 Thinking, an open-source trillion-parameter reasoning model -- 2025-11-06
  666. OpenAI asks U.S. for loan guarantees to fund $1T AI expansion -- 2025-11-06
  667. Kimi released Kimi K2 Thinking, an open-source trillion-parameter reasoning model -- 2025-11-06
  668. Reproducing the AWS Outage Race Condition with a Model Checker -- 2025-11-05
  669. Futurelock: A subtle risk in async Rust -- 2025-11-05
  670. Qwen3-VL-32B Q8 speeds in llama.cpp vs vLLM FP8 on a RTX PRO 6000 -- 2025-11-03
  671. Help me decide: EPYC 7532 128GB + 2 x 3080 20GB vs GMtec EVO-X2 -- 2025-11-03
  672. amd/Nitro-E -- 2025-11-03
  673. DeepSeek may have found a new way to improve AI’s ability to remember -- 2025-11-02
  674. Reliable Unlearning Harmful Information in LLMs with Metamorphosis Representation Projection -- 2025-11-02
  675. [Editorial] https://www.linkedin.com/posts/busiel-morley_economic-shifts-in-the-age-of-ai-ugcPost-7390349517612806144-8djS -- 2025-11-01
  676. OpenAI: gpt-oss-safeguard: two open-weight reasoning models built for safety classification (Now on Hugging Face) -- 2025-10-31
  677. briaai/FIBO -- 2025-10-31
  678. vLLM MoE Benchmark Configs for Qwen3 Coder REAP 25B & RTX Pro 6000 -- 2025-10-29
  679. Ollama supports Qwen3-VL locally! -- 2025-10-29
  680. Test results for various models' ability to give structured responses via LM Studio. Spoiler: Qwen3 won -- 2025-10-29
  681. Introducing ExecuTorch 1.0 -- 2025-10-29
  682. [Editorial] new coding model -- 2025-10-28
  683. Cerebras REAP'd GLM4.6: 25%, 30%, 40% pruned FP8 checkpoints on HF! -- 2025-10-28
  684. GLM-4.6 on fresh SWE-bench–style tasks collected in September 2025 -- 2025-10-28
  685. Best LLM for 96G RTX Pro 6000 Blackwell? -- 2025-10-27
  686. Gemma3 model differencies -- 2025-10-27
  687. AMD iGPU + dGPU : llama.cpp tensor-split not working with Vulkan backend -- 2025-10-27
  688. Is GLM 4.5 / 4.6 really sensitive to quantisation? Or is vLLM stupifying the models? -- 2025-10-27
  689. AlphaXiv,Compare the Deepseek-OCR and Mistral-OCR OCR models -- 2025-10-26
  690. Open-Bee/Bee-8B-RL -- 2025-10-26
  691. datalab-to/chandra -- 2025-10-26
  692. Unlock the power of images with AI Sheets -- 2025-10-26
  693. [Editorial] Periodic table for ai algorithms -- 2025-10-26
  694. Reverse Engineering STL Files with FreeCAD -- 2025-10-25
  695. Qwen3-VL-32B-Instruct GGUF with unofficial llama.cpp release to run it (Pre-release build) -- 2025-10-25
  696. Qwen3 Next support in llama.cpp ready for review -- 2025-10-25
  697. How's Halo Strix now ? -- 2025-10-25
  698. 20x Max Plan (€216) takes 2% of weekly Opus usage for a single Deep Research Query. That equals 50 per week if you use it ONLY for this and never continue or respond -- 2025-10-24
  699. Show HN: Cuq – Formal Verification of Rust GPU Kernels -- 2025-10-24
  700. [Editorial] We need open, uncensored, & local -- 2025-10-23
  701. 🚀 HuggingFaceChat Omni: Dynamic policy-baed routing to 115+ LLMs -- 2025-10-23
  702. Preliminary support in llama.cpp for Qualcomm Hexagon NPU -- 2025-10-23
  703. LiquidAI/LFM2-2.6B -- 2025-10-23
  704. riptideslabs/tokenex -- 2025-10-20
  705. Multi-Tenant SaaS's Wildcard TLS: An Overview of DNS-01 Challenges -- 2025-10-20
  706. From cloud to OCP? Be ready to wrangle firmware -- 2025-10-20
  707. FLOSS Weekly Episode 851: Buckets of Money -- 2025-10-20
  708. Learning Lifted Action Models From Traces of Incomplete Actions and States -- 2025-10-20
  709. volantvm/volant -- 2025-10-19
  710. Wireshark 4.6.0 Supports macOS Pktap Metadata (PID, Process Name, etc.) -- 2025-10-19
  711. A classified network of SpaceX satellites is emitting a mysterious signal -- 2025-10-19
  712. Show HN: Largest open-source multimodal AI dataset -- 2025-10-18
  713. KORMo-Team/KORMo-10B-sft -- 2025-10-18
  714. ByteDance/FaceCLIP -- 2025-10-18
  715. We built 3B and 8B models that rival GPT-5 at HTML extraction while costing 40-80x less - fully open source -- 2025-10-17
  716. inclusionAI/Ring-flash-linear-2.0 -- 2025-10-17
  717. Qwen/Qwen3-VL-8B-Instruct -- 2025-10-17
  718. Google Cloud C4 Brings a 70% TCO improvement on GPT OSS with Intel and Hugging Face -- 2025-10-17
  719. This Week in Security: ID Breaches, Code Smell, and Poetic Flows -- 2025-10-14
  720. Show HN: Rebuilt Bible search app to run 100% client-side with Transformers.js -- 2025-10-13
  721. swiss-ai/Apertus-8B-Instruct-2509 -- 2025-10-13
  722. A Childhood Dream, Created and Open Sourced -- 2025-10-13
  723. Introducing the ColBERT Nano series of models. All 3 of these models come in at less than 1 million parameters (250K, 450K, 950K) -- 2025-10-11
  724. LiquidAI/LFM2-8B-A1B -- 2025-10-11
  725. GPT-OSS from Scratch on AMD GPUs -- 2025-10-11
  726. How do I compare cost per token for serverless vs provisioned hardware? -- 2025-10-11
  727. OpenAI is good at deals -- 2025-10-11
  728. meituan-longcat/LongCat-Flash-Chat -- 2025-10-11
  729. adb1274/batchi -- 2025-10-11
  730. Granite4 Small-h 32b-A9b (Q4_K_M) at FULL 1M context window is using only 73GB of VRAM - Life is good! -- 2025-10-09
  731. Run Open AI GPT-OSS on a mobile phone (Demo) -- 2025-10-09
  732. AI21 releases Jamba 3B, the tiny model outperforming Qwen 3 4B and IBM Granite 4 Micro! -- 2025-10-09
  733. inclusionAI/Ling-mini-2.0 -- 2025-10-09
  734. Provable scaling laws of feature emergence from learning dynamics of grokking -- 2025-10-09
  735. SecureV2X: An Efficient and Privacy-Preserving System for Vehicle-to-Everything (V2X) Applications -- 2025-10-09
  736. [Editorial] The Tiny Recursive Mode -- 2025-10-08
  737. deepseek-ai/DeepSeek-V3.1-Terminus -- 2025-10-08
  738. Behavioral Modification Systems in Large Language Models: A Methodological Analysis of Long Conversation Reminders -- 2025-10-08
  739. Should I choose Ada, SPARK, or Rust over C/C++? (2024) -- 2025-10-08
  740. We built this open-source LLM Inference project to boost context generation by up to 15x and now it is being implemented by NVIDIA Dynamo! -- 2025-10-06
  741. Running Qwen3-VL-235B (Thinking & Instruct) AWQ on vLLM -- 2025-10-06
  742. Granite 4.0 Language Models - a ibm-granite Collection -- 2025-10-06
  743. NousResearch/Hermes-4-70B -- 2025-10-06
  744. ibm-granite/granite-4.0-h-small -- 2025-10-06
  745. ibm-granite/granite-4.0-h-tiny -- 2025-10-06
  746. [Editorial] Agentic Tribe -- 2025-10-06
  747. AlexanderYastrebov/onion-vanity-address -- 2025-10-06
  748. Open Printer is an open-source inkjet with DRM-free ink and no subscriptions -- 2025-10-06
  749. Yes, Gemini, A Wii Server Is Possible -- 2025-10-06
  750. Tencent-Hunyuan/Hunyuan3D-Omni -- 2025-10-04
  751. tencent/HunyuanImage-3.0 -- 2025-10-04
  752. swiss-ai/Apertus-8B-2509 -- 2025-10-04
  753. Use Remote Models on iOS with Noema -- 2025-10-04
  754. Best small model <3B for HomeAssistant -- 2025-10-04
  755. Nikity/lille-130m-instruct -- 2025-10-04
  756. Unsupervised Hallucination Detection by Inspecting Reasoning Processes -- 2025-10-03
  757. Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training -- 2025-10-03
  758. Will fine-tuning LLaMA 3.2 11B Instruct on text-only data degrade its vision capabilities? -- 2025-10-03
  759. Does Ollama immobilize GPUs / computing resources? -- 2025-10-02
  760. Any real alternatives to NotebookLM (closed-corpus only)? -- 2025-10-02
  761. Best instruct model that fits in 32gb VRAM -- 2025-10-02
  762. Bring Your Own Data (BYOD) -- 2025-09-30
  763. CohereLabs/command-a-reasoning-08-2025 -- 2025-09-30
  764. Built an MCP server for Claude Desktop to browse Reddit in real-time -- 2025-09-30
  765. 1652933138/eth-address-poisoning-tool -- 2025-09-30
  766. Inside NVIDIA GPUs: Anatomy of high performance matmul kernels -- 2025-09-29
  767. Bit is all we need: binary normalized neural networks -- 2025-09-29
  768. Accelerating Qwen3-8B Agent on Intel® Core™ Ultra with Depth-Pruned Draft Models -- 2025-09-29
  769. The MoE tradeoff seems bad for local hosting -- 2025-09-29
  770. MetalQwen3: Full GPU-Accelerated Qwen3 Inference on Apple Silicon with Metal Shaders – Built on qwen3.c - WORK IN PROGRESS -- 2025-09-28
  771. yangdongchao/UniAudio2 -- 2025-09-28
  772. jimsweb/aiMIDI -- 2025-09-28
  773. Handy – Free open-source speech-to-text app written in Rust -- 2025-09-28
  774. IndexTeam/IndexTTS-2 -- 2025-09-28
  775. beankeji-cloud/SLiteIO -- 2025-09-28
  776. GitHub - shantur/jarvis-mcp: Bring your AI to life—talk to assistants instantly in your browser. Zero hasle, No API keys, No Whisper -- 2025-09-27
  777. A1: Asynchronous Test-Time Scaling via Conformal Prediction -- 2025-09-25
  778. DeepLink-org/DeepTrace -- 2025-09-25
  779. Qwen 3 max released -- 2025-09-24
  780. Clauder, auto-updating toolkit for Claude Code, now ships with 65+ MCP servers -- 2025-09-24
  781. Show HN: I wrote inference for Qwen3 0.6B in C/CUDA -- 2025-09-24
  782. nvidia/canary-1b-v2 -- 2025-09-24
  783. baidu/ERNIE-4.5-21B-A3B-Thinking -- 2025-09-24
  784. Smol2Operator: Post-Training GUI Agents for Computer Use -- 2025-09-24
  785. Seeking Local LLM Recommendations for AST Generation (by Function Calling) -- 2025-09-24
  786. A first stab at packaging llama.cpp in a performance-optimized manner -- 2025-09-23
  787. Model: Qwen3 Next Pull Request llama.cpp -- 2025-09-23
  788. Efficient 4B parameter gpt OSS distillation without the over-censorship -- 2025-09-22
  789. [Project] I created an AI photo organizer that uses Ollama to sort photos, filter duplicates, and write Instagram captions. -- 2025-09-22
  790. Pointer Tagging in C++: The Art of Packing Bits into a Pointer -- 2025-09-22
  791. inclusionAI/Ring-mini-2.0 -- 2025-09-22
  792. Local real-time assistant that remembers convo + drafts a doc -- 2025-09-22
  793. XiaomiMiMo/MiMo-Audio-7B-Instruct -- 2025-09-21
  794. Scaling Self-Supervised Representation Learning for Symbolic Piano Performance -- 2025-09-21
  795. unsloth/Qwen3-Next-80B-A3B-Instruct -- 2025-09-19
  796. Gluon: a GPU programming language based on the same compiler stack as Triton -- 2025-09-19
  797. Tesslate/WEBGEN-OSS-20B -- 2025-09-19
  798. UUIDv47: Store UUIDv7 in DB, emit UUIDv4 outside (SipHash-masked timestamp) -- 2025-09-19
  799. Beyond Prompt-Induced Lies: Investigating LLM Deception on Benign Prompts -- 2025-09-18
  800. LFM2-1.2B safety benchmark -- 2025-09-18
  801. Has anyone successfully gotten Ollama models (or any models) to execute SQL queries through natural language in Openwebui? -- 2025-09-18
  802. ROCm 6.4.3 -> 7.0-rc1 after updating got +13.5% at 2xR9700 -- 2025-09-18
  803. Is it possible for different brand GPUs to work together? -- 2025-09-18
  804. Speculative cascades — A hybrid approach for smarter, faster LLM inference -- 2025-09-17
  805. google/embeddinggemma-300m -- 2025-09-17
  806. Running Qwen-Next (Instruct and Thinking) MLX BF16 with MLX-LM on Macs -- 2025-09-17
  807. From Research to Reality: Feasibility of Gradient Inversion Attacks in Federated Learning -- 2025-09-17
  808. Kwai-Klear/Klear-46B-A2.5B-Instruct -- 2025-09-16
  809. WestZhang/VibeVoice-Large-pt -- 2025-09-16
  810. FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference -- 2025-09-16
  811. Qwen 3 Next Series – Qwen/Qwen3 Next 80B A3B Instruct Detected -- 2025-09-16
  812. Test-time Prompt Intervention -- 2025-09-16
  813. Effecient hot-swappable LoRA variant supported in llama.cpp -- 2025-09-12
  814. Qwen/Qwen3-Next-80B-A3B-Instruct -- 2025-09-12
  815. swiss-ai/Apertus-70B-Instruct-2509 -- 2025-09-12
  816. GRASPED: Graph Anomaly Detection using Autoencoder with Spectral Encoder and Decoder (Full Version) -- 2025-09-12
  817. stepfun-ai/step3 -- 2025-09-12
  818. Exploring State-Space-Model based Language Model in Music Generation -- 2025-09-12
  819. Small Open Models Achieve Near Parity with Large Models in Low Resource Literary Translation at a Fraction of the Cost -- 2025-09-10
  820. Reasoning Vectors: Transferring Chain-of-Thought Capabilities via Task Arithmetic -- 2025-09-10
  821. Tilde AI Releases TildeOpen LLM: An Open-Source Large Language Model with Over 30 Billion Parameters and Support Most European Languages -- 2025-09-09
  822. Show HN: Open-sourcing our text-to-CAD app -- 2025-09-09
  823. bytedance-research/USO -- 2025-09-09
  824. LiquidAI/LFM2-VL-450M -- 2025-09-09
  825. Fantastic pretraining optimizers and where to find them -- 2025-09-08
  826. Qwen/Qwen3-4B-Instruct-2507 -- 2025-09-08
  827. moonshotai/Kimi-K2-Instruct-0905 -- 2025-09-08
  828. Tenstorrent p150a tested against RTX5090, RTX3090, A100, H100 by Russian blogger -- 2025-09-08
  829. unsloth/gemma-3-270m-it-GGUF -- 2025-09-08
  830. YanoljaNEXT-Rosetta: A Collection of Translation Models in Different Sizes -- 2025-09-07
  831. Vulkan back ends, what do you use? -- 2025-09-06
  832. A new OpenAI model? Could this be 5.1 or 5o? What do you think? -- 2025-09-06
  833. VibeVoice RIP? What do you think? -- 2025-09-05
  834. lodestones/Chroma1-HD -- 2025-09-05
  835. Welcome EmbeddingGemma, Google's new efficient embedding model -- 2025-09-05
  836. NousResearch/Hermes-4-405B -- 2025-09-05
  837. I locally benchmarked 41 open-source LLMs across 19 tasks and ranked them -- 2025-09-05
  838. Little SSM (RWKV7 7B) state checkpointing demo. -- 2025-09-04
  839. Need advice on how to get VLLM working with 2xR9700 + 2x7900xtx? -- 2025-09-04
  840. The Hacker's Guide to Building an AI Supercluster -- 2025-09-02
  841. CAD, From Scratch: MakerCAD -- 2025-09-02
  842. PSO-Merging: Merging Models Based on Particle Swarm Optimization -- 2025-09-02
  843. internlm/Intern-S1-mini -- 2025-09-02
  844. stepfun-ai/Step-Audio-2-mini -- 2025-09-02
  845. Claude Memory Lazy Method: The Graduation Path (From 4 Prompts to 1) -- 2025-08-30
  846. bullerwins/Wan2.2-I2V-A14B-GGUF -- 2025-08-28
  847. QuantStack/Qwen-Image-Edit-GGUF -- 2025-08-28
  848. dvlab-research/MGM-Omni -- 2025-08-28
  849. DeepSeek V3.1 dynamic Unsloth GGUFs + chat template fixes -- 2025-08-28
  850. PSA: OpenAI GPT-OSS running slow? Do not set top-k to 0! -- 2025-08-28
  851. Seamlessly bridge LM Studio and OpenWebUI with zero configuration -- 2025-08-28
  852. I built Husk, a native, private, and open-source iOS client for your local models -- 2025-08-28
  853. SQLite-Vector adds support for float16 and bfloat16 (CPU, NEON, AVX2 and SSE2) -- 2025-08-26
  854. zai-org/GLM-4.5 -- 2025-08-26
  855. rednote-hilab/dots.vlm1.inst -- 2025-08-26
  856. black-forest-labs/FLUX.1-Krea-dev -- 2025-08-26
  857. Design Patterns in MCP: Literate Reasoning -- 2025-08-22
  858. FlyMyAI/flymyai-lora-trainer -- 2025-08-22
  859. Zedless: Zed fork focused on privacy and being local-first -- 2025-08-22
  860. Show HN: Rucat – Cat for Prompt Engineers -- 2025-08-22
  861. NEW VERSION: 0.6.23 Has Just Released! - Many fixes and new features, huge changelog -- 2025-08-22
  862. nvidia/Llama-3_3-Nemotron-Super-49B-v1_5 -- 2025-08-21
  863. NVIDIA Nemotron Nano 2 and the Nemotron Pretraining Dataset v1 -- 2025-08-20
  864. deepseek-ai/DeepSeek-V3.1-Base · Hugging Face -- 2025-08-20
  865. mistralai/Devstral-Small-2507 -- 2025-08-20
  866. zai-org/GLM-4.5-Air -- 2025-08-20
  867. microsoft/Phi-4-mini-flash-reasoning -- 2025-08-20
  868. ModelTC/Qwen-Image-Lightning -- 2025-08-20
  869. Fast Type-Aware Linting in Oxlint -- 2025-08-20
  870. From Language to Logic: A Bi-Level Framework for Structured Reasoning -- 2025-08-19
  871. MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE -- 2025-08-19
  872. Distillation Scaling Laws -- 2025-08-19
  873. GPT-5, where does it shine for you? -- 2025-08-19
  874. vidore/colqwen-omni-v0.1 -- 2025-08-19
  875. Tutorial: Open WebUI and llama-swap works great together! Demo of setup, model swapping and activity monitoring. -- 2025-08-18
  876. Concurrency in open-weight/open-source models? -- 2025-08-18
  877. [Editorial] Claude Flow, Alpha 90 release -- 2025-08-17
  878. Optimizing Text gen webui (oobabooga) for MOE models (Qwen3-235b, GLM 4.5) -- 2025-08-17
  879. Chen-zexi/vllm-cli -- 2025-08-17
  880. zyfoxx/subhunter -- 2025-08-17
  881. moonshotai/Kimi-K2-Instruct -- 2025-08-17
  882. KittenML/kitten-tts-nano-0.1 -- 2025-08-17
  883. ilkerzgi/Overlay-Kontext-Dev-LoRA -- 2025-08-17
  884. JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 -- 2025-08-17
  885. Fun with RTX PRO 6000 Blackwell SE -- 2025-08-17
  886. Any tips/Advice for running gpt-oss-120b locally -- 2025-08-17
  887. LLM performance of tiny (<4B) models? -- 2025-08-17
  888. What "big" models can I run with this setup: 5070ti 16GB and 128GB ram, i9-13900k ? -- 2025-08-17
  889. HuggingFaceTB/SmolLM3-3B-Base -- 2025-08-16
  890. mistralai/Voxtral-Small-24B-2507 -- 2025-08-16
  891. TiTan - a tiny model for tags and titles -- 2025-08-16
  892. baidu/ERNIE-4.5-VL-424B-A47B-PT -- 2025-08-16
  893. character-ai/pipelining-sft -- 2025-08-15
  894. Compass-Thinker-7B Technical Report -- 2025-08-15
  895. Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning -- 2025-08-15
  896. [Editorial] GLM-4.5, enterprise use -- 2025-08-13
  897. Best local model with function calling? -- 2025-08-13
  898. agentica-org/DeepSWE-Preview -- 2025-08-13
  899. janhq/Jan-v1-4B-GGUF -- 2025-08-13
  900. Unsloth fixes chat_template (again). gpt-oss-120-high now scores 68.4 on Aider polyglot -- 2025-08-12
  901. GPT-OSS-120B runs on just 8GB VRAM & 64GB+ system RAM -- 2025-08-12
  902. How Attention Sinks Keep Language Models Stable -- 2025-08-12
  903. Mitigate Hallucinations by Fine-tuning gpt-oss-120b with One Example -- 2025-08-10
  904. uncensored gpt-oss-20b, bf16 and mxfp4 both available -- 2025-08-10
  905. LGAI-EXAONE/EXAONE-4.0-1.2B -- 2025-08-10
  906. New Open-Source Text-to-Image Model Just Dropped Qwen-Image (20B MMDiT) by Alibaba! -- 2025-08-10
  907. Experience with GLM-4.5-Air + claude code? -- 2025-08-08
  908. Just when you thought Qwen was done... -- 2025-08-08
  909. Running GPT-OSS-120B at 500 tokens per second on Nvidia GPUs -- 2025-08-08
  910. How the best image generation models work from the inside ? -- 2025-08-08
  911. AIDC-AI/Ovis-U1-3B -- 2025-08-08
  912. google/gemma-3n-E4B -- 2025-08-07
  913. IntervitensInc/pangu-pro-moe-model -- 2025-08-07
  914. Welcome GPT OSS, the new open-source model family from OpenAI! -- 2025-08-07
  915. 0.82 um 105 W diode-pumped thulium-doped all silica fiber laser -- 2025-08-06
  916. ByteDance drops Seed-Prover -- 2025-08-06
  917. naver-hyperclovax/HyperCLOVAX-SEED-Think-14B -- 2025-08-06
  918. [Editorial] a more mature phase of the AI cycle. -- 2025-08-05
  919. glm-4.5-Air appreciation poist - if you have not done so already, give this model a try -- 2025-08-05
  920. How to locally run Grok 4 with 2x AMD 7900 XTX GPUs? (24 GB VRAM x2) -- 2025-08-05
  921. zai-org/GLM-4.5 -- 2025-08-05
  922. unsloth/Qwen3-Coder-480B-A35B-Instruct-GGUF -- 2025-08-05
  923. Learn Software-Defined Radio, GNURadio, RTL-SDR and PlutoSDR with Prof Jason -- 2025-08-05
  924. Waiting on direct MCP integration—dev team, got a roadmap update? -- 2025-08-05
  925. [Help] Figma MCP Tool Execution via HTTP API - Getting 404s, Is External Tool Calling Supported? -- 2025-08-05
  926. Build an AI Shopping Assistant with Gradio MCP Servers -- 2025-08-01
  927. Wan 2.2 T2V,I2V 14B MoE Models -- 2025-07-31
  928. PowerInfer/SmallThinker-21BA3B-Instruct -- 2025-07-31
  929. haykgrigo3/TimeCapsuleLLM -- 2025-07-30
  930. unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF -- 2025-07-30
  931. PhysicsWallahAI/Aryabhata-1.0 -- 2025-07-30
  932. This year’s best open-source models and most cost-effective models -- 2025-07-29
  933. I’m looking for multimodal image input support and uncensored LLM -- 2025-07-29
  934. nvidia/audio-flamingo-3 -- 2025-07-29
  935. mistralai/Voxtral-Mini-3B-2507 -- 2025-07-29
  936. UI/UX benchmark update 7/22: Newest Qwen models added, Qwen3 takes the lead in terms of win rate (though still early) -- 2025-07-28
  937. unsloth/Qwen3-235B-A22B-Thinking-2507-GGUF -- 2025-07-28
  938. zai-org/GLM-4.5 -- 2025-07-28
  939. Tesslate/UIGEN-X-32B-0727 -- 2025-07-28
  940. had to fine-tune qwen since llama sucks at summarizing -- 2025-07-28
  941. orchestre-dev/ccproxy -- 2025-07-28
  942. Guide to PDF security -- 2025-07-28
  943. MetaMask extension bug causes 100s of GBs of extraneous data to be written -- 2025-07-28
  944. Commodore 64 on New FPGA -- 2025-07-28
  945. Running Qwen3 235B-A22B 2507 on a Threadripper 3970X + 3x RTX 3090 Machine at 15 tok/s -- 2025-07-25
  946. The Latest GPT-5 Leaks and Teasers -- 2025-07-25
  947. Qwen3-235B-A22B-Thinking-2507 released! -- 2025-07-25
  948. albozes/shotbuddy -- 2025-07-25
  949. uttam-li/dfs -- 2025-07-25
  950. OmniSVG/OmniSVG -- 2025-07-25
  951. Freezer Monitoring: Because Ice Cream Is a Dish Best Served Cold -- 2025-07-25
  952. Fast LoRA inference for Flux with Diffusers and PEFT -- 2025-07-25
  953. FreeBSD 15's installer to gain option to install a full KDE Plasma desktop -- 2025-07-24
  954. Spanish police arrest five over $542M crypto investment scheme -- 2025-07-24
  955. A Spectrophotometer Jailbreak to Resolve Colorful Disputes -- 2025-07-24
  956. Lucy: A Mobile-Capable 1.7B Reasoning Model That Rivals Jan-Nano -- 2025-07-23
  957. Recommend hardware for my use case? -- 2025-07-20
  958. Best Hardware Setup to Run DeepSeek-V3 670B Locally on $40K–$80K? -- 2025-07-20
  959. e6a5/flow -- 2025-07-20
  960. Improve Your KiCad Productivity With These Considered Shortcut Keys -- 2025-07-20
  961. Semantic chunking using LLMs -- 2025-07-20
  962. Does the OpenWebUi run the sentence transformer models locally? -- 2025-07-20
  963. Dataset for structured (JSON) output? -- 2025-07-19
  964. support for Kimi-K2 has been merged into llama.cpp -- 2025-07-19
  965. t-tech/T-pro-it-2.0 -- 2025-07-19
  966. Support for diffusion models (Dream 7B) has been merged into llama.cpp -- 2025-07-17
  967. Seq vs Seq: the Ettin Suite of Paired Encoders and Decoders -- 2025-07-17
  968. T5Gemma: A new collection of encoder-decoder Gemma models- Google Developers Blog -- 2025-07-17
  969. H-Net: a hierarchical network that replaces tokenization with a dynamic chunking process directly inside the model, automatically discovering and operating over meaningful units of data -- 2025-07-17
  970. An Open-Concept 3D Printer Using Cantilever Arms -- 2025-07-17
  971. Exploring State-Space-Model based Language Model in Music Generation -- 2025-07-16
  972. Diffusion model support in llama.cpp. -- 2025-07-16
  973. GLM-4 MoE incoming -- 2025-07-16
  974. FlexOlmo: Open Language Models for Flexible Data Use | Implications for federated training in the open source community -- 2025-07-15
  975. Tencent/AICGSecEval -- 2025-07-15
  976. microsoft/NextCoder-32B -- 2025-07-15
  977. RekaAI/reka-flash-3.1 -- 2025-07-15
  978. Qwen3-235B-A22B @ 0.7t/s. Hardware or configuration bottleneck? -- 2025-07-15
  979. Building a silent, budget 4-GPU LLM workstation—1×3090 + 3×P40, need advice -- 2025-07-15
  980. Enough resources for light AI workloads? -- 2025-07-15
  981. What can I expect from current amd igpu performance? -- 2025-07-15
  982. Kimi-K2 is a DeepSeek V3 with more experts -- 2025-07-14
  983. HuggingFaceTB/SmolLM3-3B-Base -- 2025-07-14
  984. Replication of Quantum Factorisation Records with an 8-bit Home Computer [pdf] -- 2025-07-14
  985. Why don’t we have a big torrent repo for open-source LLMs? -- 2025-07-12
  986. Local PDF Database searchable with ollama - best setup? -- 2025-07-12
  987. Tinyllama on old Mediatek G80 android device -- 2025-07-12
  988. I used Ollama to build a Cursor for PDFs -- 2025-07-12
  989. Advice on switching to LLM -- 2025-07-12
  990. Building the Hugging Face MCP Server -- 2025-07-11
  991. Support for the upcoming IBM Granite 4.0 has been merged into llama.cpp -- 2025-07-11
  992. support for Falcon-H1 model family has been merged into llama.cpp -- 2025-07-11
  993. [Tool Release] Finetune & Quantize 1–3B LLMs on 8GB RAM using LoFT CLI (TinyLlama + QLoRA + llama.cpp) -- 2025-07-11
  994. Qwen3-8B-BitNet -- 2025-07-11
  995. How are commercial dense models so much faster? -- 2025-07-09
  996. pola-rs/polars -- 2025-07-09
  997. Looking for an upgrade from Meta-Llama-3.1-8B-Instruct-Q4_K_L.gguf, especially for letter parsing. Last time I looked into this was a very long time ago (7 months!) What are the best models nowadays? -- 2025-07-08
  998. Best models by size? -- 2025-07-08
  999. Planning a 7–8B Model Benchmark on 8GB GPU — What Should I Test & Measure? -- 2025-07-08
  1000. Smallest & best OCR model that can read math & code? -- 2025-07-07
  1001. Qwen/WorldPM-72B -- 2025-07-07
  1002. black-forest-labs/FLUX.1-Kontext-dev-onnx -- 2025-07-07
  1003. MCPVerse – An open playground for autonomous agents to publicly chat, react, publish, and exhibit emergent behavior -- 2025-07-06
  1004. I made a free iOS app for people who run LLMs locally. It’s a chatbot that you can use away from home to interact with an LLM that runs locally on your desktop Mac. -- 2025-07-06
  1005. Using local models with Void -- 2025-07-05
  1006. Intel GPU vLLM Docker Compose Bootstrap with Phi-lthy4 on A770 -- 2025-07-05
  1007. Accelerated LLM Inference on AMD Instinct™ GPUs with vLLM 0.9.x and ROCm -- 2025-07-05
  1008. Gemma 3n fully available in the open-source ecosystem! -- 2025-07-05
  1009. 5060ti 16gb or 9060xt 16gb for small llm server -- 2025-07-05
  1010. Qwen3 models in MLX format! -- 2025-07-05
  1011. Best local coding model right now? -- 2025-07-05
  1012. Wallbleed: A Memory Disclosure Vulnerability in the Great Firewall of China -- 2025-07-05
  1013. 0-Pierced Triangles within a Poisson Overlay -- 2025-07-05
  1014. 1000 days of lowest frequency emission from the low-luminosity GRB 171205A -- 2025-07-05
  1015. Found a Web3 LLM That Actually Gets DeFi Right -- 2025-07-03
  1016. apple/DiffuCoder-7B-cpGRPO -- 2025-07-03
  1017. Ollama - Windows 11 > LXC Docker - Openwebui = constant BSOD with RTX 5090 Ventus on driver 576.80 -- 2025-07-02
  1018. Running Open WebUI with NVIDIA GPU Support? -- 2025-07-02
  1019. Cursor 1.0 -- 2025-06-30
  1020. Help me design a robust on-prem Llama 3 70B infrastructure for 30 users – Complete hardware/software list wanted -- 2025-06-30
  1021. Jan-nano, a 4B model that can outperform 671B on MCP -- 2025-06-30
  1022. Models that are good and fast at Long Document Processing -- 2025-06-30
  1023. I am making an AI batteries included Web Framework (like Django but for AI) -- 2025-06-30
  1024. [New Features & Better] Tabulens: A Vision-LLM Powered PDF Table Extractor -- 2025-06-30
  1025. I tested 10 LLMs locally on my MacBook Air M1 (8GB RAM!) – Here's what actually works- -- 2025-06-30
  1026. Chatbot without ChatGPT -- 2025-06-30
  1027. What's the best way to save and manage different text files for the models to reference? PRD, cursor rules, tech stack, design reference, etc? -- 2025-06-30
  1028. Bzip2 crate switches from C to 100% Rust -- 2025-06-30
  1029. Litestream: Revamped -- 2025-06-30
  1030. Stop using REST for state synchronization (2024) -- 2025-06-30
  1031. Announcing `mcp-protocol-sdk`: A New Enterprise grade Rust SDK for AI Tool Calling (Model Context Protocol) -- 2025-06-30
  1032. THU-KEG/LongWriter-Zero-32B -- 2025-06-30
  1033. 100 Gbps Indoor Access and 4.8 Gbps Outdoor Point-to-Point LiFi Transmission Systems using Laser-based Light Sources -- 2025-06-30
  1034. (0,4) brane box models -- 2025-06-30
  1035. Built memX: a shared memory backend for LLM agents (demo + open-source code) -- 2025-06-29
  1036. Automatically Evaluating AI Coding Assistants with Each Git Commit (Open Source) -- 2025-06-29
  1037. Secure Minions: private collaboration between Ollama and frontier models -- 2025-06-29
  1038. Privacy implications of sending data to OpenRouter -- 2025-06-29
  1039. Exploring Practical Uses for Small Language Models (e.g., Microsoft Phi) -- 2025-06-29
  1040. LLM with OCR capabilities -- 2025-06-29
  1041. How to create a speech recognition model from scratch -- 2025-06-29
  1042. Arch 0.3.0 is out - I added support for the Claude family of LLMs in the proxy server framework for agents 🚀 -- 2025-06-29
  1043. Gemini Cli MCP Agent just released ! -- 2025-06-29
  1044. Freeplane xml mind maps locally: only Qwen3 and Phi4 Reasoning Plus can create them in one shot? -- 2025-06-29
  1045. Reinforcement Pre-Training -- 2025-06-29
  1046. unsloth/gemma-3n-E4B-it-GGUF -- 2025-06-29
  1047. chandar-lab/NeoBERT -- 2025-06-29
  1048. unsloth/Mistral-Small-3.2-24B-Instruct-2506-GGUF -- 2025-06-26
  1049. jinaai/jina-embeddings-v4 -- 2025-06-26
  1050. Intelligent-Internet/II-Medical-8B-1706 -- 2025-06-26
  1051. 0-concordance of knotted surfaces and Alexander ideals -- 2025-06-26
  1052. 100% of the zeros of the Riemann zeta-function are on the critical line -- 2025-06-25
  1053. 100% of odd hyperelliptic Jacobians have no rational points of small height -- 2025-06-25
  1054. deepseek-ai/DualPipe -- 2025-06-23
  1055. 100 Particles Quantum Heat Engine: Exploring the Impact of Criticality on Efficiency -- 2025-06-23
  1056. 0-Auslander correspondence -- 2025-06-23
  1057. nvidia/Cosmos-Predict2-2B-Text2Image -- 2025-06-22
  1058. 1000-10,000 M$_\odot$ Primordial Stars Created the Nitrogen Excess in the Galaxy GS 3073 at $z = 5.55$ -- 2025-06-21
  1059. $0^+$ to $2^+$ neutrinoless double-$β$ decay of $^{76}$Ge, $^{82}$Se, $^{130}$Te and $^{136}$Xe in the microscopic interacting boson model} -- 2025-06-21
  1060. resemble-ai/chatterbox -- 2025-06-21
  1061. google/magenta-realtime -- 2025-06-21
  1062. 0-1 laws for pattern occurrences in phylogenetic trees and networks -- 2025-06-20
  1063. meta-llama/Llama-3.1-8B-Instruct -- 2025-06-19
  1064. MiniMaxAI/MiniMax-M1-80k -- 2025-06-19
  1065. Demo Video of AutoBE, Backend Vibe Coding Agent Achieving 100% Compilation Success (Open Source) -- 2025-06-19
  1066. How to set up local llms on a 6700 xt -- 2025-06-19
  1067. Jetson Orin AGX 32gb -- 2025-06-19
  1068. AMD GPU support -- 2025-06-19
  1069. Much lower performance for Mistral-Small 24B on RTX 3090 and from deepinfra API -- 2025-06-19
  1070. Extract Website Information -- 2025-06-19
  1071. Looking for a verified copy of big-lama.ckpt (181MB) used in the original LaMa inpainting model trained on Places2. -- 2025-06-19
  1072. Is it true that all tools like Cline/Copilot Agent/Roo Code/Windsurf/Claude Code/Cursor are roughly the same thing? -- 2025-06-19
  1073. SkyRoof: New Ham Satellite Tracking and SDR Receiver Software -- 2025-06-19
  1074. MiniMaxAI/MiniMax-M1-40k -- 2025-06-18
  1075. Qwen/Qwen3-Reranker-4B -- 2025-06-18
  1076. ckanthony/openapi-mcp -- 2025-06-18
  1077. Ta0ing/MCP-SecurityTools -- 2025-06-18
  1078. 100% Elimination of Hallucinations on RAGTruth for GPT-4 and GPT-3.5 Turbo -- 2025-06-17
  1079. unsloth/Magistral-Small-2506-GGUF -- 2025-06-12
  1080. mistralai/Magistral-Small-2506 -- 2025-06-12
  1081. rednote-hilab/dots.llm1.base -- 2025-06-12
  1082. [Tool] rvn-convert: OSS Rust-based SafeTensors to GGUF v3 converter (single-shard, fast, no Python) -- 2025-06-12
  1083. GuidedQuant: Boost LLM layer-wise PTQ methods using the end loss guidance (Qwen3, Gemma3, Llama3.3 / 2~4bit Quantization) -- 2025-06-12
  1084. I built a memory MCP that understands you (so Sam Altman can't). -- 2025-06-12
  1085. Built an open source desktop app to easily play with local LLMs and MCP -- 2025-06-12
  1086. mtmd : support Qwen 2.5 Omni (input audio+vision, no audio output) by ngxson · Pull Request #13784 · ggml-org/llama.cpp -- 2025-06-12
  1087. i got tired of the errors, so automated debugging using Ollama -- 2025-06-12
  1088. Ablating Gemma 3 27B variants with synthetic data from Sonnet 4 (Few-shot vs LoRA) -- 2025-06-12
  1089. The LLM Gateway gets a major upgrade: becomes a data-plane for Agents. -- 2025-06-12
  1090. Introducing stronger dependencies on systemd -- 2025-06-12
  1091. How we decreased GitLab repo backup times from 48 hours to 41 minutes -- 2025-06-12
  1092. The Quest for 100k - LLAMA.CPP Setting for a Noobie -- 2025-06-12
  1093. News publishers call Google's AI Mode 'theft' -- 2025-06-11
  1094. Qwen/Qwen3-Embedding-4B -- 2025-06-10
  1095. The Unreliability of LLMs and What Lies Ahead -- 2025-06-10
  1096. AnythingLLM RAG with Gemma 3:12b & BGE-m3-F16: LM Studio vs. Ollama Embedding Discrepancies - Same GGUF, Different Results? -- 2025-06-09
  1097. What Models for C/C++? -- 2025-06-09
  1098. Help with guardrails ai and local ollama model -- 2025-06-09
  1099. Setup Recommendation for University (H200 vs RTX 6000 Pro) -- 2025-06-09
  1100. LlamaFirewall: framework open source per rilevare e mitigare i rischi per la sicurezza incentrati sull'intelligenza artificiale - Help Net Security -- 2025-06-09
  1101. Are autoencoders really need for anomaly detection in time series? -- 2025-06-09
  1102. Backdoored malware repos traced to single GitHub user -- 2025-06-09
  1103. Improper Access Control Allows All Users to View Private Content. Am I doing it wrong ? -- 2025-06-09
  1104. Agno Now Supports Dual Model Output (Reasoning + Structure) -- 2025-06-09
  1105. 100 Gbps Quantum-safe IPsec VPN Tunnels over 46 km Deployed Fiber -- 2025-06-09
  1106. Qwen/Qwen3-Embedding-0.6B -- 2025-06-08
  1107. Qwen/Qwen3-Embedding-8B -- 2025-06-06
  1108. nvidia/Llama-3.1-Nemotron-Nano-VL-8B-V1 -- 2025-06-05
  1109. hexgrad/Kokoro-82M -- 2025-06-05
  1110. Qwen/Qwen3-Embedding-0.6B-GGUF -- 2025-06-05
  1111. Gradient-Based Program Repair: Fixing Bugs in Continuous Program Spaces -- 2025-06-05
  1112. 0-dimensional Homology Preserving Dimensionality Reduction with TopoMap -- 2025-06-05
  1113. NousResearch/atropos -- 2025-06-04
  1114. 100,000 Podcasts: A Spoken English Document Corpus -- 2025-06-03
  1115. nvidia/Nemotron-Research-Reasoning-Qwen-1.5B -- 2025-06-03
  1116. Intelligent-Internet/II-Medical-8B -- 2025-06-03
  1117. Gen-Verse/MMaDA -- 2025-06-01
  1118. osmosis-ai/Osmosis-Structure-0.6B -- 2025-06-01
  1119. simplescaling/s1 -- 2025-05-31
  1120. FractalAIResearch/Fathom-R1-14B -- 2025-05-31
  1121. unsloth/DeepSeek-R1-0528-Qwen3-8B-GGUF -- 2025-05-31
  1122. deepseek-ai/DeepSeek-R1-0528-Qwen3-8B -- 2025-05-31
  1123. I made Model Version Control Protocol for AI agents -- 2025-05-31
  1124. AI Baby Monitor – fully local Video-LLM nanny (beeps when safety rules are violated) -- 2025-05-31
  1125. LMStudio - llama.cpp - vLLM -- 2025-05-31
  1126. Built an ADK Agent that finds Jobs based on your Resume -- 2025-05-31
  1127. Should I resize the image before sending it to Qwen VL 7B? Would it give better results? -- 2025-05-31
  1128. How to start a LLM project? -- 2025-05-31
  1129. Beware of Fast-Math -- 2025-05-31
  1130. facebook/OMol25 -- 2025-05-30
  1131. EdinburghNLP/MMLongBench -- 2025-05-29
  1132. Show HN: Model2vec-Rs – Fast Static Text Embeddings in Rust -- 2025-05-29
  1133. unsloth/DeepSeek-R1-0528-GGUF -- 2025-05-29
  1134. QuantStack/Wan2.1-VACE-14B-GGUF -- 2025-05-29
  1135. nvidia/Llama-3.1-Nemotron-Nano-4B-v1.1 -- 2025-05-28
  1136. Tongyi-Zhiwen/QwenLong-L1-32B -- 2025-05-28
  1137. PKU-DS-LAB/FairyR1-32B -- 2025-05-28
  1138. google/medgemma-4b-pt -- 2025-05-28