Reasoning Models

Chain of thought, thinking models, math and logic reasoning

455 articles across 142 editions

Articles

  1. Amazon Unveils Strands Decider 2B: Free, Fast, Open-Source Decision Model -- 2026-10-02
  2. Mapika/decider on GitHub -- 2026-10-02
  3. nokia-applied-research/AnyJev -- 2026-10-02
  4. llama, server: add /v1/systemone API (models: laya, julia-1, lev, openjev, kev) by ngxson · Pull Request #29818 · ggml-org/llama.cpp -- 2026-10-02
  5. Trained locally: ultra-fast 0.8B/2B System 1 decision models that match Jev on benchmarks and Doom, ~30 ms per decision (open weights) -- 2026-09-30
  6. Jeeves. Reasoning improves Jev-like decision models -- 2026-09-30
  7. BAAI/AREX-2 - 27B - Agent model based on Qwen3.8 27B -- 2026-09-30
  8. 400+ LLM agents living in a 2004-era MMO server, all local on Qwen3-4B -- 2026-09-30
  9. modified qwen 3.8 27b modifies windows credential dumper to bypass EDR detection -- 2026-09-29
  10. NVIDIA OpenShell: safe, private kernel-level sandbox runtime for autonomous AI agents -- 2026-09-29
  11. Jev is the FIRST of a Whole New Class of AI Models (Here's How to Actually Use It) -- 2026-09-29
  12. Run decision models on vLLM and Red Hat AI using DiffusionGemma -- 2026-09-29
  13. RedHatAI/diffusiongemma-26B-A4B-it-FP8-dynamic on Hugging Face -- 2026-09-29
  14. You can use any LLM just like JEV -- 2026-09-22
  15. [Editorial] mizorewww/laya-mlx -- 2026-09-22
  16. Laya project site (laya.convaiinnovations.com) -- 2026-09-21
  17. Laya: typed decision models (convaiinnovations/laya on Hugging Face) -- 2026-09-21
  18. mizorewww/laya-mlx -- 2026-09-21
  19. OpenJev: open-source Jev implementation (razorback16/openjev) -- 2026-09-21
  20. OpenJev -- 2026-09-21
  21. Jev Reproductions Tracker (Hugging Face Space by multimodalart) -- 2026-09-21
  22. [Editorial] pd-bridge -- 2026-09-17
  23. 3k$ 128GB VRAM + 256GB RAM DDR4 Server -- 2026-09-17
  24. Qwen 3.8 27B UD-IQ4_XS even faster on 16GB CUDA -- 2026-09-17
  25. New tensor type layouts for my GGUF uploads -- 2026-09-17
  26. Nvidia announces native GPU programming in Rust -- 2026-09-17
  27. Spomin - Live KV cache compaction (Experimental for Qwen) -- 2026-09-15
  28. I made a way to migrate between embedding models without re-embedding your entire corpus -- 2026-09-15
  29. Nemoryn — open-source memory backend for Open WebUI (looking for testers) -- 2026-09-15
  30. NInfer vs llama.cpp vs vLLM: quality + speed comparison for Qwen3.8-27B NVFP4 on RTX 5090 -- 2026-09-07
  31. Qwen3.8 Flash AP Quants -- 2026-09-07
  32. Introducing Quartermaster, an open source local AI platform designed for ease of use that does not sacrifice customizability -- 2026-09-07
  33. [Editorial] -- 2026-09-07
  34. Introducing K2 Horizon: Frontier Performance, Radically Open -- 2026-09-07
  35. Microsoft VibeVoice-ASR-Streaming Released -- 2026-09-07
  36. Editor's video pick (YouTube gasgivVCl2U) -- 2026-09-02
  37. [Editorial] robertelee78/hf2q -- 2026-08-21
  38. [Editorial] hf2q.us -- 2026-08-21
  39. I pushed Qwen3.8-27B limits again... Dflash2 - 134 tps on a RTX 3090 -- 2026-08-20
  40. NInfer RTX 4090 for Qwen 3.8 27B update - up to 250-350K tokens context in VRAM -- 2026-08-20
  41. llama.cpp adaptive MTP PR#27210 -- 2026-08-20
  42. Qwen 3.8 2.4T at 288k tokens/s on Nvidia GB300 NVL72 -- 2026-08-20
  43. Lightricks/LTX-2.3 -- 2026-08-20
  44. Token Budget Saturation and Mechanistic Early Detection of Reasoning Non-Convergence in Chain-of-Thought Models -- 2026-08-18
  45. Decoding Claude's DNA: Comparing System Prompts Across Fable 5, Opus 5/4.8/4.6, Sonnet 5 & Haiku 4.5 -- 2026-08-18
  46. Adaptive and Explicit safe: Triggering Latent Safety Awareness in Large Reasoning Models -- 2026-08-14
  47. [Editorial] LinkedIn Post (gPEg3ZXB) -- 2026-08-14
  48. WorldClaw Agentic 3D open-world generation at scale -- 2026-08-14
  49. [Editorial] Reuters Video Report -- 2026-08-14
  50. What sort of maths are LLMs good at? -- 2026-08-13
  51. Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence -- 2026-08-13
  52. Efficient LLM-Generated Shuttling Compilers for Complex Trapped-Ion Architectures -- 2026-08-13
  53. Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning -- 2026-08-12
  54. [2606.05682] Beyond Output Matching: Preserving Internal Geometry in NVFP4 LLM Distillation -- 2026-08-12
  55. torchtune: PyTorch native post-training library -- 2026-08-12
  56. ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents -- 2026-08-06
  57. I built an MIT-licensed MCP server so my agent can do PDF work without the file ever leaving my machine -- 2026-08-06
  58. aws/context-ontology-accelerator -- 2026-08-06
  59. Smaller, faster, safer: running Kimi and GLM at scale on Cloudflare -- 2026-08-04
  60. Huawei open-sources openPangu-2.0-Pro: 505B-A18B MoE on Ascend -- 2026-08-04
  61. NousResearch ships Hermes 0.20 agent framework -- 2026-08-04
  62. [Editorial] WASTE – Stream Kimi K3 from NVMe -- 2026-08-03
  63. WASTE: Run Kimi K3 beyond available RAM by streaming from NVMe -- 2026-08-03
  64. [Editorial] Kimi K3 Implemented in C -- 2026-08-03
  65. [Editorial] DeltaFin -- 2026-08-03
  66. Why we write our own C and C++ inference engines -- 2026-08-03
  67. mistral.rs v0.9.0: up to 1.8x faster CPU decode than llama.cpp on x86 and ARM -- 2026-07-14
  68. Implemented Qwen3.5's hybrid Mamba-Transformer architecture from scratch in Rust -- 2026-07-14
  69. llama.cpp Agentic Workflows Context Checkpoints Fix -- 2026-07-14
  70. DyCon: Dynamic Reasoning Control via Evolving Difficulty Modeling -- 2026-07-09
  71. Particle Scattering Sampler for llama.cpp -- 2026-07-06
  72. [Editorial] NVIDIA Just Open-Sourced a Polyamorous AI -- 2026-07-06
  73. [Paper] Multi-Resolution Flow Matching: Training-Free Diffusion Acceleration via Staged Sampling -- 2026-07-06
  74. 55 LLMs blind-grade each other: 22K judgments reveal systematic same-family bias -- 2026-06-30
  75. If LLMs Have Human-Like Attributes, Then So Does Age of Empires II -- 2026-06-30
  76. The Doorman's Fallacy in action -- 2026-06-30
  77. LFM2.5 230M running in-browser at 1,400 tok/s using custom WebGPU kernels -- 2026-06-26
  78. Got GLM-5.2 + MTP speculative decode running on 4x DGX Spark (GB10) — and the build piece the public recipe is missing -- 2026-06-26
  79. Run a vLLM Server on HF Jobs in One Command -- 2026-06-26
  80. Running Sonnet 4.6 on every Instagram DM for a 7-location restaurant. 97% cache hit is the only reason it's affordable -- 2026-06-26
  81. [Editorial] rupixel -- 2026-06-26
  82. OpenDeepThink: Parallel Reasoning via Bradley-Terry Aggregation -- 2026-06-19
  83. Beyond LoRA: Can you beat the most popular fine-tuning technique? -- 2026-06-19
  84. [Editorial] Can LLMs Be Computers? -- 2026-06-15
  85. [Editorial] Reasoning Language Models — RLM and MinRLM -- 2026-06-15
  86. [Editorial] Research Paper — AI/ML Methods -- 2026-06-15
  87. [Editorial] Steve Yegge on Services and Complexity -- 2026-06-15
  88. Holo3.1: Fast & Local Computer Use Agents -- 2026-06-03
  89. Fused MoE dispatch kernel in pure Triton: 89-131% of Megablocks, runs on AMD with zero code changes -- 2026-06-03
  90. ReAligned-Qwen3.5 Release -- 2026-06-03
  91. KANX: A production-ready Kolmogorov-Arnold Network library -- 2026-06-03
  92. SWE-rebench Leaderboard (March, April and May 2026): GPT-5.5, Opus 4.7, Cursor (Composer 2.5), Kimi K2.6 and More -- 2026-05-28
  93. The frontier reasoning race is starting to look like a crowded subway station -- 2026-05-28
  94. MiMo-V2.5-coder -- 2026-05-28
  95. An OpenAI model has disproved a central conjecture in discrete geometry -- 2026-05-21
  96. I've joined Anthropic -- 2026-05-21
  97. Need a second pair of eyes, this Qwen3.6 27B quant recipe consistently thinks less and is correct -- 2026-05-20
  98. llama: avoid copying logits during prompt decode in MTP by am17an - Pull Request #23198 - ggml-org/llama.cpp -- 2026-05-20
  99. Extension idea: llama-server with custom samplers -- 2026-05-20
  100. Simpler self hosted alt to Open WebUI -- 2026-05-20
  101. EMO: Pretraining mixture of experts for emergent modularity -- 2026-05-13
  102. [Editorial] -- 2026-05-13
  103. 2.5x Faster Inference with Qwen 3.6 27B Using MTP — Complete Hardware Guide -- 2026-05-08
  104. Atlas Is Now Open Source — Pure Rust+CUDA Inference Engine for Blackwell -- 2026-05-08
  105. antirez/ds4 — DeepSeek 4 Flash Local Inference Engine for Metal -- 2026-05-08
  106. [Editorial] dflash — Flash Inference Tool -- 2026-05-08
  107. Heretic 1.3: Reproducible Abliterated Models, Integrated Benchmarking, Reduced VRAM -- 2026-05-08
  108. Ryzen AI Max+ 495 (Gorgon Halo) with 192GB VRAM! -- 2026-05-06
  109. noonghunna/club-3090 — Community LLM serving recipes for RTX 3090 -- 2026-05-06
  110. Mistral Medium 3.5 128B and Qwen 3.5 122B A10B on 4x RTX 3080 20GB -- 2026-05-06
  111. Decoupled Attention from Weights - Gemma 4 26B -- 2026-05-06
  112. Writing an LLM Compiler from Scratch: PyTorch to CUDA -- 2026-05-04
  113. Karpathy's MicroGPT Running at 50,000 Tokens/Second on an FPGA -- 2026-05-04
  114. Hipfire: Full AMD Architecture Validation Across RDNA 1–4, Strix Halo, and BC250 -- 2026-05-04
  115. Qwen 3.6-35B KV Cache Benchmark: f16 vs q8_0 vs turbo3 vs turbo4 from 0 to 1M Context -- 2026-05-04
  116. llama.cpp DeepSeek v4 Flash experimental inference -- 2026-05-01
  117. llama.cpp benchmark native vs. non native NVFP4 on Blackwell — summary -- 2026-05-01
  118. Speculative decoding with Gemma-4-31B + Gemma-4-E2B enables 120-200 tok/s -- 2026-05-01
  119. The 4B class of 2026 (benchmark) -- 2026-05-01
  120. Study: 2x+ coding performance of 7B model without touching the coding agent -- 2026-05-01
  121. [Editorial] RuVLLM ESP32 v0.3.0-rc2 — LLM Inference on Microcontrollers -- 2026-05-01
  122. Zyphra/ZUNA — New Model on Hugging Face -- 2026-05-01
  123. [Editorial] Schmidhuber: The World Model Boom -- 2026-04-20
  124. [Editorial] Schmidhuber: World Models, Planning & Curiosity (1990 Origins) -- 2026-04-20
  125. [Editorial] The Ontology Problem (Technical) — Kurt Cagle -- 2026-04-20
  126. All elementary functions from a single binary operator -- 2026-04-14
  127. Falcon Perception -- 2026-04-02
  128. TRL v1.0: Post-Training Library Built to Move with the Field -- 2026-04-02
  129. lucas-maes/le-wm -- 2026-04-02
  130. Toward explaining why traditional ablation/abliteration works -- 2026-04-02
  131. The Cognitive Dark Forest -- 2026-03-30
  132. $500K/year pharmacovigilance platform replicated in a weekend with Claude Code -- 2026-03-30
  133. Google Stitch is insane -- 2026-03-30
  134. [Editorial] Time Traveled from Frontier AI -- 2026-03-30
  135. [Editorial] Ambient Intelligence -- 2026-03-30
  136. Controllable Reasoning Models Are Private Thinkers -- 2026-03-26
  137. Nvidia built a silent opinion engine into NemotronH to gaslight you and they're not the only ones doing it -- 2026-03-26
  138. Gemini thoughts turned very violent -- 2026-03-26
  139. Towards a Neural Debugger for Python -- 2026-03-26
  140. Mathematics behind extreme quantization of Microsoft's BitNet -- 2026-03-26
  141. A.T.L.A.S - Adaptive Test-time Learning and Autonomous Specialization -- 2026-03-26
  142. [Editorial] -- 2026-03-25
  143. [Editorial] -- 2026-03-25
  144. [Editorial] -- 2026-03-25
  145. [Editorial] -- 2026-03-25
  146. [Editorial] -- 2026-03-25
  147. [Editorial] arxiv:2603.15371 -- 2026-03-23
  148. [Editorial] The Open World / Closed World Conundrum -- 2026-03-16
  149. [Editorial] W3C Context Graph Community Group -- 2026-03-16
  150. [Editorial] LLM vs LRM -- 2026-03-10
  151. To everyone using still ollama/lm-studio... llama-swap is the real deal -- 2026-03-10
  152. [Editorial] Markov Chains -- 2026-03-10
  153. [Editorial] Speeding One Cog Breaks the Machine -- 2026-03-10
  154. [Editorial] KatanaLarp -- 2026-03-07
  155. [Editorial] Research Paper -- 2026-03-07
  156. [Editorial] The Builders PRD -- 2026-03-07
  157. [Editorial] Our Design Docs Write Themselves -- 2026-03-05
  158. [Editorial] Claude Code: Brilliant Until the Repo... -- 2026-03-05
  159. [Editorial] Context Is King, But Bad Context Is Poison -- 2026-03-05
  160. [Editorial] Semantic Anchors for LLM Coding -- 2026-03-05
  161. Turbo Connection: Reasoning as Information Flow from Higher to Lower Layers -- 2026-02-25
  162. O(1) Inference and Causal Monoid State Compression in Spartacus-1B -- 2026-02-25
  163. Sink-Aware Pruning for Diffusion Language Models -- 2026-02-25
  164. Consistency of Large Reasoning Models Under Multi-Turn Attacks -- 2026-02-16
  165. [Editorial] AI Testing and Quality Engineering -- 2026-02-16
  166. [Editorial] https://github.com/GMaN1911/claude-cognitive -- 2026-01-02
  167. SA-RAG: Using spreading activation to improve multi-hop retrieval in RAG systems -- 2026-01-02
  168. Is there a way to see what is trashing my context? -- 2026-01-02
  169. A zero-setup agent that benchmarks multiple open / closed source LLMs on your specific problem / data -- 2026-01-02
  170. What is a good model for assisting with patching source code? -- 2026-01-02
  171. Just got an RTX Pro 6000 - need recommendations for processing a massive dataset with instruction following -- 2026-01-02
  172. MiniMaxAI/MiniMax-M2.1 -- 2026-01-02
  173. [Editorial] https://www.linkedin.com/posts/stuart-winter-tear_realist-and-pluralist-conceptions-of-intelligence-activity-7397231918871703554-FmSP -- 2025-12-11
  174. VulnLLM-R: Specialized Reasoning LLM with Agent Scaffold for Vulnerability Detection -- 2025-12-11
  175. Nanbeige4-3B: Lightweight with strong reasoning capabilities -- 2025-12-10
  176. mistralai/Devstral-2-123B-Instruct-2512 -- 2025-12-10
  177. MDAR: A Multi-scene Dynamic Audio Reasoning Benchmark -- 2025-12-04
  178. PrimeIntellect/INTELLECT-3 -- 2025-12-04
  179. cerebras/MiniMax-M2-REAP-162B-A10B -- 2025-12-04
  180. 62-day fixed-prompt probe on Grok-4: strong semantic attractors, thematic inversion, and refusal onset (1,242 samples, fully public) -- 2025-12-03
  181. I built an open-source "Passport" for Claude Agents (MCP) so they can cryptographically sign their own actions -- 2025-12-01
  182. Implemented Anthropic's Programmatic Tool Calling with Langchain so you can use it with any models and tune it for your own use case -- 2025-12-01
  183. CodeModeToon -- 2025-12-01
  184. WeiboAI/VibeThinker-1.5B -- 2025-11-28
  185. [Editorial] https://ai.google.dev/gemini-api/docs/prompting-strategies#agentic-si-template -- 2025-11-28
  186. An explainer blog on attention, KV-caching, continuous batching -- 2025-11-28
  187. I built an open-source CLI that generates context.json bundles for React/TypeScript projects -- 2025-11-28
  188. GraphLite: An Embeddable Graph Database with ISO Graph Query Language Support -- 2025-11-26
  189. allenai/Olmo-3-32B-Think -- 2025-11-25
  190. tencent/HunyuanOCR -- 2025-11-25
  191. peteromallet/Qwen-Image-Edit-InScene -- 2025-11-25
  192. [Editorial] https://www.linkedin.com/posts/stuart-winter-tear_decision-making-amid-information-based-threats-activity-7396539314815533056-4pVx -- 2025-11-18
  193. [Editorial] https://gist.github.com/ruvnet/d6d2739400943037443b78c3ef86d8a5 -- 2025-11-18
  194. [Editorial] https://github.com/mrwadams/stride-gpt/blob/master/docs/operationalization-guide.md -- 2025-11-18
  195. janhq/Jan-v2-VL-high -- 2025-11-18
  196. Loong: Synthesize Long Chain-of-Thoughts at Scale through Verifiers -- 2025-11-18
  197. [Editorial] https://arxiv.org/pdf/2506.21734 -- 2025-11-11
  198. OPERA: A Reinforcement Learning--Enhanced Orchestrated Planner-Executor Architecture for Reasoning-Oriented Multi-Hop Retrieval -- 2025-11-11
  199. [Editorial] https://www.linkedin.com/posts/andriyburkov_this-paper-shows-a-27-million-parameter-model-activity-7393432619365052416-SFLO -- 2025-11-10
  200. Trajectory Distillation for Foundation Models -- 2025-11-10
  201. sail-sg/Precision-RL -- 2025-11-10
  202. inclusionAI/LLaDA2.0-flash-preview -- 2025-11-10
  203. AI Agents Reasoning Collapse Imminent (CMU, Berkeley) -- 2025-11-02
  204. Natural Language Programming: Run Natural Language as Script -- 2025-11-02
  205. Claude Code is a Beast – Tips from 6 Months of Hardcore Use -- 2025-11-02
  206. Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs -- 2025-10-31
  207. [Editorial] https://www.linkedin.com/posts/anthony-alcaraz-b80763155_ai-agents-cant-reason-without-semantic-structure-activity-7389222435906244608-QMMc -- 2025-10-31
  208. Raezil/lattice-agent -- 2025-10-31
  209. driaforall/mem-agent -- 2025-10-31
  210. Memp: Exploring Agent Procedural Memory -- 2025-10-31
  211. I built the HuggingChat Omni Router 🥳 🎈 -- 2025-10-28
  212. Claude Code 2.0.27 -- 2025-10-28
  213. ltjed/freephdlabor -- 2025-10-28
  214. Show HN: Whatdidido – CLI to summarize your work from Jira/Linear -- 2025-10-28
  215. Reasoning should be thought of as a drawback, not a feature -- 2025-10-21
  216. inclusionAI/Ring-1T-preview -- 2025-10-21
  217. Learning Lifted Action Models From Traces of Incomplete Actions and States -- 2025-10-20
  218. I got tired of OpenAI dependency. Built a multi-LLM control center instead. -- 2025-10-19
  219. Turn ChatGPT into a real-time meeting assistant (via MCP + Apps SDK) -- 2025-10-19
  220. Claude Code taking a coffee break 🤔 -- 2025-10-19
  221. Show HN: Cmux – Coding Agent Multiplexer -- 2025-10-19
  222. [Editorial] Sqlite vector -- 2025-10-17
  223. Meta Superintelligence group publishes paper on new RAG technique -- 2025-10-17
  224. [Editorial] ReasoningBank is a self-learning, local-first memory system -- 2025-10-16
  225. [Editorial] ReasoningBank is a self-learning, local-first memory system -- 2025-10-16
  226. I tested if tiny LLMs can self-improve through memory: Qwen3-1.7B gained +8% accuracy on MATH problems -- 2025-10-16
  227. Tested 9 RAG query transformation techniques – HydE is absurdly underrated -- 2025-10-16
  228. GPT-OSS from Scratch on AMD GPUs -- 2025-10-11
  229. How do I compare cost per token for serverless vs provisioned hardware? -- 2025-10-11
  230. OpenAI is good at deals -- 2025-10-11
  231. meituan-longcat/LongCat-Flash-Chat -- 2025-10-11
  232. adb1274/batchi -- 2025-10-11
  233. What are the best models for legal work in Oct 2025? -- 2025-10-07
  234. [Update] FamilyBench: New models tested - Claude Sonnet 4.5 takes 2nd place, Qwen 3 Next breaks 70%, new Kimi weirdly below the old version, same for GLM 4.6 -- 2025-10-07
  235. princeton-pli/RLMT -- 2025-10-07
  236. TGPO: Tree-Guided Preference Optimization for Robust Web Agent Reinforcement Learning -- 2025-10-07
  237. DSpAST: Disentangled Representations for Spatial Audio Reasoning with Large Language Models -- 2025-10-05
  238. swiss-ai/Apertus-8B-2509 -- 2025-10-04
  239. Qwen3-Omni thinking model running on local H100 (major leap over 2.5) -- 2025-09-30
  240. Seeking Advice: Best Model + Framework for Max Tokens/sec on Dual L40S (Testing Rig) -- 2025-09-30
  241. For local models, has anyone benchmarked tool calling protocols performance? -- 2025-09-30
  242. A step by step guide on how to build a LLM from scratch -- 2025-09-28
  243. YannQi/R-4B -- 2025-09-28
  244. New Agent benchmark from Meta Super Intelligence Lab and Hugging Face -- 2025-09-27
  245. evalops/dspy-micro-agent -- 2025-09-27
  246. nvidia/NVIDIA-Nemotron-Nano-9B-v2 -- 2025-09-27
  247. inclusionAI/Ling-flash-2.0 -- 2025-09-27
  248. Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model -- 2025-09-26
  249. CogniSQL-R1-Zero: Lightweight Reinforced Reasoning for Efficient SQL Generation -- 2025-09-26
  250. A1: Asynchronous Test-Time Scaling via Conformal Prediction -- 2025-09-25
  251. DeepLink-org/DeepTrace -- 2025-09-25
  252. GLM 4.5 Air Template Breaking llamacpp Prompt Caching -- 2025-09-25
  253. Tracking prompt evolution for RAG systems - anyone else doing this? -- 2025-09-25
  254. MAESTRO v0.1.6 Update: Better support for models that struggle with JSON mode (DeepSeek, Kimi K2, etc.) -- 2025-09-25
  255. Dead-simple example code for Ollama function calling. -- 2025-09-25
  256. nvidia/NVIDIA-Nemotron-Nano-12B-v2 -- 2025-09-23
  257. DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning -- 2025-09-22
  258. Atom-Searcher: Enhancing Agentic Deep Research via Fine-Grained Atomic Thought Reward -- 2025-09-22
  259. support for the upcoming Olmo3 model has been merged into llama.cpp -- 2025-09-21
  260. Running Nvidia CUDA Pytorch/vLLM projects and pipelines on AMD with no modifications -- 2025-09-21
  261. A Quick Look At The AMD Instinct MI355X With ROCm 7.0 -- 2025-09-21
  262. Uncensored AI model for from 4b Max 8b -- 2025-09-21
  263. GPT-OSS-120B Performance Benchmarks and Provider Trade-Offs -- 2025-09-20
  264. Why are there three different Codex variants? -- 2025-09-20
  265. zli12321/Vision-SR1 -- 2025-09-19
  266. lrzjason/Comfyui-QwenEditUtils -- 2025-09-19
  267. Mini-o3/Mini-o3 -- 2025-09-19
  268. vLLM is kinda awesome -- 2025-09-19
  269. Public AI on Hugging Face Inference Providers 🔥 -- 2025-09-19
  270. HyST: LLM-Powered Hybrid Retrieval over Semi-Structured Tabular Data -- 2025-09-19
  271. GPT-OSS:20b & Qwen 4b are a match made in heaven for 24GB VRAM builds -- 2025-09-18
  272. Was working in RAG recently got to know how well Gemma3 4B performs -- 2025-09-18
  273. [Editorial] which patterns truly survived compression -- 2025-09-16
  274. [Editorial] AI Kill Chain -- 2025-09-16
  275. TsinghuaC3I/Unify-Post-Training -- 2025-09-16
  276. [Editorial] Tricks from OpenAI gpt-oss YOU can use with transformers -- 2025-09-15
  277. openbmb/MiniCPM4.1-8B -- 2025-09-15
  278. nunchaku-tech/nunchaku-qwen-image -- 2025-09-15
  279. ggml-org/gpt-oss-20b-GGUF -- 2025-09-15
  280. MBZUAI releases K2 Think. 32B reasoning model based on Qwen 2.5 32B backbone, focusing on high performance in math, coding and science. -- 2025-09-14
  281. unsloth/Qwen3-Coder-30B-A3B-Instruct-1M-GGUF -- 2025-09-14
  282. [vllm] Hints to run Qwen3-235B MoE on 8x AMD mixed cards! -- 2025-09-12
  283. Inference for 24 people with a 5000€ budget -- 2025-09-12
  284. $142 upgrade kit and spare modules turn Nvidia RTX 4090 24GB to 48GB AI card -- 2025-09-12
  285. Tricks from OpenAI gpt-oss YOU 🫵 can use with transformers -- 2025-09-12
  286. Small Open Models Achieve Near Parity with Large Models in Low Resource Literary Translation at a Fraction of the Cost -- 2025-09-10
  287. Reasoning Vectors: Transferring Chain-of-Thought Capabilities via Task Arithmetic -- 2025-09-10
  288. Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search -- 2025-09-10
  289. Introducing FineVision: a huge open-source dataset for training SOTA Vision Language Models -- 2025-09-10
  290. wildminder/ComfyUI-VibeVoice -- 2025-09-10
  291. bytedance/USO -- 2025-09-10
  292. Wan-AI/Wan2.2-I2V-A14B -- 2025-09-10
  293. [Editorial] Update from Anthropic regarding their poor perfomance of late -- 2025-09-09
  294. LiquidAI/LFM2-VL-450M -- 2025-09-09
  295. An LLM-powered Natural-to-Robotic Language Translation Framework with Correctness Guarantees -- 2025-09-09
  296. Qwen3 30B A3B 2507 Hybrid Deep Reasoning Showcase -- 2025-09-08
  297. Is the "cost of inference" going up or down? -- 2025-09-08
  298. Smartphone Sensors Unlocked: Turn Your Phone into a Physics Lab -- 2025-09-08
  299. UniSLU: Unified Spoken Language Understanding from Heterogeneous Cross-Task Datasets -- 2025-09-08
  300. Voice cloning -- 2025-09-08
  301. haasonsaas/dspy-0to1-guide -- 2025-09-06
  302. 16 reproducible failures → upgraded into a 300+ page Global Fix Map. one link inside, feedback wanted -- 2025-09-06
  303. Show HN: Entropy-Guided Loop – How to make small models reason -- 2025-09-06
  304. Kwaipilot/KAT-V1-40B -- 2025-09-06
  305. Beyond Scaling Law: A Data-Efficient Distillation Framework for Reasoning -- 2025-09-06
  306. nasa-ibm-ai4science/Surya-1.0 -- 2025-09-06
  307. Context Reasoning Benchmarks: GPT-5, Claude, Gemini, Grok on Real Tasks -- 2025-09-05
  308. The CLAUDE.md Framework: A Guide to Structured AI-Assisted Work (prompts included) -- 2025-09-05
  309. Team-intN18-SoybeanSeclab/Typhon -- 2025-09-05
  310. DatarusAI/Datarus-R1-14B-preview -- 2025-09-05
  311. Training & Querying 3 Ollama Models with Zer00logy: Symbolic Cognition Framework and Void-Math OS -- 2025-09-04
  312. I'm building local, open-source, fast, efficient, minimal, and extendible RAG library I always wanted to use -- 2025-09-03
  313. Creating the brain behind dumb models -- 2025-09-03
  314. 🌟Introducing Art-0-8B: Reasoning the way you want it to with Adaptive Thinking🌟 -- 2025-09-02
  315. Fine Tune Model for Home Assistant? -- 2025-09-02
  316. DeepSeek V3.1 improves on the multiplayer Step Game social reasoning benchmark -- 2025-08-31
  317. I built Husk, a native, private, and open-source iOS client for your local models -- 2025-08-31
  318. Would a “Knowledge Coverage Audit” tool be useful for RAG/chatbot builders? -- 2025-08-30
  319. baichuan-inc/Baichuan-M2-32B -- 2025-08-28
  320. Hierarchical Reasoning Model (HRM) implementation for text generation -- 2025-08-27
  321. Datarus-R1-14B-Preview, an adaptive multi-step reasoning LLM for automated data analysis -- 2025-08-24
  322. Fully Open source, serverless, community-driven MCP alternative built in Python, TS and Go -- 2025-08-24
  323. unsloth/Kimi-K2-Instruct-GGUF -- 2025-08-24
  324. DeepSeek V3.1 Reasoner improves over DeepSeek R1 on the Extended NYT Connections benchmark -- 2025-08-24
  325. DeepSeek-V3.1 (Thinking and Non Thinking) -- 2025-08-22
  326. Modify <think> to explore the impact on <answer> -- 2025-08-22
  327. Tiny finance “thinking” model (Gemma-3 270M) with verifiable rewards (SFT → GRPO) — structured outputs + auto-eval (with code) -- 2025-08-22
  328. Qwen/Qwen3-30B-A3B-Thinking-2507 -- 2025-08-22
  329. tencent/Hunyuan-7B-Instruct -- 2025-08-22
  330. 🐧 llama.cpp on Steam Deck (Ubuntu 25.04) with GPU (Vulkan) — step-by-step that actually works -- 2025-08-22
  331. Running Qwen3-Coder-30B-A3 Q4_LM in Cursor with Agent Mode unlocked -- 2025-08-22
  332. Docker now support AI Models, anyone using it? -- 2025-08-22
  333. Why does gpt-oss 120b run slower in ollama than in LM Studio in my setup? -- 2025-08-22
  334. Speculative decoding in archgw candidate release 0.4.0. Could use feedback, -- 2025-08-16
  335. Nvidia Tilus: A Tile-Level GPU Kernel Programming Language -- 2025-08-16
  336. SynapseRoute: An Auto-Route Switching Framework on Dual-State Large Language Model -- 2025-08-16
  337. HoML: vLLM's speed + Ollama like interface -- 2025-08-15
  338. HoML vs. Ollama: A Deep Dive into Performance -- 2025-08-15
  339. Sampler Settings for GLM 4.5-Air -- 2025-08-15
  340. Fully verbal LLM program for OSX using whisper, ollama & XTTS -- 2025-08-15
  341. Closing the Modality Gap for Mixed Modality Search -- 2025-08-07
  342. ByteDance drops Seed-Prover -- 2025-08-06
  343. naver-hyperclovax/HyperCLOVAX-SEED-Think-14B -- 2025-08-06
  344. Context Management by Trimming Conversation -- 2025-08-06
  345. Exploiting Primacy Effect To Improve Large Language Models -- 2025-08-06
  346. [Editorial] HRM -- 2025-08-03
  347. How are people running an MLX-compatible OpenAI API server locally? -- 2025-08-03
  348. I built the perfect MCP client for broke developers (Ollama powered) -- 2025-08-03
  349. character-ai/pipelining-sft -- 2025-08-03
  350. CoexistAI – LLM-Powered Research Assistant (Now with MCP, Vision, Local File Chat, and More) -- 2025-08-02
  351. Best <2B open-source LLMs for European languages? -- 2025-08-02
  352. Local TTS quality -- 2025-08-02
  353. [Editorial] The Anatomy of a Modern LLM -- 2025-07-31
  354. PowerInfer/SmallThinker-21BA3B-Instruct -- 2025-07-31
  355. Has vLLM made Ollama and llama.cpp redundant? -- 2025-07-30
  356. [Editorial] Alternative to vector db rag -- 2025-07-30
  357. How are people extracting system prompts? -- 2025-07-29
  358. cherrydra/mcpurl -- 2025-07-29
  359. [Editorial] neural networks don’t need to be giant to be powerful -- 2025-07-27
  360. Qwen/Qwen3-235B-A22B-Thinking-2507 -- 2025-07-27
  361. mistralai/Magistral-Small-2507 -- 2025-07-27
  362. Running Qwen3 235B-A22B 2507 on a Threadripper 3970X + 3x RTX 3090 Machine at 15 tok/s -- 2025-07-25
  363. The Latest GPT-5 Leaks and Teasers -- 2025-07-25
  364. Qwen3-235B-A22B-Thinking-2507 released! -- 2025-07-25
  365. From chaotic prompting to structured workflow: My Claude evolution -- 2025-07-24
  366. Never Come Up Empty: Adaptive HyDE Retrieval for Improving LLM Developer Support -- 2025-07-24
  367. MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models -- 2025-07-24
  368. Building an MCP Server and Client with FastMCP 2.0 -- 2025-07-24
  369. Does LLM architecture allow for injecting some more input tokens in the middle of token generation? -- 2025-07-24
  370. Lucy: A Mobile-Capable 1.7B Reasoning Model That Rivals Jan-Nano -- 2025-07-23
  371. microsoft/Phi-4-mini-flash-reasoning -- 2025-07-22
  372. LGAI-EXAONE/EXAONE-4.0-32B -- 2025-07-22
  373. Replacing thinking with tool usage enables reasoning in small language models -- 2025-07-22
  374. A Request for Comments (RFC) for MCP-alternative Universal Tool Calling Protocol (UTCP) was created -- 2025-07-22
  375. How to use the same context across LLMs and Agents -- 2025-07-22
  376. new models from NVIDIA: OpenReasoning-Nemotron 32B/14B/7B/1.5B -- 2025-07-21
  377. OpenAI Places Second Behind Human Coder at AtCoder Progmming Event -- 2025-07-21
  378. HelpingAI/Dhanishtha-2.0-preview -- 2025-07-21
  379. Probing for Arithmetic Errors in Language Models -- 2025-07-21
  380. Struggling to Generate Polished UI with Claude Code -- 2025-07-20
  381. IMO 2025 LLM Mathematical Reasoning Evaluation -- 2025-07-20
  382. A comprehensive study of LLM-based argument classification: from LLAMA through GPT-4o to Deepseek-R1 -- 2025-07-19
  383. Madness, the ignorant's question. Would it be possible to lighten an LLM model? -- 2025-07-18
  384. Open source and free iOS app to chat with your LLMs when you are away from home. -- 2025-07-16
  385. Requirements and architecture for a good enough model with scientific papers RAG -- 2025-07-16
  386. Excited to share updates to Open WebUI Starter! New docs, Docker support, and templates for everyone -- 2025-07-16
  387. OpenAI's open source LLM is a reasoning model, coming Next Thursday! -- 2025-07-14
  388. The BastionRank Showdown: Crowning the Best On-Device AI Models of 2025 -- 2025-07-14
  389. Local Llama with Home Assistant Integration and Multilingual-Fuzzy naming -- 2025-07-14
  390. Podcast generation app -- works with Ollama -- 2025-07-14
  391. support for Jamba hybrid Transformer-Mamba models has been merged into llama.cpp -- 2025-07-13
  392. Asynchronous Robot Inference: Decoupling Action Prediction and Execution -- 2025-07-13
  393. Suggestion: Grayscale-First Hack to Optimize Image Recognition in Grok—Save Compute Without Losing Accuracy? -- 2025-07-13
  394. Upskill your LLMs with Gradio MCP Servers -- 2025-07-09
  395. AGI is not multimodal -- 2025-07-09
  396. How Do Vision-Language Models Process Conflicting Information Across Modalities? -- 2025-07-09
  397. High Precision -- 2025-07-09
  398. skt/A.X-4.0 -- 2025-07-09
  399. Medical language model - for STT and summarize things -- 2025-07-09
  400. Ollama alternatives -- 2025-07-09
  401. Dealing with tool_calls hallucinations -- 2025-07-09
  402. SynapseRoute: An Auto-Route Switching Framework on Dual-State Large Language Model -- 2025-07-07
  403. i made a commit message generator that can be used offline and for free -- 2025-07-05
  404. THUDM/GLM-4.1V-9B-Thinking -- 2025-07-05
  405. baidu/ERNIE-4.5-VL-424B-A47B-Base-Paddle -- 2025-07-05
  406. LoRA Fine-Tuning Without GPUs: A CPU-Efficient Meta-Generation Framework for LLMs -- 2025-07-05
  407. skt/A.X-4.0-Light -- 2025-07-04
  408. ChatDOC/OCRFlux-3B -- 2025-07-04
  409. Is there a local model that can solve this text decoding riddle? -- 2025-07-03
  410. Seven replies to the viral Apple reasoning paper and why they fall short -- 2025-07-03
  411. R1-0528 won't stop thinking -- 2025-07-03
  412. Running Deepseek R1 0528 q4_K_M and mlx 4-bit on a Mac Studio M3 -- 2025-07-02
  413. Hoshinonyaruko/Gensokyo-MCP -- 2025-07-01
  414. THU-KEG/AdaptThink -- 2025-06-28
  415. modelcontextprotocol/registry -- 2025-06-27
  416. Skywork/Skywork-SWE-32B -- 2025-06-25
  417. moonshotai/Kimi-VL-A3B-Thinking-2506 -- 2025-06-25
  418. POLARIS-Project/Polaris-4B-Preview -- 2025-06-25
  419. XiaomiMiMo/MiMo -- 2025-06-22
  420. nvidia/AceReason-Nemotron-1.1-7B -- 2025-06-22
  421. Menlo/Jan-nano -- 2025-06-22
  422. MiniMax-AI/SynLogic -- 2025-06-15
  423. The Fractured Entangled Representation Hypothesis -- 2025-06-15
  424. mistralai/Magistral-Small-2506_gguf -- 2025-06-14
  425. Ruminate: From All-or-Nothing to Just-Right Reasoning in LLMs -- 2025-06-14
  426. [update] Restructured repo under rvn-tools — modular CLI for LLM formats -- 2025-06-14
  427. Testing Quant Quality for Shisa V2 405B -- 2025-06-14
  428. Old model, new implementation -- 2025-06-14
  429. Ollama vs Llamacpp: Different output for same model -- 2025-06-14
  430. How to improve my ViT model -- 2025-06-14
  431. From RPC to transactions and durable executions -- 2025-06-14
  432. Flattening Rust’s learning curve -- 2025-06-14
  433. Async from scratch 3: Pinned against the wall -- 2025-06-14
  434. How to get the most out of my AMD 7900XT? -- 2025-06-14
  435. typelevel/cats -- 2025-06-11
  436. wesm/pydata-book -- 2025-06-11
  437. lerobot/smolvla_base -- 2025-06-10
  438. jedisct1/openapi-mcp -- 2025-06-09
  439. sarvamai/sarvam-m -- 2025-06-07
  440. Qwen/Qwen3-Reranker-0.6B -- 2025-06-07
  441. arcee-ai/Homunculus -- 2025-06-04
  442. PRIME-RL/Entropy-Mechanism-of-RL -- 2025-06-02
  443. Atlas: Learning to Optimally Memorize the Context at Test Time -- 2025-06-02
  444. Gen-Verse/MMaDA -- 2025-06-01
  445. osmosis-ai/Osmosis-Structure-0.6B -- 2025-06-01
  446. 0-1 phase transitions in sparse spiked matrix estimation -- 2025-06-01
  447. 0-Step Capturability, Motion Decomposition and Global Feedback Control of the 3D Variable Height-Inverted Pendulum -- 2025-06-01
  448. simplescaling/s1 -- 2025-05-31
  449. FractalAIResearch/Fathom-R1-14B -- 2025-05-31
  450. unsloth/DeepSeek-R1-0528-Qwen3-8B-GGUF -- 2025-05-31
  451. deepseek-ai/DeepSeek-R1-0528-Qwen3-8B -- 2025-05-31
  452. I made Model Version Control Protocol for AI agents -- 2025-05-31
  453. AI Baby Monitor – fully local Video-LLM nanny (beeps when safety rules are violated) -- 2025-05-31
  454. LMStudio - llama.cpp - vLLM -- 2025-05-31
  455. Built an ADK Agent that finds Jobs based on your Resume -- 2025-05-31
  456. Should I resize the image before sending it to Qwen VL 7B? Would it give better results? -- 2025-05-31
  457. How to start a LLM project? -- 2025-05-31
  458. Beware of Fast-Math -- 2025-05-31