Benchmarks & Evaluation

Leaderboards, evaluation frameworks, model comparison

496 articles across 149 editions

Articles

  1. AgentCyberRange: benchmarking frontier AI agents in realistic cyber ranges -- 2026-10-08
  2. AI Security Bootcamp (AISB): open curriculum repo for securing frontier AI systems -- 2026-10-08
  3. dealignai/GLM-5.3-CYBERSECURITY-FP8 (trending on Hugging Face) -- 2026-10-08
  4. Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s -- 2026-10-06
  5. Two local Qwen ( 3.8 27b unsloth Q6 and Qwen flash next strata coder ) models vs Claude Opus 4.6 on the same 3 coding tasks. One of them tied it. Not here to start a fight, just sharing numbers -- 2026-10-06
  6. Qwen4Exp: add MTP by am17an · Pull Request #29761 · ggml-org/llama.cpp -- 2026-10-06
  7. [Discussion] A 5KB pure x86-64 assembly engine for Gemma-2B (FP16, 4.6 tok/s on CPU) -- 2026-10-06
  8. DDR4/PCIe4 vs DDR5/PCIe5 for LLMs- I benchmarked them for pre-training. What are your thoughts? -- 2026-10-06
  9. [Editorial] Post by @ashxhart on X -- 2026-10-06
  10. EmbeddingGemma 2 running locally in-browser on WebGPU -- 2026-10-06
  11. I leaked a deliberately wrong answer key to an LLM and told it not to use it. It matched the key in 63% of answers - and denied it 47 out of 47 times when asked. -- 2026-10-06
  12. Optimal Adversarial Testing: Extracting Honest Test Results from Dishonest Test Takers -- 2026-10-06
  13. reddit.com -- 2026-10-01
  14. [Editorial] -- 2026-10-01
  15. [Editorial] -- 2026-10-01
  16. [Editorial] -- 2026-10-01
  17. reddit.com -- 2026-10-01
  18. Ember-1 -- 2026-09-29
  19. 42x Faster Prompt Lookup Drafting in llama.cpp -- 2026-09-29
  20. MiaAI-Lab/DeepSeek-v4.1-Flash-EXL3-2x-DGX-Sparks -- 2026-09-29
  21. ESP32S3 cluster running 1.58-bit (BitNet) Language model -- 2026-09-29
  22. [Editorial] -- 2026-09-28
  23. [Editorial] -- 2026-09-28
  24. hearim(헤아림): Maybe you don’t need a special model for Jev — ordinary local LLMs already have the capability -- 2026-09-28
  25. SupersonicLabs/Julia-1 · Hugging Face -- 2026-09-28
  26. FIDES: A Concordance Protocol for LLM-Generated Trading Strategies -- 2026-09-25
  27. [Editorial] Contrastive-LM/CLM: contrastive language modeling -- 2026-09-25
  28. Opus 5.5 is the real deal: same accuracy as Fable 5.1 on 25 graded tasks, faster and way cheaper -- 2026-09-25
  29. [Editorial] tee.public.computer: trusted execution for public compute -- 2026-09-25
  30. AMD's random number generator can't generate a 0? -- 2026-09-25
  31. ZuckOff is a free app that sees Meta glasses before they see you -- 2026-09-25
  32. Claude Opus 5.5 -- 2026-09-23
  33. OpenAI GPT–6 Astra breaks Enigma message that has resisted solution since 2005 -- 2026-09-23
  34. Dream RSI (editorial pick) -- 2026-09-23
  35. How much of F-Droid is LLM generated? -- 2026-09-15
  36. Qwen3.8-27B-Uncensored-Genesis-V1-GGUF -- 2026-09-15
  37. Astra and Fable still hack on simple variants of alignment evals from 2025 -- 2026-09-14
  38. DS 4.1 and the new Harness -- 2026-09-14
  39. Should coding agents scan tool results before putting them into context? -- 2026-09-14
  40. decionis/docker -- 2026-09-14
  41. What You Can't See Is What You Learn: Restricted Evidence Visibility Favors Compositional Generalization in Shared-Genome Language-Model Societies -- 2026-09-14
  42. Async GRPO with LoRA across HF Jobs: a bucket, a proxy, and no NCCL -- 2026-09-14
  43. Junchao-cs/SolarWM -- 2026-09-14
  44. Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra -- 2026-09-10
  45. Qwen3.8 Flash Next - Templates Comparison -- 2026-09-10
  46. Qwen 3.8 27b with PI agent - pushed to its 3D graphic game limits -- 2026-09-10
  47. OpenAI alleged of stealing mathematicians work -- 2026-09-10
  48. [Editorial] YouTube video (tbXKZsodiiw) -- 2026-09-10
  49. [Editorial] YouTube video (OZng1eydHJ8) -- 2026-09-10
  50. Closed AI doesn't like biological research, user turns to open weight models -- 2026-09-10
  51. Tristan Buckmaster's statement on the AI-assisted Navier-Stokes proof and the OpenAI/Anthropic dispute -- 2026-09-08
  52. Is mathematics about to enter the conservatory? -- 2026-09-08
  53. BenchMIRT: What are LLM benchmarks actually measuring? -- 2026-09-07
  54. A Four-Axis Trustworthiness Benchmark for LLM-as-Judge in Principle-Based Regulation -- 2026-09-07
  55. [Editorial] -- 2026-09-07
  56. [Editorial] -- 2026-09-07
  57. Path to Astra: critical capabilities and frontier safeguards -- 2026-09-04
  58. Gemini 3.8 Flash and 3.8 Flash Cyber -- 2026-09-04
  59. Extremely Fast and Accurate Transcription with Granite Speech 5.0 Turbo CTC -- 2026-08-28
  60. I implemented a modern LLM in 700 lines of C -- 2026-08-28
  61. Every Model Cheats -- 2026-08-24
  62. [Editorial] DreamLab-AI Loom — Research Paper v4 -- 2026-08-24
  63. [P] synthfin-aml: A graph generator to test if your models actually learn topology (and not just tabular leakage) -- 2026-08-21
  64. H-EmbodVis/TurboVLA -- 2026-08-21
  65. The August 17 outage -- 2026-08-21
  66. [Editorial] RepoRadar -- 2026-08-21
  67. [Editorial] Event-Horizon (ruvnet) -- 2026-08-21
  68. GIMP Development Update -- 2026-08-21
  69. [Editorial] music.cognitum.one -- 2026-08-21
  70. Qwen3.8 27B Q2 vs Q3 vs Qwen3.6 35B-A3B MoE on 12GB VRAM -- 2026-08-17
  71. [Editorial] Huihui Qwen3.8-27B Abliterated Q4_K_M GGUF -- 2026-08-17
  72. Mimir: Did the vikings train a 1.7B killer model? -- 2026-08-17
  73. Revision Prompting: Trades slow (decoded) output tokens for cheap (prefilled) input tokens. -- 2026-08-17
  74. [Editorial] Jason Haddix: DEF CON Is Dead, Security Is Cooked -- 2026-08-17
  75. [Editorial] Video feature -- 2026-08-17
  76. [Editorial] DeepSWE by Datacurve -- 2026-08-13
  77. Anthropic: Introducing The Conceptual Reasoning Index -- 2026-08-13
  78. pathwaycom/arc-task-gen -- 2026-08-13
  79. Meta releases open weights for Muse Glimmer-30B -- 2026-08-13
  80. Qwen3.6 35B (2 min) vs Muse Glimmer 30B (4 min) on custom Llama.cpp build (RTX 5080) -- 2026-08-13
  81. Qwen 35B-A3B MoE vs 27B dense in local coding tests: ~4× faster, much smaller quality gap than I expected -- 2026-08-13
  82. I asked DeepSeek-V4-Flash to work with Muse-Glimmer for Vision ability in PI agent and it produced this -- 2026-08-13
  83. Adaptive Identity Anchoring: Closed-Loop Keyframe Placement for Synthetic Paired Supervision in Video Face Swapping -- 2026-08-13
  84. guillaumemeyer/watermarks-remover -- 2026-08-13
  85. Soul-AILab/SoulX-Singer -- 2026-08-13
  86. A 2.6B model with tool calling and 128K context now runs at 30 tok/s on a phone -- 2026-08-06
  87. inclusionAI/Ling-3.0-flash weights are up on Hugging Face — MIT, BF16 plus an official FP8 -- 2026-08-06
  88. They almost catched up on Frontier performance, so now catching up on prices -- 2026-08-06
  89. Deepseek v4 flash 0731 still not holding up. -- 2026-08-06
  90. Show HN: Fine-tune an 8B model on a 4GB laptop GPU -- 2026-08-04
  91. llama.cpp adds MTP / DSpark support for DeepSeek V4 Flash -- 2026-08-04
  92. DeepSeek V4 Flash 0731 running on M5 Air 32GB with streamed experts trick -- 2026-08-04
  93. Uncensored Multi-Model Releases: LongCat-Flash-Lite, Jamba2-Mini, Nikusui 9B/27B with MTPs -- 2026-08-04
  94. Vacuum 16T: A 16.5-trillion-parameter model that contains nothing -- 2026-08-03
  95. I benchmarked classic vector RAG vs Google's new OKF format vs both combined -- 2026-08-03
  96. [Editorial] Awesome Systematic Trading -- 2026-08-03
  97. [Editorial] -- 2026-07-23
  98. [Editorial] -- 2026-07-23
  99. Are AI labs pelicanmaxxing? -- 2026-07-23
  100. [Editorial] -- 2026-07-16
  101. [Editorial] -- 2026-07-16
  102. [Editorial] -- 2026-07-16
  103. [Editorial] -- 2026-07-16
  104. Goedel-Architect: Streamlining Formal Theorem Proving with Blueprint Generation and Refinement -- 2026-07-15
  105. Looped SSMs: Depth-Recurrence and Input Reshaping for Time Series Classification -- 2026-07-15
  106. Misalignment Between Backpropagation and the Hierarchy of Brain Responses to Images -- 2026-07-15
  107. LeMario: Training a JEPA World Model on Super Mario Bros -- 2026-07-15
  108. Claude Sonnet 5 -- 2026-07-03
  109. A Red-Team Study of Anthropic Fable 5 & Opus 4.8 Models -- 2026-07-03
  110. [Editorial] ruvnet Technical Reference -- 2026-07-03
  111. Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers -- 2026-07-02
  112. New bench designed for smaller models: ObviousBench.com -- 2026-07-02
  113. I built an autonomous dev pipeline and ran the same project head to head: a 27B local on a modded 4090, then again on cheap cloud LLMs -- 2026-07-02
  114. Stdlib or Third-Party? Empirical Performance and Correctness of LLM-Assisted Zero-Dependency Python Libraries -- 2026-07-02
  115. StarTrail-org/PixelRAG -- 2026-06-25
  116. georgebuilds/anneal -- 2026-06-25
  117. numind/NuExtract3 -- 2026-06-25
  118. [Editorial] Explainer Agent Harness Generator -- 2026-06-25
  119. [Editorial] Chainguard Scans Source Code for Malware and Greyware -- 2026-06-12
  120. Are insecure code completions in PyCharm a vulnerability? -- 2026-06-12
  121. ShieldNet-360/prompt-gate -- 2026-06-12
  122. [Editorial] Staris Tech -- 2026-06-12
  123. [Editorial] Hallucinations -- 2026-06-11
  124. Lines of Code Got a Better Publicist -- 2026-06-11
  125. How's Linear so fast? A technical breakdown -- 2026-06-11
  126. Port React Compiler to Rust -- 2026-06-11
  127. Show HN: Extend UI – open-source UI kit for modern document apps -- 2026-06-11
  128. [Editorial] LiveContainer — iOS App Sideloading Container -- 2026-06-03
  129. [Editorial] SideStore — Alternative iOS App Store -- 2026-06-03
  130. [Editorial] idevice_pair — Rust iOS Device Pairing -- 2026-06-03
  131. Welcome NVIDIA Cosmos 3: The First Open Omni-model for Physical AI Reasoning and Action -- 2026-06-02
  132. Nvidia announces new AI chip for personal computers -- 2026-06-02
  133. Nvidia LocateAnything - Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding. (10x faster than Qwen3-VL) -- 2026-06-02
  134. New DeepSWE benchmark finds Claude Opus cheats -- 2026-05-29
  135. ITBench-AA: Frontier Models Score Below 50% on Enterprise IT Tasks — by Artificial Analysis and IBM -- 2026-05-29
  136. Context, Reasoning, and Hierarchy: Cost-Performance Study of Compound LLM Agent Design in an Adversarial POMDP -- 2026-05-29
  137. 10 years of AI robustness tricks (PGD, RLHF, Data Augmentation) are actually computing the same hidden matrix -- 2026-05-29
  138. OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization -- 2026-05-29
  139. Shard — getting to 10× KV cache compression -- 2026-05-29
  140. LightVLM: Efficient inference toolkit for vision-language models -- 2026-05-29
  141. Speech Tokenizer Arena: Side-by-side benchmarking for discrete speech tokenizers -- 2026-05-29
  142. DeepSeek just popped the American AI bubble. -- 2026-05-29
  143. DeepSeek V4 Flash at 8.4 tok/s on 3×3090 — patching GGUF metadata for cchuter's fork -- 2026-05-29
  144. Why are the AI Companies spreading F.U.D. about AI? -- 2026-05-29
  145. Running DeepSeek-V4 locally with 4x legacy RTX 2080 Ti ($2k budget setup). Custom Turing kernels, W8A8 quantization, and 255 prefill tok/s! -- 2026-05-20
  146. Ran the same models across Strix Halo, RTX 3090, and RTX 5070 because I wanted my own numbers -- 2026-05-20
  147. Intel's Crescent Island PCB Leaks, Showing a Massive Xe3P GPU, 16-Pin Connector, 160GB LPDDR5X as Intel Sidesteps the HBM Shortage -- 2026-05-20
  148. Sipeed's K3 RISC-V SBCs can run 30B-parameter LLMs 60 TOPS (INT4), Supports BF16/FP16/INT4 -- 2026-05-20
  149. club-5060ti: practical RTX 5060 Ti local LLM notes and configs -- 2026-05-20
  150. [Editorial] -- 2026-05-18
  151. [Editorial] -- 2026-05-18
  152. [Editorial] -- 2026-05-18
  153. [Editorial] -- 2026-05-18
  154. [Editorial] Synaptic-Tuner — LLM Tuning Framework -- 2026-05-11
  155. [Editorial] Video — AI Tools & Frameworks -- 2026-05-11
  156. A C++ port of Echo-TTS -- 2026-05-11
  157. [Editorial] -- 2026-05-07
  158. ProgramBench: Can we really rebuild huge binaries from scratch? (doesn't look like it) -- 2026-05-07
  159. Adding Benchmaxxer Repellant to the Open ASR Leaderboard -- 2026-05-07
  160. AI Evals Are Becoming the New Compute Bottleneck -- 2026-05-04
  161. Function Calling Harness 2: Schema-Driven CoT Compliance from 9.91% to 100% -- 2026-05-04
  162. Microsoft and OpenAI end their exclusive and revenue-sharing deal -- 2026-05-01
  163. [Editorial] Video: AI Development Insights -- 2026-05-01
  164. Talkie: a 13B vintage language model from 1930 -- 2026-05-01
  165. SWE-bench Verified no longer measures frontier coding capabilities -- 2026-04-29
  166. Opus 4.7: Are these first signs of model collapse? -- 2026-04-29
  167. Ternary Bonsai: Top Intelligence at 1.58 Bits -- 2026-04-22
  168. Personal Eval: Gemma4 26B MoE vs Qwen3.5 27B Dense vs Gemma4 31B Dense Compared -- 2026-04-22
  169. NVIDIA Nemotron-3-Super-120B-A12B-FP8 -- 2026-04-22
  170. High-Fidelity KV Cache Summarization Using Entropy and Low-Rank Reconstruction -- 2026-04-21
  171. Density-Guided Response Optimization: Community-Grounded Alignment via Implicit Acceptance Signals -- 2026-04-21
  172. Building a Fast Multilingual OCR Model with Synthetic Data -- 2026-04-21
  173. Physical Simulator In-the-Loop Video Generation -- 2026-04-21
  174. LLM Novice Uplift on Dual-Use Biology Tasks — 4x Accuracy Boost Bypasses Safeguards -- 2026-04-10
  175. [Editorial] Your AI Is Developing Capabilities Nobody Tested -- 2026-04-10
  176. 1-bit llms on device?! -- 2026-04-07
  177. Running SmolLM2-360M on a Samsung Galaxy Watch 4 (380MB RAM) – 74% RAM reduction in llama.cpp -- 2026-04-07
  178. Built my 10x NVidia V100 AI Server - 320gb vram - vLLM Testing Linux Headless -- 2026-04-07
  179. [Editorial] arxiv:2603.15569 -- 2026-03-30
  180. TinyLoRA: LoRA training works at just 13 parameters -- 2026-03-30
  181. KV rotation PR: q8 quants tank performance on AIME25, recovered with rotation -- 2026-03-30
  182. [Editorial] AI ASIC for LLMs -- 2026-03-30
  183. [Editorial] Heretic -- 2026-03-30
  184. mlx-snn: Spiking Neural Network library for Apple MLX -- 2026-03-30
  185. ARC-AGI-3 -- 2026-03-27
  186. LABSHIELD: A Multimodal Benchmark for Safety-Critical Reasoning and Planning in Scientific Laboratories -- 2026-03-27
  187. [Editorial] IAWG — AI Governance Working Group -- 2026-03-18
  188. Antrophic CEO says 50% entry-level white-collar jobs will be eradicated within 3 years -- 2026-03-18
  189. GPT-5.4 -- 2026-03-09
  190. [Editorial] -- 2026-03-09
  191. How do you automate end to end testing without coding when you vibe coded the whole app -- 2026-03-03
  192. [Editorial] Visual Learning for AI Coding -- 2026-03-03
  193. darrenburns/dv -- 2026-03-03
  194. ReasonDB – open-source document DB where the LLM navigates a tree instead of vector search (RAG alternative) -- 2026-03-03
  195. [Editorial] Claude Code Nano-Banana Plugin -- 2026-03-03
  196. Mercury 2: Fast reasoning LLM powered by diffusion -- 2026-02-26
  197. I Benchmarked Opus 4.6 vs Sonnet 4.6 on agentic PR review and browser QA the results weren't what I expected -- 2026-02-26
  198. [Editorial] Bullshit meter :) -- 2026-02-26
  199. The Qwen team verified that there are serious problems with the data quality of the GPQA and HLE test sets. -- 2026-02-25
  200. Qwen 3.5 craters on hard coding tasks — tested all Qwen3.5 models (And Codex 5.3) on 70 real repos so you don't have to. -- 2026-02-25
  201. ChatGPT isn't the only chatbot pulling answers from Elon Musk's Grokipedia -- 2026-02-25
  202. [Editorial] Benchmarking LLMs for Voice Agent Use Cases -- 2026-02-21
  203. Claude Opus 4.6 Surges Past Forecasts on METR's 50% Time-Horizon Benchmark with Exponential Gains -- 2026-02-21
  204. [Editorial] Unsloth: MiniMax M2.5 Fine-Tuning Guide -- 2026-02-21
  205. [Editorial] When everyone can build software, who learns well? -- 2026-02-19
  206. Sonnet 4.6 feels like Opus 4.5 at Sonnet pricing -- 2026-02-19
  207. Anthropic Raises $30,000,000,000 As Run-Rate Revenue Grew 10x Annually Over Three Years -- 2026-02-19
  208. REASONING AUGMENTED RETRIEVAL (RAR) is the production-grade successor to single-pass RAG -- 2026-02-19
  209. [Editorial] Antigravity Awesome Skills -- 2026-02-18
  210. I built an MCP that connects your agent to 8,000+ skills with zero setup -- 2026-02-18
  211. Is the Nvidia T4 actually viable for 70B (EXL2) daily driving, or is it just pure cope compared to dual 3090s? -- 2026-02-13
  212. Open weight kimi k2.5 overtakes opus 4.5 non thinking on arena -- 2026-02-13
  213. When did we go from 400k to 256k? -- 2026-02-13
  214. [Editorial] https://github.com/d-Rickyy-b/certstream-server-go?tab=readme-ov-file -- 2026-02-13
  215. [Editorial] https://www-cdn.anthropic.com/f21d93f21602ead5cdbecb8c8e1c765759d9e232.pdf -- 2026-02-12
  216. [Editorial] https://d3lm.medium.com/overly-agentic-why-anthropic-is-worried-about-opus-4-6-17eee0f8e5cd -- 2026-02-12
  217. [Editorial] https://www.linkedin.com/posts/avipil_i-got-my-first-bill-after-switching-to-claude-activity-7427320523870629889-vM5K -- 2026-02-12
  218. Pros/Cons and use case for bypassing permissions -- 2026-02-12
  219. [Editorial] https://www.linkedin.com/posts/dragan-spiridonov_agentic-qe-competitive-landscape-2026-activity-7427362099175211010-pd1J -- 2026-02-11
  220. jmuncor/sherlock -- 2026-02-11
  221. [Editorial] https://github.com/mitkox/megacode -- 2026-02-10
  222. [Editorial] https://www.marktechpost.com/2026/02/07/google-ai-introduces-paperbanana-an-agentic-framework-that-automates-publication-ready-methodology-diagrams-and-statistical-plots -- 2026-02-10
  223. [Editorial] https://www.linkedin.com/posts/ryansmith108_frank-lee-amplitude-skills-are-now-indexed-activity-7426777024284893184-8eTf -- 2026-02-10
  224. [Editorial] https://arxiv.org/abs/2602.04118 -- 2026-02-10
  225. Measuring output stability across LLM runs (JSON drift problem) -- 2026-02-09
  226. BalatroBench - Benchmark LLMs' strategic performance in Balatro -- 2026-02-09
  227. ykushch/ask -- 2026-02-05
  228. zaolin/vanguard -- 2026-02-05
  229. Tadpole – A modular and extensible DSL built for web scraping -- 2026-02-05
  230. Coding assistants are solving the wrong problem -- 2026-02-05
  231. How Vibe Coding is Killing Open Source -- 2026-02-05
  232. [Editorial] https://github.com/mondweep/vibe-cast/tree/claude/claude-code-v3-skill-KucJF/claude-code-v3-qe-skill -- 2026-02-04
  233. [Editorial] https://forge-quality.dev/articles/case-of-passing-tests-investigation -- 2026-02-02
  234. MultiX0/last-archive -- 2026-01-28
  235. roborev-dev/roborev -- 2026-01-28
  236. Why I Stopped Using Nbdev -- 2026-01-21
  237. VectorDBZ update: Pinecone, pgvector, custom embeddings, search stats -- 2026-01-19
  238. Prompt tool I built/use with Ollama daily - render prompt variations without worrying about text files -- 2026-01-19
  239. Need people to get excited part 2 -- 2026-01-19
  240. Binary Fuse Filters: Fast and Smaller Than XOR Filters -- 2026-01-19
  241. Read_once(), Write_once(), but Not for Rust -- 2026-01-19
  242. Show HN: HTTP:COLON – A quick HTTP header/directive inspector and reference -- 2026-01-19
  243. [Editorial] https://www.linkedin.com/posts/daniel-cuthbert0x_last-year-i-spent-most-of-my-time-reviewing-activity-7414597548050665472-dYjg -- 2026-01-08
  244. Anyone tried IQuest-Coder-V1 yet? The 40B numbers look wild -- 2026-01-06
  245. open-thoughts/OpenThinker-Agent-v1 -- 2026-01-06
  246. zrougamed/orion-belt -- 2026-01-06
  247. Leo-Mu/montecarlo-ip-searcher -- 2026-01-06
  248. wkhtmltopdf - Convert HTML to PDF Using QtWebKit (2021) -- 2026-01-06
  249. zakirkun/guardian-cli -- 2026-01-05
  250. orneryd/NornicDB -- 2026-01-02
  251. Build a Deep Learning Library -- 2026-01-02
  252. Liquid CO2 For Grid Scale Energy Storage Isn’t Just Hot Air -- 2026-01-02
  253. How llama.cpp implements 2.9x faster top-k sampling with bucket sort -- 2025-12-31
  254. Built an offline-first vector database (v0.2.0) looking for real-world feedback -- 2025-12-31
  255. Linux 7.0 Expected to Bring IO_uring Iopoll Polling Improvements -- 2025-12-31
  256. rix4uni/subhijack -- 2025-12-30
  257. Worktrunk – CLI for Git worktree management -- 2025-12-30
  258. [Editorial] https://github.com/JohannesLks/CVE-2025-14558 -- 2025-12-29
  259. batterdaysahead/cipher0 -- 2025-12-29
  260. MongoBleed -- 2025-12-29
  261. dsl-learn/cutile-learn -- 2025-12-18
  262. Errors in Rust: A Deep Dive -- 2025-12-18
  263. Plug Into USB, Read Hostname and IP Address -- 2025-12-18
  264. Gouryella/drip -- 2025-12-17
  265. Koko-boya/Comfyui-Z-Image-Utilities -- 2025-12-17
  266. Show HN: Generate Passwords from Regex Constraints -- 2025-12-17
  267. Generating synthetic test data for LLM applications (our approach) -- 2025-12-12
  268. Benchmarked A100 vs H100 local storage for Multi-GPU loading. The Gen4 bottleneck is brutal for cold starts. -- 2025-12-11
  269. [Toolkit] TinyLlama Fine-Tuning + RAG Lab (Full FT / LoRA / QLoRA | T4-friendly | Unified pipeline) -- 2025-12-04
  270. Introducing Lynkr — an open-source Claude-style AI coding proxy built specifically for Databricks model endpoints 🚀 -- 2025-12-04
  271. AI Runner v5.0.5 -- 2025-12-04
  272. Which local model for 3090 5069 TI combo -- 2025-12-04
  273. I built a macOS app to monitor all my Claude Code sessions at once -- 2025-12-04
  274. nvidia/Orchestrator-8B · Hugging Face -- 2025-12-03
  275. Llamacpp Parameters Tuning -- 2025-12-02
  276. 4xRTX 4000 Pro Blackwell vs 1x6000 RTX Pro -- 2025-12-02
  277. ardanlabs/kronk -- 2025-12-02
  278. Zig Book – An open, technical and introductory book for Zig -- 2025-12-02
  279. Arcee Trinity Mini: US-Trained Moe Model -- 2025-12-02
  280. Build Your Own Glasshole Detector -- 2025-12-02
  281. Askimo: Open source of Ollama native desktop client -- 2025-12-01
  282. Created 24 Claude Code learning units (beginner → power user) - Free on GitHub -- 2025-12-01
  283. You can now do FP8 reinforcement learning locally! (<5GB VRAM) -- 2025-12-01
  284. A Repository with 44 Years of Unix Evolution -- 2025-11-28
  285. Strix Halo batching with tensor parallel and pipeline parallel using vllm benchmarked -- 2025-11-28
  286. RTX 3090 vs RX 7900 with ROCm, also Vulcan -- 2025-11-26
  287. moonshotai/Kimi-K2-Thinking -- 2025-11-26
  288. Ollama Not Using GPU on RTX 5070 Ti (Blackwell) -- 2025-11-25
  289. PCIE Bifurcation - More than 4 GPUs on a consumer motherboard -- 2025-11-18
  290. Qual a melhor GPU para o llama 3(.1 ou .3) -- 2025-11-18
  291. PyTorch 2.10.0a0 w/ Blackwell (sm_120) Support — Patched & Packaged for One-Command Install -- 2025-11-17
  292. Half-trillion parameter model on a machine with 128 GB RAM + 24 GB VRAM -- 2025-11-17
  293. Real-Time BART in a Box Smaller Than Your Coffee Mug -- 2025-11-17
  294. etalazz/vsa -- 2025-11-13
  295. Pi Compute Modules Make for Compact Cluster -- 2025-11-13
  296. antarys-ai/antarys -- 2025-11-11
  297. [Editorial] https://www.linkedin.com/posts/daniel-cuthbert0x_a-month-ago-gadi-evron-and-i-set-about-building-ugcPost-7393643597729845248-TSTD -- 2025-11-11
  298. Breakdown of New RunC Vulnerabilities -- 2025-11-11
  299. When Your Hash Becomes a String: Hunting Ruby's Million-to-One Memory Bug -- 2025-11-07
  300. Maude 3 Manual -- 2025-11-07
  301. [Editorial] Frequently wrong, but never in doubt’ -- 2025-11-05
  302. The Zero Freeze Formula: Teaching Local LLaMA Real Physics Through Python (SU(3) Mass Gap Simulation) to solve the Yang–Mills Mass Gap -- 2025-11-05
  303. Audio Sound Capture Project Needs Help -- 2025-11-05
  304. [Editorial] https://blog.peerllm.com/2025/11/02/announcing-v0.7.6.html -- 2025-11-04
  305. Faster llama.cpp ROCm performance for AMD RDNA3 (tested on Strix Halo/Ryzen AI Max 395) -- 2025-11-04
  306. KTransformers Open Source New Era: Local Fine-tuning of Kimi K2 and DeepSeek V3 -- 2025-11-04
  307. FlashPack: High-throughput tensor loading for PyTorch -- 2025-11-01
  308. M5 Neural Accelerator benchmark results from Llama.cpp -- 2025-11-01
  309. Kafka is Fast – I'll use Postgres -- 2025-11-01
  310. ZOZO's Contact Solver for physics-based simulations -- 2025-11-01
  311. Need advice on building a GPU-based render/Al compute setup: Unsure about hardware direction -- 2025-11-01
  312. [Editorial] https://pivot-to-ai.com/2025/10/15/ai-is-not-popular-and-ai-users-are-unpleasant-asshats/ -- 2025-10-30
  313. [Editorial] Developer machine part of attack chain -- 2025-10-29
  314. DGX SPARK Compiled llama.cpp Benchmarks Compared to M4 MAX (non-MLX) -- 2025-10-21
  315. perplexityai/search_evals -- 2025-10-21
  316. Hetzner: The Simple Cloud just got more flexible and more affordable -- 2025-10-21
  317. A new, super simple LLM benchmark for testing changes across models, quants, parameters, samplers, engines, etc -- 2025-10-21
  318. Significant speedup for local models -- 2025-10-20
  319. Cursor tricking paid users with fake Claude Sonnet 4.5 -- 2025-10-20
  320. inclusionAI/Ring-1T -- 2025-10-20
  321. 1r0BIT/TaskHound -- 2025-10-18
  322. armai92/goauth -- 2025-10-18
  323. Chinese gang used ArcGIS as a backdoor for a year – and no one noticed -- 2025-10-18
  324. We built 3B and 8B models that rival GPT-5 at HTML extraction while costing 40-80x less - fully open source -- 2025-10-17
  325. Comparing Popular AI Evaluation Platforms for 2025 -- 2025-10-17
  326. State of AI Report 2025 -- 2025-10-17
  327. Signed Backdoor Hiding in Plain Sight on Framework Devices -- 2025-10-15
  328. Three ways formally verified code can go wrong in practice -- 2025-10-15
  329. Jeep pushed software update that bricked all 2024 Wrangler 4xe models -- 2025-10-15
  330. junron/agar -- 2025-10-15
  331. A modern approach to preventing CSRF in Go -- 2025-10-15
  332. Stop flexing Pass@N — show Pass-all-N -- 2025-10-11
  333. Architecting a project for optimal AI coding, any tips? -- 2025-10-11
  334. Basekick-Labs/arc -- 2025-10-11
  335. ServiceNow-AI/Apriel-1.5-15b-Thinker -- 2025-10-11
  336. meituan-longcat/LongCat-Flash-Chat -- 2025-10-11
  337. Did anyone try out GLM-4.5-Air-GLM-4.6-Distill ? -- 2025-10-10
  338. Thank you Anthropic & this community! Our little side project just hit 1M visits and even made it on National TV! -- 2025-10-10
  339. Sneak Preview: Ollama Bench -- 2025-10-08
  340. When Curl Works but IntelliJ Doesn't: The Ollama Connection Mystery -- 2025-10-08
  341. Local Open Deep Research with Offline Wikipedia Search Source -- 2025-10-07
  342. Ollama drops MI50 support -- 2025-10-07
  343. CoexistAI Now Supports Docker Setup, Also now you can turn any text into Podcasts and Speech Easily -- 2025-10-07
  344. MCP_File_Generation_Tool - v0.6.0 Update! -- 2025-10-07
  345. How do I help Codex critique my ideas rather than just go along with it everytime? -- 2025-10-06
  346. Plan with Codex, code with Sonnet 4.5. What's your simple workflow here? -- 2025-10-06
  347. aminofox/zentrox -- 2025-10-06
  348. Linus Torvalds Vents over "Completely Crazy Rust Format Checking" -- 2025-10-06
  349. vllm setup for nvidia (can use llama) -- 2025-10-05
  350. Full-fine tuning doesn't require much vRAM with gradient checkpointing... -- 2025-10-05
  351. Qwen/Qwen3-Omni-30B-A3B-Thinking -- 2025-10-05
  352. inclusionAI/Ring-mini-linear-2.0 -- 2025-10-05
  353. llama.cpp: Quantizing from bf16 vs f16 -- 2025-10-05
  354. GLM 4.6 is nice -- 2025-10-04
  355. NVFP4 or MXFP4 MOE on sm120 (RTX 5900 RTX 6000 PRO) -- 2025-10-04
  356. K2-Think 32B - Reasoning model from UAE -- 2025-10-03
  357. MoonshotAI/checkpoint-engine -- 2025-10-03
  358. Whither the Chip Shortage? -- 2025-10-02
  359. A tiny receipt per AI run: κ (stress), Δhol (drift), and guards—in plain JSON. -- 2025-10-02
  360. Microsoft Agent Framework (Preview): Making AI Agents Simple for Every Developer -- 2025-10-02
  361. How bad to have RTX Pro 6000 run at PCIE x8? -- 2025-09-24
  362. A.S.E: A Repository-Level Benchmark for Evaluating Security in AI-Generated Code -- 2025-09-24
  363. SWE-Bench Pro -- 2025-09-23
  364. Investigating Training Data Detection in AI Coders -- 2025-09-23
  365. Comparison H100 vs RTX 6000 PRO with VLLM and GPT-OSS-120B -- 2025-09-23
  366. Answer Matching Outperforms Multiple Choice for Language Model Evaluation -- 2025-09-21
  367. facebook/MobileLLM-R1-950M -- 2025-09-20
  368. OpenGVLab/InternVL3_5-241B-A28B -- 2025-09-20
  369. KBlueLeaf/HDM-xut-340M-anime -- 2025-09-20
  370. Definitive proof openai/gpt-oss-20b is dumb as hell -- 2025-09-19
  371. Qwen3‑Next‑80B‑A3B‑Instruct (FP8) on Windows 11 WSL2 + vLLM + Docker (Blackwell) -- 2025-09-19
  372. Free 10%+ Speedup for CPU/Hybrid Inference on Intel CPUs with Efficiency Cores -- 2025-09-17
  373. PSA/RFC: KV Cache quantization forces excess processing onto CPU in llama.cpp -- 2025-09-15
  374. native tool calling support for DeepSeek V3.1 just merged in llama.cpp -- 2025-09-15
  375. model : add grok-2 support by CISC · Pull Request #15539 · ggml-org/llama.cpp -- 2025-09-15
  376. [Editorial] Defeating Nondeterminism in LLM Inference -- 2025-09-14
  377. Nvidia Unveils Rubin CPX Amidst Chart-Topping Blackwell Ultra MLPerf Results -- 2025-09-14
  378. Repair-R1: Better Test Before Repair -- 2025-09-14
  379. Jupyter Agents: training LLMs to reason with notebooks -- 2025-09-14
  380. Intel Files Patent for "Software Defined Super Cores" -- 2025-09-04
  381. I tried almost every tts model on my ryzen 7 5000 series 16gb ram rtx 3060 laptop 6-8GB Vram -- 2025-09-02
  382. devnen/Kitten-TTS-Server -- 2025-09-02
  383. internlm/Intern-S1-mini -- 2025-09-02
  384. stepfun-ai/Step-Audio-2-mini -- 2025-09-02
  385. QuEST/Quartet authors discuss their work on SOTA 4-bit training optimizations -- 2025-09-01
  386. F-Stack – A network development kit with high performance based on DPDK -- 2025-09-01
  387. An Empirical Study of Knowledge Distillation for Code Understanding Tasks -- 2025-09-01
  388. A Comparative Analysis of Vision Language Models for Scientific Data Interpretation -- 2025-08-31
  389. CaddyManager 0.0.1 – Web UI for managing Caddy servers -- 2025-08-30
  390. [Editorial] AI interfaces for future -- 2025-08-29
  391. I’ve Debugged 100+ RAG/LLM Pipelines. These 16 Bugs Always Come Back. (70 days, 800 stars) -- 2025-08-29
  392. Updates to Consumer Terms and Privacy Policy -- 2025-08-29
  393. LLM speedup breakthrough? 53x faster generation and 6x prefilling from NVIDIA -- 2025-08-29
  394. Intel Granite Rapids CPU on sale at Newegg up to 65% off MSRP -- 2025-08-29
  395. unsloth/DeepSeek-V3.1-GGUF -- 2025-08-29
  396. Deepseek V3.1 benchmarks released -- 2025-08-25
  397. Is openrouters tokens per second reading super bugged? -- 2025-08-22
  398. It’s a Pi, But it’s not Quite a Raspberry Pi -- 2025-08-22
  399. nvidia/Llama-3_3-Nemotron-Super-49B-v1_5 -- 2025-08-21
  400. Mistral 7B fine tuning training loss stagnant after adding more fine tuning prompts -- 2025-08-20
  401. Detecting Hallucinations in LLM Function Calling with Entropy (Part 2) -- 2025-08-20
  402. Anyone have the deets on ROCM 7.0's 3x perf claims? -- 2025-08-19
  403. Rust in 2025: Targeting foundational software -- 2025-08-19
  404. I built a small cli tool to execute agentic workflows -- 2025-08-19
  405. AvatarNova - Local AI companion -- 2025-08-19
  406. 🤖 Built an AI-powered DOCX viewer that extracts & analyzes images with Ollama! -- 2025-08-19
  407. Davincible/claude-code-open -- 2025-08-19
  408. OpenVINO GenAI 2025.2 adds a GGUF reader (preview) -- 2025-08-18
  409. CLI Agent that Supports Multiple Models? -- 2025-08-18
  410. Is there a standard oci image format for models? -- 2025-08-18
  411. moonshotai/Kimi-K2-Instruct -- 2025-08-17
  412. KittenML/kitten-tts-nano-0.1 -- 2025-08-17
  413. ilkerzgi/Overlay-Kontext-Dev-LoRA -- 2025-08-17
  414. JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 -- 2025-08-17
  415. PSA: Don't waste time trying Gemma 3 27B on V100s - it's architecturally impossible -- 2025-08-16
  416. People with MacBook Pro with 36gb of memory, which models you are running for coding? -- 2025-08-16
  417. GLiNER2: An Efficient Multi-Task Information Extraction System with Schema-Driven Interface -- 2025-08-13
  418. TextQuests: How Good are LLMs at Text-Based Video Games? -- 2025-08-13
  419. Introducing Trackio: A Lightweight Experiment Tracking Library from Hugging Face -- 2025-08-09
  420. [Editorial] https://unhypedai.substack.com/p/unhyped-ai-week-4-digest -- 2025-08-04
  421. 100+ AI Benchmarks list -- 2025-08-04
  422. google/langextract -- 2025-08-04
  423. Chain-GPT/Solidity-LLM -- 2025-07-25
  424. Anyone interested in adding their fine-tuned / open source models to this benchmark? -- 2025-07-25
  425. What kind of rig would you build with a 5k budget for local LLM? -- 2025-07-16
  426. What is your "perfect" £10,000 for Local LLM, Gaming, plex with the following conditional and context. -- 2025-07-16
  427. How to use Claude code -- 2025-07-16
  428. unsloth/Kimi-K2-Instruct-GGUF -- 2025-07-16
  429. moonshotai/Kimi-K2-Base -- 2025-07-16
  430. It's been a while, I'm out of date, suggest me a model -- 2025-07-16
  431. i need the best local llm i can run on my gaming pc -- 2025-07-16
  432. Import of chatgbt Export Zip File with Images of entire previous chats -- 2025-07-14
  433. Which is the best small local LLM models for tasks like doing research and generating insights -- 2025-07-10
  434. Looking for practical advice with my MSc thesis “On-Premise Orchestration of SLMs” (OpenWebUI + SLM v LLM benchmarking on multiple GPUs) -- 2025-07-10
  435. Deep Research with local LLM and local documents -- 2025-07-10
  436. WikipeQA : An evaluation dataset for both web-browsing agents and vector DB RAG systems -- 2025-07-10
  437. Looking for advice. -- 2025-07-10
  438. OpenAI to release open-source model this summer - everything we know so far -- 2025-07-09
  439. Accelerating Docker Builds by Halving EC2 Boot Time -- 2025-06-30
  440. Show HN: Inspect and extract files from MSI installers directly in your browser -- 2025-06-28
  441. Meet Mistral Devstral, SOTA open model designed specifically for coding agents -- 2025-06-26
  442. 1.93bit Deepseek R1 0528 beats Claude Sonnet 4 -- 2025-06-26
  443. DeepSeek R1 05/28 performance on five independent benchmarks -- 2025-06-26
  444. Few-Shot Examples: Overfitting / Leakage -- 2025-06-26
  445. Finetune a model to think and use tools -- 2025-06-26
  446. I need help using open web UI with Ollama. Help installing and getting it running win 11 -- 2025-06-26
  447. I built/am building a micro-transformer for learning and experimentation -- 2025-06-26
  448. I shipped more code yesterday with Claude 4 than the last 3 weeks combined -- 2025-06-26
  449. A deep dive into self-improving AI and the Darwin-Gödel Machine -- 2025-06-26
  450. Shisa V2 405B: The strongest model ever built in Japan! (JA/EN) -- 2025-06-25
  451. Is it possible to give Gemma 3 or any other model on-device screen awareness? -- 2025-06-25
  452. Chonkie update. -- 2025-06-25
  453. Memory Layer Compatible with Local Llama -- 2025-06-25
  454. Ollama/AnythingLLM on Windows 11 with AMD RX 6600: GPU Not Utilized for LLM Inference - Help! -- 2025-06-25
  455. How can synthetic data improve a model if the model was the thing that generated that data? -- 2025-06-25
  456. After reading OpenAI's GPT-4.1 prompt engineering cookbook, I created this comprehensive Python coding template -- 2025-06-25
  457. The 55% Regret Club: How AI-First Companies Are Learning Lessons the Hard Way -- 2025-06-25
  458. Microsoft-backed Builder.ai enters insolvency proceedings -- 2025-06-25
  459. Show HN: FaynoSync Self-Hosted API for Automatic App Updates -- 2025-06-23
  460. identicallead/mse6 -- 2025-06-22
  461. GCC 13.4 Released with 129 additional bug fixes -- 2025-06-22
  462. Databricks acquires Neon -- 2025-06-22
  463. Java Virtual Threads Ate My Memory: A Web Crawler's Tale of Speed vs. Memory -- 2025-06-20
  464. Show HN: Zeekstd – Rust Implementation of the ZSTD Seekable Format -- 2025-06-20
  465. DeepSeek R1 05 28 Tested. It finally happened. The ONLY model to score 100% on everything I threw at it. -- 2025-06-17
  466. ubergarm/DeepSeek-R1-0528-GGUF -- 2025-06-17
  467. LLM training on RTX 5090 -- 2025-06-17
  468. [DEMO] I created a coding agent that can do dynamic, runtime debugging. -- 2025-06-17
  469. Is anyone productively using Aider and Ollama together? -- 2025-06-17
  470. For everyone who's still confused by Attention... I made this spreadsheet just for you(FREE) -- 2025-06-17
  471. What setup/model do you use and what’s your monthly spend? -- 2025-06-17
  472. Xiaomi released an updated 7B reasoning model and VLM version claiming SOTA for their size -- 2025-06-17
  473. hhftechnology/middleware-manager -- 2025-06-17
  474. Show HN: McWig – A modal, Vim-like text editor written in Go -- 2025-06-17
  475. The Unreliability of LLMs and What Lies Ahead -- 2025-06-10
  476. 007: Democratically Finding The Cause of Packet Drops -- 2025-06-08
  477. langtalks/swe-agent -- 2025-06-08
  478. wey-gu/py-pglite -- 2025-06-08
  479. 0-$π$ qubit in one Josephson junction -- 2025-06-07
  480. 100-kT Magnetic field generation using paisley targets by femtosecond laser-plasma interactions -- 2025-06-07
  481. fileshare-go/fileshare -- 2025-06-04
  482. Rust Coreutils 0.1.0 Release -- 2025-06-04
  483. Show HN: Samchika – A Java Library for Fast, Multithreaded File Processing -- 2025-06-04
  484. 100 Drivers, 2200 km: A Natural Dataset of Driving Style toward Human-centered Intelligent Driving Systems -- 2025-06-04
  485. 100+ Metrics for Software Startups - A Multi-Vocal Literature Review -- 2025-06-04
  486. Building a plug-and-play vector store for any data stream (text, audio, video, etc.)—searchable by your LLM via MCP -- 2025-05-29
  487. Building a real-world LLM agent with open-source models—structure > prompt engineering -- 2025-05-29
  488. New LocalLLM Hardware complete -- 2025-05-29
  489. Parameter-Efficient Fine-Tuning (PEFT) Explained -- 2025-05-29
  490. LLM help for recovering deleted data? -- 2025-05-29
  491. AI Runner v4.10.0 Release Notes -- 2025-05-29
  492. Unpopular opinion: RAG is actively hurting your coding agents -- 2025-05-29
  493. Teal – A statically-typed dialect of Lua -- 2025-05-29
  494. I think it's time to give Nix a chance -- 2025-05-29
  495. deepseek-ai/DeepSeek-R1-0528 -- 2025-05-29
  496. AM5 or TRX4 for local LLMs? -- 2025-05-29