Quantization & Efficiency
Model compression, GGUF, efficient inference, optimization
425 articles across 119 editions
Articles
- [Editorial] Benchmarking LLMs for Voice Agent Use Cases -- 2026-02-21
- Claude Opus 4.6 Surges Past Forecasts on METR's 50% Time-Horizon Benchmark with Exponential Gains -- 2026-02-21
- [Editorial] Unsloth: MiniMax M2.5 Fine-Tuning Guide -- 2026-02-21
- Let your coding agent benchmark llama.cpp for you (auto-hunt the fastest params per model) -- 2026-02-06
- GGML implementation of Qwen3-ASR -- 2026-02-06
- Running LLMs & VLMs Fully On-Device on iPhone(6GB RAM) — Offline, Privacy-Focused, Real-Time Performance -- 2026-02-06
- We benchmarked every 4-bit quantization method in vLLM 👀 -- 2026-01-12
- Gpu inference with model that does not fit in one GPU -- 2026-01-12
- Llama.cpp rpc experiment -- 2026-01-12
- Performance improvements in llama.cpp over time -- 2026-01-12
- [Editorial] https://docs.rs/crate/bitchat-qudag/latest -- 2026-01-02
- [Editorial] https://github.com/permissionlesstech/bitchat/blob/main/WHITEPAPER.md -- 2026-01-02
- [Editorial] https://www.npmjs.com/package/@ruvector/edge-net -- 2026-01-02
- Why I Ditched Serverless Neptune/OpenSearch for Dockerized Neo4j/pgvector on EC2 (60% Cost Cut) -- 2025-12-30
- Llama-3.3-8B-Instruct -- 2025-12-30
- Benchmarking local llms for speed with CUDA and vulkan, found an unexpected speedup for select models -- 2025-12-30
- Why Kimi K2 Thinking choose Int4 QAT, from infra enginner of KImi -- 2025-12-30
- Help RTX 5090 + llama.cpp crashes after 2-3 inferences (VFIO passthrough, SM120 CUDA) -- 2025-12-30
- AI-Doomsday-Toolbox Distributed inference + workflows -- 2025-12-30
- [Tool] imesde: Zero-GPU, In-Memory Vector Engine for Real-Time Local RAG -- 2025-12-22
- I built a Rust-based HTML-to-Markdown converter to save RAG tokens (Self-Hosted / API) -- 2025-12-22
- Golang optimizations for high‑volume services -- 2025-12-12
- PaCoRe: The first open-source deep think 8B model beats GPT-5 on HMMT25 -- 2025-12-11
- RnJ-1-Instruct FP8 Quantization -- 2025-12-10
- Optical Context Compression Is Just (Bad) Autoencoding -- 2025-12-10
- Masked Diffusion Models as Energy Minimization -- 2025-12-10
- Miles + FSDP2 = Megatron-Level Performance with More Flexibility -- 2025-12-10
- P4nda0s/IDA-NO-MCP -- 2025-12-09
- Toyota unintended acceleration and the big bowl of "spaghetti" code (2013) -- 2025-12-09
- https://huggingface.co/Doradus/Hermes-4.3-36B-FP8 -- 2025-12-09
- Support for rnj-1 now in llama.cpp -- 2025-12-09
- Comfy-Org/flux2-dev -- 2025-12-09
- baidu/ERNIE-4.5-VL-28B-A3B-Thinking -- 2025-12-09
- I built a personal assistant script, and the CPU inference speed beats my Llama setup. -- 2025-12-08
- Semantic Compression (2014) -- 2025-12-08
- A Deep Dive into Using PIO and DMA on the RP2350 -- 2025-12-05
- Free yourself from the Spotify desktop client with spotifyd -- 2025-12-04
- I cooked abliterated gemma3-27b-it with norm-preserving technique -- 2025-12-04
- Qwen3 VL built from scratch with PyTorch -- 2025-12-03
- EmbeddingGemma: Powerful and Lightweight Text Representations -- 2025-12-03
- Z-Image: Powerful and highly efficient image generation model with 6B parameters -- 2025-12-03
- RTX 5090 + Qwen 30B MoE @ 135 tok/s in NVFP4 - Full guide with C++ patches -- 2025-12-02
- [Editorial] https://huggingface.co/blog/grimjim/norm-preserving-biprojected-abliteration -- 2025-12-02
- Optimizing Token Generation in llama.cpp's CUDA Backend -- 2025-12-01
- [Editorial] https://arxiv.org/html/2511.09030v1 -- 2025-11-28
- You're using HuggingFace wrong. Stop downloading pre-quantized GGUFs and start building hardware-optimized, domain-specific models. Here's the pipeline I built to do it. -- 2025-11-26
- Binary Quantization For LLMs Through Dynamic Grouping -- 2025-11-26
- dx8152/Relight -- 2025-11-26
- Question About Motherboards -- 2025-11-26
- [Release] DragonMemory: 16× semantic compression for local RAG context (open-source, AGPL) -- 2025-11-25
- ORPO-Distill: Mixed-Policy Preference Optimization for Cross-Architecture LLM Distillation -- 2025-11-25
- Continuous batching from first principles -- 2025-11-25
- Can an expert chime in and explain what is holding Vulkan back from becoming the standard API for ML? -- 2025-11-25
- luozijian1990/network-traffic-ebpf-exporter -- 2025-11-24
- We found cryptography bugs in the elliptic library using Wycheproof -- 2025-11-24
- Show HN: Cynthia – Reliably play MIDI music files – MIT / Portable / Windows -- 2025-11-24
- Browser Fingerprinting and Why VPNs Won’t Make You Anonymous -- 2025-11-24
- [Release] Memory-Isolated Recursive Compression (MIRC). A local-first probabilistic compression utility for Apple Silicon. Research Preview (Open Source) -- 2025-11-21
- Read long podcasts locally with Whisper + LLM, open sourced -- 2025-11-21
- Local all-in-one AI system (Local multimodal AI) -- 2025-11-21
- JMS1717/8mb.local -- 2025-11-21
- Mimir Memory Bank now uses llama.cpp! -- 2025-11-21
- Quantum physicists have shrunk and "de-censored" DeepSeek R1 -- 2025-11-20
- Built a tool to solve the "how much GPU do I actually need?" problem for LLM deployment -- 2025-11-20
- New Parameter Browser added to Llamacpp Model Launcher! experimental model parameter tuning(window/cuda only) -- 2025-11-20
- cuda device list mismatch - ggml_cuda_init / ubuntu - significance to using --main-gpu flag -- 2025-11-20
- What Size of LLM Can 4x RTX 5090 Handle? (96GB VRAM) -- 2025-11-20
- Gain 60% performance on RDNA 4 using this fix -- 2025-11-19
- wildminder/ComfyUI-DyPE -- 2025-11-19
- lightx2v/Autoencoders -- 2025-11-19
- Scale-out is the silent killer of LLM applications. Are we solving the wrong problem? -- 2025-11-19
- PyTorch 2.10.0a0 w/ Blackwell (sm_120) Support — Patched & Packaged for One-Command Install -- 2025-11-17
- Half-trillion parameter model on a machine with 128 GB RAM + 24 GB VRAM -- 2025-11-17
- [Editorial] https://www.linkedin.com/posts/andriyburkov_when-you-train-a-model-on-one-dataset-it-activity-7392804316769701888-166x/ -- 2025-11-14
- xRouter: Training Cost-Aware LLMs Orchestration System via Reinforcement Learning -- 2025-11-14
- [Editorial] Balancing order, freedom, and technology -- 2025-11-12
- AMD warns the Intel and Nvidia partnership is a risk to its business -- 2025-11-12
- A Pentium In Your Hand -- 2025-11-12
- [Editorial] https://www.linkedin.com/posts/ismaelvelasco_theres-an-ai-text-model-comparable-to-sota-activity-7393850964731912192-nZT1 -- 2025-11-12
- Last week in Multimodal AI - Local Edition -- 2025-11-12
- Apache Iggy is a high-performance, persistent message streaming platform -- 2025-11-07
- [P] Training Better LLMs with 30% Less Data – Entropy-Based Data Distillation -- 2025-11-06
- I fine tuned a (small) model to help with reasoning backfill on old/non-reasoning datasets -- 2025-11-06
- Superhuman AI for Multiplayer Poker -- 2025-11-06
- cerebras/GLM-4.5-Air-REAP-82B-A12B -- 2025-11-06
- Retrieval Enhanced Feedback via In-context Neural Error-book -- 2025-11-06
- Kimi release Kimi K2 Thinking, an open-source trillion-parameter reasoning model -- 2025-11-06
- OpenAI asks U.S. for loan guarantees to fund $1T AI expansion -- 2025-11-06
- Kimi released Kimi K2 Thinking, an open-source trillion-parameter reasoning model -- 2025-11-06
- [Editorial] Frequently wrong, but never in doubt’ -- 2025-11-05
- The Zero Freeze Formula: Teaching Local LLaMA Real Physics Through Python (SU(3) Mass Gap Simulation) to solve the Yang–Mills Mass Gap -- 2025-11-05
- Audio Sound Capture Project Needs Help -- 2025-11-05
- [D] It turns out WDDM driver mode is making our RAM - GPU transfer extremely slower compared to TCC or MCDM mode. Anyone has figured out the bypass NVIDIA software level restrictions? -- 2025-11-05
- GLaDOS TTS finetuning on MLX from the original game files -- 2025-11-04
- zeusftk/FTK_CANVAS_AGENT_for_Comfyui -- 2025-11-04
- guyyariv/DyPE -- 2025-11-04
- Qwen3-VL-32B Q8 speeds in llama.cpp vs vLLM FP8 on a RTX PRO 6000 -- 2025-11-03
- Help me decide: EPYC 7532 128GB + 2 x 3080 20GB vs GMtec EVO-X2 -- 2025-11-03
- amd/Nitro-E -- 2025-11-03
- CISA and NSA share tips on securing Microsoft Exchange servers -- 2025-11-02
- The Smol Training Playbook: The Secrets to Building World-Class LLMs -- 2025-11-02
- Latest Update from Anthropic's new model - Neptune V6 -- 2025-11-02
- AI "Phone Farm" Startup Gets Funding from Marc Andreessen to Flood Social Media With Spam -- 2025-11-02
- FlashPack: High-throughput tensor loading for PyTorch -- 2025-11-01
- M5 Neural Accelerator benchmark results from Llama.cpp -- 2025-11-01
- Kafka is Fast – I'll use Postgres -- 2025-11-01
- [Editorial] https://www.linkedin.com/posts/busiel-morley_economic-shifts-in-the-age-of-ai-ugcPost-7390349517612806144-8djS -- 2025-11-01
- Analog Surround Sound Was Everywhere, But You Probably Didn’t Notice -- 2025-11-01
- US Gas Turbine Shortage Likely to Slow AI Demand Growth -- 2025-10-31
- The Supercon 2025 Badge is Built to be Customized -- 2025-10-31
- Experimenting with Qwen3-VL for Computer-Using Agents -- 2025-10-30
- Built a full voice AI assistant running locally on my RX 6700 with Vulkan - Proof AMD cards excel at LLM inference -- 2025-10-30
- Streaming datasets: 100x More Efficient -- 2025-10-30
- Cerebras REAP'd GLM4.6: 25%, 30%, 40% pruned FP8 checkpoints on HF! -- 2025-10-28
- Qwen/Qwen3-VL-30B-A3B-Instruct-FP8 -- 2025-10-28
- lightx2v/Wan2.2-Distill-Loras -- 2025-10-28
- [Editorial] Periodic table for ai algorithms -- 2025-10-26
- Need help understanding OpenAIs API usage for text-embedding -- 2025-10-26
- Qwen3 Next support in llama.cpp ready for review -- 2025-10-25
- GLM Air REAP tool call problems -- 2025-10-25
- Reverse Engineering STL Files with FreeCAD -- 2025-10-25
- Un-LOCC (Universal Lossy Optical Context Compression), Achieve Up To 3× context compression with 93.65% Accuracy. -- 2025-10-24
- LiquidAI/LFM2-1.2B-RAG -- 2025-10-24
- zai-org/GLM-4.6 -- 2025-10-24
- inference-net/Schematron-3B -- 2025-10-21
- [By GLM Team] Glyph: Scaling Context Windows via Visual-Text Compression -- 2025-10-21
- DGX SPARK Compiled llama.cpp Benchmarks Compared to M4 MAX (non-MLX) -- 2025-10-21
- perplexityai/search_evals -- 2025-10-21
- Hetzner: The Simple Cloud just got more flexible and more affordable -- 2025-10-21
- A new, super simple LLM benchmark for testing changes across models, quants, parameters, samplers, engines, etc -- 2025-10-21
- riptideslabs/tokenex -- 2025-10-20
- Multi-Tenant SaaS's Wildcard TLS: An Overview of DNS-01 Challenges -- 2025-10-20
- From cloud to OCP? Be ready to wrangle firmware -- 2025-10-20
- FLOSS Weekly Episode 851: Buckets of Money -- 2025-10-20
- Significant speedup for local models -- 2025-10-20
- Cursor tricking paid users with fake Claude Sonnet 4.5 -- 2025-10-20
- inclusionAI/Ring-1T -- 2025-10-20
- volantvm/volant -- 2025-10-19
- Wireshark 4.6.0 Supports macOS Pktap Metadata (PID, Process Name, etc.) -- 2025-10-19
- A classified network of SpaceX satellites is emitting a mysterious signal -- 2025-10-19
- linkedlist771/SoraWatermarkCleaner -- 2025-10-19
- Qwen/Qwen-Image-Edit-2509 -- 2025-10-19
- The Entire Process of Building an Open Source Analog ASIC -- 2025-10-15
- Built a 1288x RTFx Parakeet Speech-to-Text server... Enjoy! -- 2025-10-13
- Novel OpenGL Pixel Shader Dewarping -- 2025-10-13
- lovis93/next-scene-qwen-image-lora-2509 -- 2025-10-13
- Beyond Token Count: Our Research Suggests "Contextual Weight" is a Key Limiter on Large Context Windows -- 2025-10-13
- FractalAIResearch/Fathom-Search-4B -- 2025-10-13
- LLM Robustness Leaderboard v1 --Technical report -- 2025-10-13
- Hide and Seek with LLMs: An Adversarial Game for Sneaky Error Generation and Self-Improving Diagnosis -- 2025-10-13
- Preference optimization with ORPO and LoRA -- 2025-10-12
- [Show] SpiralTorch: A Rust-based PyTorch-style autograd engine (Python 3.14-ready) -- 2025-10-12
- Qwen3-VL-30B-A3B-Thinking GGUF with llama.cpp patch to run it -- 2025-10-10
- Did anyone try out GLM-4.5-Air-GLM-4.6-Distill ? -- 2025-10-10
- What and when 7900xtx is boosted? -- 2025-10-10
- Modelfile. Do I need these tags PER prompt? -- 2025-10-10
- Divining Air Quality With A Cheap Computer Vision Device -- 2025-10-09
- Awesome Local LLM Speech-to-Speech Models & Frameworks -- 2025-10-08
- FabioSarracino/VibeVoice-Large-Q8 -- 2025-10-08
- CapRecover: A Cross-Modality Feature Inversion Attack Framework on Vision Language Models -- 2025-10-08
- Mitigating Watermark Stealing Attacks in Generative Models via Multi-Key Watermarking -- 2025-10-08
- How to make the AI Bot to understand the exact design and App flow -- 2025-10-04
- Creating a Full Stack App W/Cloudflare Works and BetterAuth -- 2025-10-04
- Comprehension debt: A ticking time bomb of LLM-generated code -- 2025-10-04
- deepseek-ai/DeepSeek-V3.2-Exp -- 2025-10-03
- moondream/moondream3-preview -- 2025-10-03
- [Editorial] https://github.com/emcie-co/parlant -- 2025-10-02
- Built a persistent memory system for LLMs - 3 months testing with Claude/Llama -- 2025-10-02
- Do I need to run /init on a repo if I already have AGENTS.md? -- 2025-10-02
- Inside NVIDIA GPUs: Anatomy of high performance matmul kernels -- 2025-09-29
- Bit is all we need: binary normalized neural networks -- 2025-09-29
- Accelerating Qwen3-8B Agent on Intel® Core™ Ultra with Depth-Pruned Draft Models -- 2025-09-29
- This $5,999 RTX PRO 6000 Ebay listing is a scam, right? -- 2025-09-26
- Accelerating Local AI on Consumer GPUs: A Hardware-Aware Dynamic Strategy for YOLOv10s -- 2025-09-26
- Efficient 4B parameter gpt OSS distillation without the over-censorship -- 2025-09-22
- [Project] I created an AI photo organizer that uses Ollama to sort photos, filter duplicates, and write Instagram captions. -- 2025-09-22
- Pointer Tagging in C++: The Art of Packing Bits into a Pointer -- 2025-09-22
- inclusionAI/Ring-mini-2.0 -- 2025-09-22
- Local real-time assistant that remembers convo + drafts a doc -- 2025-09-22
- XiaomiMiMo/MiMo-Audio-7B-Instruct -- 2025-09-21
- Scaling Self-Supervised Representation Learning for Symbolic Piano Performance -- 2025-09-21
- Uncensor Qwen3 models without retraining -- 2025-09-20
- Depth upscaling? -- 2025-09-20
- Qwen3‑Next‑80B‑A3B‑Instruct (FP8) on Windows 11 WSL2 + vLLM + Docker (Blackwell) -- 2025-09-19
- unsloth/Qwen3-Next-80B-A3B-Instruct -- 2025-09-19
- The AI-Scraping Free-for-All Is Coming to an End -- 2025-09-18
- Visible Watermarking with Gradio -- 2025-09-18
- xiaomi-research/q-frame -- 2025-09-17
- google/embeddinggemma-300m -- 2025-09-17
- 3-month Claude Code Max user review - considering alternatives -- 2025-09-15
- Chesars/whatsapp-mcp -- 2025-09-15
- Claude’s memory architecture is the opposite of ChatGPT’s -- 2025-09-15
- The Internet Will Be More Dead Than Alive Within 3 Years, Trend Shows | All signs point to a future internet where bot-driven interactions far outnumber human ones. -- 2025-09-15
- New "speech" mode in Imagine... -- 2025-09-15
- I made local RAG, web search, and voice mode on iPhones completely open source, private, and free -- 2025-09-08
- jwest33/jam_model_memory -- 2025-09-08
- How was your experience with Claude vs Codex? -- 2025-09-08
- [Project/Code] Fine-Tuning LLMs on Windows with GRPO + TRL -- 2025-09-07
- nvidia/NVIDIA-Nemotron-Nano-12B-v2-Base -- 2025-09-07
- huihui-ai/Huihui-gpt-oss-20b-BF16-abliterated -- 2025-09-07
- RX570 compatibility issues -- 2025-09-07
- Continue.dev setup -- 2025-09-07
- Little SSM (RWKV7 7B) state checkpointing demo. -- 2025-09-04
- Need advice on how to get VLLM working with 2xR9700 + 2x7900xtx? -- 2025-09-04
- pwnfuzz/diffrays -- 2025-09-04
- Chromium Hardening Guide -- 2025-09-04
- roomkangali/dursgo -- 2025-09-04
- QuEST/Quartet authors discuss their work on SOTA 4-bit training optimizations -- 2025-09-01
- F-Stack – A network development kit with high performance based on DPDK -- 2025-09-01
- An Empirical Study of Knowledge Distillation for Code Understanding Tasks -- 2025-09-01
- A Comparative Analysis of Vision Language Models for Scientific Data Interpretation -- 2025-08-31
- Sparrow: Custom language model architecture for microcontrollers like the ESP32 -- 2025-08-30
- Password only for this week: Welcome to Hugston -- 2025-08-26
- Prism MCP Rust SDK v0.1.0 - Production-Grade Model Context Protocol Implementation -- 2025-08-26
- Compute Where It Counts: High Quality Sparsely Activated LLMs -- 2025-08-25
- moonshotai/Kimi-K2-Base -- 2025-08-25
- unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF -- 2025-08-25
- BlueLM-2.5-3B Technical Report -- 2025-08-25
- lightx2v/Qwen-Image-Lightning -- 2025-08-25
- Made Chatterbox TTS a bit faster again on CUDA (155it/s on 3090) -- 2025-08-25
- KittenML/KittenTTS -- 2025-08-25
- city96/Qwen-Image-gguf -- 2025-08-23
- Menlo/Lucy-128k -- 2025-08-23
- NVIDIA just accelerated output of OpenAI’s gpt-oss-120B by nearly 2x -- 2025-08-23
- COMponent-Aware Pruning for Accelerated Control Tasks in Latent Space Models -- 2025-08-23
- Speculative decoding in archgw candidate release 0.4.0. Could use feedback, -- 2025-08-16
- Nvidia Tilus: A Tile-Level GPU Kernel Programming Language -- 2025-08-16
- SynapseRoute: An Auto-Route Switching Framework on Dual-State Large Language Model -- 2025-08-16
- New Tool for Finding Why Your LLM Inference is Slow -- 2025-08-14
- I ran OpenAI’s GPT-OSS 20B locally on a 16GB Mac with Ollama — setup, gotchas, and mini demo -- 2025-08-14
- GLM 4.5 Air - Optimizing - Vulkan vs. CUDA? -- 2025-08-14
- GLiNER2: An Efficient Multi-Task Information Extraction System with Schema-Driven Interface -- 2025-08-13
- TextQuests: How Good are LLMs at Text-Based Video Games? -- 2025-08-13
- Mitigate Hallucinations by Fine-tuning gpt-oss-120b with One Example -- 2025-08-10
- uncensored gpt-oss-20b, bf16 and mxfp4 both available -- 2025-08-10
- LGAI-EXAONE/EXAONE-4.0-1.2B -- 2025-08-10
- New Open-Source Text-to-Image Model Just Dropped Qwen-Image (20B MMDiT) by Alibaba! -- 2025-08-10
- Kitten TTS Web Demo -- 2025-08-09
- Show HN: I built a tool to replace capcut audio transcription -- 2025-08-09
- Whispers From The Void, Transcribed With AI -- 2025-08-09
- The Tape Speed Keyboard -- 2025-08-08
- 0.82 um 105 W diode-pumped thulium-doped all silica fiber laser -- 2025-08-06
- GLM-4.5 llama.cpp PR is nearing completion -- 2025-08-05
- glm-4.5-Air appreciation poist - if you have not done so already, give this model a try -- 2025-08-05
- Amazon's AI Coding Revealed a Dirty Little Secret -- 2025-08-02
- On the Interaction of Compressibility and Adversarial Robustness -- 2025-08-02
- realtime-ai/blastoff-llm -- 2025-08-02
- Quantize your own GGUFs the same way as your fav Unsloth Dynamic GGUFs -- 2025-08-01
- unsloth/Qwen3-235B-A22B-Instruct-2507-GGUF -- 2025-08-01
- On the Predictive Power of Representation Dispersion in Language Models -- 2025-08-01
- Wan 2.2 T2V,I2V 14B MoE Models -- 2025-07-31
- PowerInfer/SmallThinker-21BA3B-Instruct -- 2025-07-31
- Ollama + Open WebUI -- is there a way for the same query to run through the same model multiple times (could be 3 times, could be 100 times), then gather all the answers together to summarise/count? -- 2025-07-25
- WGRAMMAR: Leverage Prior Knowledge to Accelerate Structured Decoding -- 2025-07-25
- Semantic chunking using LLMs -- 2025-07-20
- Does the OpenWebUi run the sentence transformer models locally? -- 2025-07-20
- Dataset for structured (JSON) output? -- 2025-07-19
- support for Kimi-K2 has been merged into llama.cpp -- 2025-07-19
- t-tech/T-pro-it-2.0 -- 2025-07-19
- Madness, the ignorant's question. Would it be possible to lighten an LLM model? -- 2025-07-18
- ETH Zurich and EPFL will release a fully open-source LLM developed on public infrastructure. Trained on the “Alps” supercomputer at the Swiss National Supercomputing Centre (CSCS). Trained on 60% english/40% non-english, it will be released in 8B and 70B sizes. -- 2025-07-17
- Moonshot AI’s open source Kimi K2 outperforms GPT-4 in key benchmarks -- 2025-07-17
- Advice Needed: Best way to replace Together API with self-hosted LLM for high-concurrency app -- 2025-07-17
- baidu/ERNIE-4.5-0.3B-PT -- 2025-07-17
- LiquidAI/LFM2-700M -- 2025-07-17
- RekaAI/reka-flash-3.1 · Hugging Face -- 2025-07-17
- Seq vs Seq: the Ettin Suite of Paired Encoders and Decoders -- 2025-07-17
- T5Gemma: A new collection of encoder-decoder Gemma models- Google Developers Blog -- 2025-07-17
- H-Net: a hierarchical network that replaces tokenization with a dynamic chunking process directly inside the model, automatically discovering and operating over meaningful units of data -- 2025-07-17
- How I build software quickly -- 2025-07-16
- RekaAI/reka-flash-3.1 -- 2025-07-15
- What kind of throughput can I expect with Llama 3.1 on a H200? -- 2025-07-15
- MIDI-VALLE: Improving Expressive Piano Performance Synthesis Through Neural Codec Language Modelling -- 2025-07-14
- Replication of Quantum Factorisation Records with an 8-bit Home Computer [pdf] -- 2025-07-14
- Local llms works great! -- 2025-07-12
- LiquidAI/LFM2-350M -- 2025-07-12
- QPART: Adaptive Model Quantization and Dynamic Workload Balancing for Accuracy-aware Edge Inference -- 2025-07-12
- Issues with Qwen 3 Embedding models (4B and 0.6B) -- 2025-07-12
- [Tool Release] Finetune & Quantize 1–3B LLMs on 8GB RAM using LoFT CLI (TinyLlama + QLoRA + llama.cpp) -- 2025-07-11
- Qwen3-8B-BitNet -- 2025-07-11
- Megakernel doubles Llama-1B inference speed for batch size 1 -- 2025-07-09
- Smallest & best OCR model that can read math & code? -- 2025-07-07
- Qwen/WorldPM-72B -- 2025-07-07
- Code single file with multiple LLM models -- 2025-07-07
- Gen-Verse/CURE -- 2025-07-07
- Run Deepseek locally on a 24g GPU: Quantizing on our Giga Computing 6980P Xeon -- 2025-07-01
- I built a document workflow system using VLMs: processes complex docs end-to-end (runs locally!!) -- 2025-07-01
- Jan Nano + Deepseek R1: Combining Remote Reasoning with Local Models using MCP -- 2025-07-01
- Query Classifier for RAG - Save your $$$ and users from irrelevant responses -- 2025-07-01
- Building a memory-heavy AI agent — looking for local-first storage & recall solutions -- 2025-07-01
- Is there any easy way to get up and running with chatgpt-like capabilities at home? -- 2025-07-01
- No recognition of slavic characters. English characters recognized are separate singular characters, not a block of text when using PaddleOCR. -- 2025-07-01
- Tired of copy-pasting from ChatGPT for coding? I am building an open-source tool (Athanor) to fix that - Alpha testers/feedback wanted! -- 2025-07-01
- VideoGameBench from Princeton: Can vision-language models play 90s video games? -- 2025-07-01
- New band surges to 500k listeners on Spotify, but turns out it's AI slop -- 2025-07-01
- How we cut CKEditor's bundle size by 40% -- 2025-07-01
- My VSCode → AI chat website connector extension just got 3 new features! -- 2025-07-01
- 100 Gbps Indoor Access and 4.8 Gbps Outdoor Point-to-Point LiFi Transmission Systems using Laser-based Light Sources -- 2025-06-30
- (0,4) brane box models -- 2025-06-30
- Cursor 1.0 -- 2025-06-30
- Help me design a robust on-prem Llama 3 70B infrastructure for 30 users – Complete hardware/software list wanted -- 2025-06-30
- Jan-nano, a 4B model that can outperform 671B on MCP -- 2025-06-30
- Models that are good and fast at Long Document Processing -- 2025-06-30
- I am making an AI batteries included Web Framework (like Django but for AI) -- 2025-06-30
- [New Features & Better] Tabulens: A Vision-LLM Powered PDF Table Extractor -- 2025-06-30
- I tested 10 LLMs locally on my MacBook Air M1 (8GB RAM!) – Here's what actually works- -- 2025-06-30
- Chatbot without ChatGPT -- 2025-06-30
- What's the best way to save and manage different text files for the models to reference? PRD, cursor rules, tech stack, design reference, etc? -- 2025-06-30
- Bzip2 crate switches from C to 100% Rust -- 2025-06-30
- Litestream: Revamped -- 2025-06-30
- Stop using REST for state synchronization (2024) -- 2025-06-30
- Announcing `mcp-protocol-sdk`: A New Enterprise grade Rust SDK for AI Tool Calling (Model Context Protocol) -- 2025-06-30
- Reinforcement Pre-Training -- 2025-06-29
- unsloth/gemma-3n-E4B-it-GGUF -- 2025-06-29
- chandar-lab/NeoBERT -- 2025-06-29
- tencent/Hunyuan-A13B-Instruct -- 2025-06-27
- maya-research/Veena -- 2025-06-27
- Meet Mistral Devstral, SOTA open model designed specifically for coding agents -- 2025-06-26
- 1.93bit Deepseek R1 0528 beats Claude Sonnet 4 -- 2025-06-26
- DeepSeek R1 05/28 performance on five independent benchmarks -- 2025-06-26
- Few-Shot Examples: Overfitting / Leakage -- 2025-06-26
- Finetune a model to think and use tools -- 2025-06-26
- I need help using open web UI with Ollama. Help installing and getting it running win 11 -- 2025-06-26
- I built/am building a micro-transformer for learning and experimentation -- 2025-06-26
- I shipped more code yesterday with Claude 4 than the last 3 weeks combined -- 2025-06-26
- A deep dive into self-improving AI and the Darwin-Gödel Machine -- 2025-06-26
- 100% of the zeros of the Riemann zeta-function are on the critical line -- 2025-06-25
- 100% of odd hyperelliptic Jacobians have no rational points of small height -- 2025-06-25
- deepseek-ai/DualPipe -- 2025-06-23
- 100 Particles Quantum Heat Engine: Exploring the Impact of Criticality on Efficiency -- 2025-06-23
- 0-Auslander correspondence -- 2025-06-23
- Advanced Time Manipulation with GDB -- 2025-06-21
- Practical SDR: Getting started with software-defined radio -- 2025-06-21
- 1000-10,000 M$_\odot$ Primordial Stars Created the Nitrogen Excess in the Galaxy GS 3073 at $z = 5.55$ -- 2025-06-21
- $0^+$ to $2^+$ neutrinoless double-$β$ decay of $^{76}$Ge, $^{82}$Se, $^{130}$Te and $^{136}$Xe in the microscopic interacting boson model} -- 2025-06-21
- 0-1 laws for pattern occurrences in phylogenetic trees and networks -- 2025-06-20
- 100ps time resolution with thin silicon pixel detectors and a SiGe HBT amplifier -- 2025-06-18
- 0-$\pi$ quantum transition in a carbon nanotube Josephson junction: universal phase dependence and orbital degeneracy -- 2025-06-18
- openbmb/MiniCPM4-8B -- 2025-06-17
- lym00/Wan2.1-T2V-1.3B-Self-Forcing-VACE-Addon-Experiment -- 2025-06-17
- DeepSeek R1 05 28 Tested. It finally happened. The ONLY model to score 100% on everything I threw at it. -- 2025-06-17
- ubergarm/DeepSeek-R1-0528-GGUF -- 2025-06-17
- LLM training on RTX 5090 -- 2025-06-17
- [DEMO] I created a coding agent that can do dynamic, runtime debugging. -- 2025-06-17
- Is anyone productively using Aider and Ollama together? -- 2025-06-17
- For everyone who's still confused by Attention... I made this spreadsheet just for you(FREE) -- 2025-06-17
- What setup/model do you use and what’s your monthly spend? -- 2025-06-17
- Xiaomi released an updated 7B reasoning model and VLM version claiming SOTA for their size -- 2025-06-17
- UPDATE: Inference needs nontrivial amount of PCIe bandwidth (8x RTX 3090 rig, tensor parallelism) -- 2025-06-16
- IQ1_Smol_Boi -- 2025-06-16
- Qwen releases official MLX quants for Qwen3 models in 4 quantization levels: 4bit, 6bit, 8bit, and BF16 -- 2025-06-16
- Seeking Help Setting Up a Local LLM Assistant for TTRPG Worldbuilding + RAG on Windows 11 -- 2025-06-16
- New VS Code Pair Programming Extension, Need Help Testing -- 2025-06-16
- Claude-Trace -- 2025-06-16
- Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training -- 2025-06-16
- best fine tuned local LLM for Github Copilot Agent specificaly -- 2025-06-16
- A Simulation in C++ of Joseph Weizenbaum's 1966 Eliza -- 2025-06-16
- 0D-2D Heterostructure for making very Large Quantum Registers using itinerant Bose-Einstein Condensate of Excitons -- 2025-06-16
- 100-mJ class, sub-two-cycle, carrier-envelope phase-stable dual-chirped optical parametric amplification -- 2025-06-16
- [Update] Rensa: added full CMinHash + OptDensMinHash support (fast MinHash in Rust for dataset deduplication / LLM fine-tuning) -- 2025-06-15
- Open Source Unsiloed AI Chunker (EF2024) -- 2025-06-15
- ether0 - Mistral 24B with RL on several molecular design tasks in chemistry -- 2025-06-15
- Need selfhosted AI to generate better bash scripts and ansible playbooks -- 2025-06-15
- How do I finetune Devstral with vision support? -- 2025-06-15
- What's the best approach for including niche dependency source files and associated documentation reference material in context? -- 2025-06-15
- Airlines Don't Want You to Know They Sold Your Flight Data to DHS -- 2025-06-15
- John Deere Must Face Second Right to Repair Lawsuit -- 2025-06-15
- What vector database and embeddings are y'all using -- 2025-06-15
- Turn based two model critique for rounds to refine answer - any examples or FOSS projects? -- 2025-06-15
- mistralai/Magistral-Small-2506_gguf -- 2025-06-14
- Ruminate: From All-or-Nothing to Just-Right Reasoning in LLMs -- 2025-06-14
- [update] Restructured repo under rvn-tools — modular CLI for LLM formats -- 2025-06-14
- Testing Quant Quality for Shisa V2 405B -- 2025-06-14
- Old model, new implementation -- 2025-06-14
- Ollama vs Llamacpp: Different output for same model -- 2025-06-14
- How to improve my ViT model -- 2025-06-14
- From RPC to transactions and durable executions -- 2025-06-14
- Flattening Rust’s learning curve -- 2025-06-14
- Async from scratch 3: Pinned against the wall -- 2025-06-14
- How to get the most out of my AMD 7900XT? -- 2025-06-14
- Faulty 120W charger analysis (Anker GAN Prime) [video] -- 2025-06-13
- unsloth/Magistral-Small-2506-GGUF -- 2025-06-12
- mistralai/Magistral-Small-2506 -- 2025-06-12
- rednote-hilab/dots.llm1.base -- 2025-06-12
- [Tool] rvn-convert: OSS Rust-based SafeTensors to GGUF v3 converter (single-shard, fast, no Python) -- 2025-06-12
- GuidedQuant: Boost LLM layer-wise PTQ methods using the end loss guidance (Qwen3, Gemma3, Llama3.3 / 2~4bit Quantization) -- 2025-06-12
- I built a memory MCP that understands you (so Sam Altman can't). -- 2025-06-12
- Built an open source desktop app to easily play with local LLMs and MCP -- 2025-06-12
- mtmd : support Qwen 2.5 Omni (input audio+vision, no audio output) by ngxson · Pull Request #13784 · ggml-org/llama.cpp -- 2025-06-12
- i got tired of the errors, so automated debugging using Ollama -- 2025-06-12
- Ablating Gemma 3 27B variants with synthetic data from Sonnet 4 (Few-shot vs LoRA) -- 2025-06-12
- The LLM Gateway gets a major upgrade: becomes a data-plane for Agents. -- 2025-06-12
- Introducing stronger dependencies on systemd -- 2025-06-12
- How we decreased GitLab repo backup times from 48 hours to 41 minutes -- 2025-06-12
- The Quest for 100k - LLAMA.CPP Setting for a Noobie -- 2025-06-12
- Clipjacking: Hacked by copying text – Clickjacking but better -- 2025-06-11
- 0/1 Deep Neural Networks via Block Coordinate Descent -- 2025-06-10
- turbulentdrom/sing-srs-converter -- 2025-06-10
- abi/screenshot-to-code -- 2025-06-10
- Qwen/Qwen3-Embedding-4B -- 2025-06-10
- 100 Gbps Quantum-safe IPsec VPN Tunnels over 46 km Deployed Fiber -- 2025-06-09
- Qwen/Qwen3-Embedding-0.6B -- 2025-06-08
- 0-$π$ qubit in one Josephson junction -- 2025-06-07
- 100-kT Magnetic field generation using paisley targets by femtosecond laser-plasma interactions -- 2025-06-07
- 100 GHz Micrometer compact broadband Monolithic ITO Mach Zehnder Interferometer Modulator enabling 3500 times higher Packing Density -- 2025-06-06
- 0-$\pi$ phase-controllable $thermal$ Josephson junction -- 2025-06-06
- Precomputing Transparency Order in 3D -- 2025-06-06
- ban6cat6/aparecium -- 2025-06-03
- 0-Gaps on 3D Digital Curves -- 2025-06-03
- Reports of Deno's Demise Have Been Greatly Exaggerated -- 2025-06-02
- Comparing Parallel Functional Array Languages: Programming and Performance -- 2025-06-02
- What Every Programmer Should Know About Enumerative Combinatorics -- 2025-06-02
- DuckLake: SQL as a Lakehouse Format -- 2025-05-31
- 1000x Faster Camera and Machine Vision with Ordinary Devices -- 2025-05-31
- 0.75 Gbit/s high-speed classical key distribution with mode-shift keying chaos synchronization of Fabry-Perot lasers -- 2025-05-31
- EdinburghNLP/MMLongBench -- 2025-05-29
- Show HN: Model2vec-Rs – Fast Static Text Embeddings in Rust -- 2025-05-29
- unsloth/DeepSeek-R1-0528-GGUF -- 2025-05-29
- QuantStack/Wan2.1-VACE-14B-GGUF -- 2025-05-29
- 100,000 frames-per-second compressive imaging with a conventional rolling-shutter camera by random point-spread-function engineering -- 2025-05-29
- 1,000-Fold Enhancement of Light-Induced Magnetism in Plasmonic Au Nanoparticles -- 2025-05-29
- nvidia/Llama-3.1-Nemotron-Nano-4B-v1.1 -- 2025-05-28
- Tongyi-Zhiwen/QwenLong-L1-32B -- 2025-05-28
- PKU-DS-LAB/FairyR1-32B -- 2025-05-28
- google/medgemma-4b-pt -- 2025-05-28