Voice & Audio

Text-to-speech, speech recognition, voice cloning

105 articles across 46 editions

Articles

  1. [Editorial] Anthropic — Mythos 5 Incident Transcript -- 2026-09-17
  2. [Editorial] Video Feature (OhOmLqR5nN4) -- 2026-09-17
  3. [Editorial] Video Feature (98syxABbUPk) -- 2026-09-17
  4. reddit.com -- 2026-09-17
  5. One note in three: a verified census of three deployed AI scribes, and the instrument that counted it -- 2026-09-01
  6. [Editorial] -- 2026-09-01
  7. Smartphone LED detects hidden cameras with AI -- 2026-09-01
  8. [Editorial] OpenAI: Jalapeño First Results -- 2026-08-26
  9. Apple introduces new Mac Studio with M5 Max and M5 Ultra - up to 512GB of unified memory -- 2026-08-26
  10. I hosted Kimi K3 (2.8T parameters) using 8 B300s. 92 tok/s, $190 per million tokens -- 2026-08-26
  11. We quantized Qwen 3.8 27B and compared the quants on an RTX 6000 -- 2026-08-26
  12. WorldClaw Agentic 3D open-world generation at scale -- 2026-08-14
  13. [Editorial] Reuters Video Report -- 2026-08-14
  14. Scenema Audio Comes to ComfyUI, Runs on 8GB VRAM -- 2026-08-12
  15. CheshireMew/VoxWeave -- 2026-08-12
  16. [Editorial] -- 2026-07-28
  17. [Editorial] -- 2026-07-28
  18. I distilled an 8B teacher into a 0.6B student on my Mac (MLX). The 0.6B went from 36% to 100% on the task, but few-shot prompting actively made it worse. -- 2026-07-28
  19. Fractale-350M-base: memory as trained behaviour instead of long context, a fully open research release -- 2026-07-23
  20. We built NeuTTS-2E, an open-source on-device TTS model with 7 controllable emotions -- 2026-07-23
  21. audio.cpp: 12 audio models (Qwen3-TTS, PocketTTS, VeVo2 etc) in 1 C++/ggml runtime — TTS up to 5x faster than Python on CUDA -- 2026-06-29
  22. ONOTE: Benchmarking Omnimodal Notation Processing for Expert-level Music Intelligence -- 2026-06-29
  23. PP-OCRv6 on Hugging Face: 50-Language OCR from 1.5M to 34.5M Parameters -- 2026-06-29
  24. MolmoMotion: Language-guided 3D motion forecasting -- 2026-06-29
  25. k2-fsa/OmniVoice — High-Quality Voice Cloning TTS for 600+ Languages -- 2026-04-16
  26. Show HN: Sub-500ms latency voice agent from scratch -- 2026-03-05
  27. PKU-YuanGroup/Helios: Real Real-Time Long Video Generation Model -- 2026-03-04
  28. StyleStream: Real-Time Zero-Shot Voice Style Conversion -- 2026-03-04
  29. KokoClone: Kokoro TTS, but it clones voices now -- 2026-03-04
  30. Speech to text via LLM -- 2026-01-16
  31. kyutai-labs/pocket-tts -- 2026-01-16
  32. zai-org/GLM-ASR-Nano-2512 -- 2025-12-12
  33. zai-org/GLM-TTS -- 2025-12-11
  34. openbmb/VoxCPM1.5 -- 2025-12-11
  35. MDAR: A Multi-scene Dynamic Audio Reasoning Benchmark -- 2025-12-04
  36. nvidia/parakeet_realtime_eou_120m-v1 -- 2025-12-03
  37. Qwen/Qwen3-VL-4B-Instruct -- 2025-11-20
  38. Soul-AILab/SoulX-Podcast-1.7B -- 2025-11-20
  39. Last week in Multimodal AI - Local Edition -- 2025-11-12
  40. pnnbao97/VieNeu-TTS -- 2025-11-12
  41. FocalCodec-Stream: Streaming Low-Bitrate Speech Coding via Causal Distillation -- 2025-11-12
  42. GLaDOS TTS finetuning on MLX from the original game files -- 2025-11-04
  43. zeusftk/FTK_CANVAS_AGENT_for_Comfyui -- 2025-11-04
  44. guyyariv/DyPE -- 2025-11-04
  45. Esonhugh/go-rex-java -- 2025-10-27
  46. SuperSonic – SuperCollider's audio engine in a Web AudioWorklet -- 2025-10-27
  47. 3-way FTP: Pushing files around with silly and unusual methods -- 2025-10-27
  48. HRV Gets Home Automation Upgrades -- 2025-10-27
  49. Open source streaming STT (Parakeet + Silero + Pipecat Smart Turn) -- 2025-10-19
  50. Turn ChatGPT into a real-time meeting assistant (via MCP + Apps SDK) -- 2025-10-19
  51. BASICODE: A Bit Like Java, But From The 1980s -- 2025-10-18
  52. Audio transcription with llama.cpp multimodal -- 2025-10-18
  53. I built a fully automated AI podcast generator that connects to ollama -- 2025-10-18
  54. Chinny (iOS/MacOS): offline, on-device voice cloning with an optimized Chatterbox model -- 2025-10-12
  55. herimor/voxtream -- 2025-10-12
  56. microsoft/VibeVoice-Large -- 2025-10-12
  57. chetwinlow1/Ovi -- 2025-10-12
  58. Phr00t/Qwen-Image-Edit-Rapid-AIO -- 2025-10-12
  59. kyomber/CVE-2025-8088 -- 2025-10-08
  60. This Week in Security: CVSS 0, Chwoot, and Not in the Threat Model -- 2025-10-08
  61. I created the cheapest possible AI voice agent (over 30x less expensive than Elevenlabs and OpenAI Realtime). Check out the Github repo below if you want to try it for yourself! -- 2025-10-07
  62. MaximeRivest/maivi -- 2025-10-07
  63. nineninesix/kani-tts-370m -- 2025-10-07
  64. We just open-sourced Kroko ASR: a fast, streaming alternative to Whisper. It’s early days, we’d love testers, feedback, and contributors. -- 2025-10-04
  65. Chaos96/NTPP -- 2025-09-27
  66. We made a new AI interface that is compatible with Ollama -- 2025-09-24
  67. if-ai/ComfyUI_HunyuanVideoFoley -- 2025-09-24
  68. Show HN: Inferencer – Run and deeply control local AI models (macOS release) -- 2025-09-24
  69. tencent/HunyuanWorld-Voyager -- 2025-09-24
  70. FireRedTeam/FireRedTTS2 -- 2025-09-24
  71. OpenBMB/VoxCPM -- 2025-09-22
  72. voicepowered-ai/VibeVoice-finetuning -- 2025-09-22
  73. Why is the name of a wireless mouse hard-coded into Windows Bluetooth drivers? -- 2025-09-17
  74. Qwen3-Coder-480B Q2_K_XL same speed as Qwen3-235b-instruct Q3_K_XL WHY? -- 2025-09-09
  75. Renting GPUs is hilariously cheap -- 2025-09-09
  76. Ex-Miner Turned Local LLM Enthusiast, now I have a Dilemma -- 2025-09-09
  77. Tencent-Hunyuan/HunyuanWorld-Voyager -- 2025-09-09
  78. Smartphone Sensors Unlocked: Turn Your Phone into a Physics Lab -- 2025-09-08
  79. UniSLU: Unified Spoken Language Understanding from Heterogeneous Cross-Task Datasets -- 2025-09-08
  80. Voice cloning -- 2025-09-08
  81. TencentARC/ToonComposer -- 2025-09-04
  82. MeiGen-AI/InfiniteTalk -- 2025-09-04
  83. RELEASED: ComfyUI Wrapper for Microsoft’s new VibeVoice TTS (voice cloning in seconds) -- 2025-09-03
  84. High-Logic/Genie -- 2025-09-03
  85. Has someone used OWebUi with Docling to talk to pdfs with visualizations? -- 2025-09-01
  86. THU-BPM/Omni-SafetyBench -- 2025-09-01
  87. AIDC-AI/Ovis2.5-9B -- 2025-09-01
  88. TTS VibeVoice FastAPI -- 2025-08-30
  89. Microsoft VibeVoice TTS : Open-Sourced, Supports 90 minutes speech, 4 distinct speakers at a time -- 2025-08-29
  90. tencent/HunyuanVideo-Foley -- 2025-08-29
  91. Made Chatterbox TTS a bit faster again on CUDA (155it/s on 3090) -- 2025-08-25
  92. KittenML/KittenTTS -- 2025-08-25
  93. Kitten TTS Web Demo -- 2025-08-09
  94. Show HN: I built a tool to replace capcut audio transcription -- 2025-08-09
  95. Whispers From The Void, Transcribed With AI -- 2025-08-09
  96. kyutai/tts-voices -- 2025-08-08
  97. Explore KittenTTS with Gradio: Easy Text-to-Speech model -- 2025-08-06
  98. Introcuding KokoroDoki a Local, Open-Source and Real-Time TTS. -- 2025-07-19
  99. Voxtral – Frontier open source speech understanding models -- 2025-07-19
  100. AI can now translate brain scans to text -- 2025-07-19
  101. Suggestions to build local voice assistant -- 2025-07-03
  102. google/gemma-3n-E4B -- 2025-07-03
  103. openai/whisper-large-v3 -- 2025-06-23
  104. Audio-Foundation-Models/ConversationTTS -- 2025-06-18
  105. ResembleAI/chatterbox -- 2025-06-06