Image & Video Generation

Diffusion models, Stable Diffusion, ComfyUI, text-to-image/video

322 articles across 114 editions

Articles

  1. reddit.com -- 2026-10-08
  2. reddit.com -- 2026-10-08
  3. reddit.com -- 2026-10-08
  4. ByteDance Seedance 2.5: one-take creation with flexible referencing -- 2026-09-29
  5. AntLing open sourced the Ming-Image-0.1-Design family -- 2026-09-29
  6. apple/LensVLM-9B · Hugging Face -- 2026-09-29
  7. [Editorial] YouTube: N1rjtDs8blY -- 2026-09-24
  8. letorig/video-generator-client -- 2026-09-24
  9. [Editorial] dgreenheck/tidewater (GitHub) -- 2026-09-24
  10. NeoMME: an efficient Multimodal-native and Multilingual Encoder -- 2026-09-03
  11. InstructMesh: Selective Refinement of Generative 3D Models for Fabrication -- 2026-09-03
  12. krea/Krea-2-Raw -- 2026-09-03
  13. SlopTV: an infinite livestream of AI slop generated from youtube chat comments, Minimax H3 on 2x5090 -- 2026-09-03
  14. novel-to-game: Turn any novel into a playable game via 7-skill AI pipeline -- 2026-08-04
  15. [Editorial] Osmantic/ODS -- 2026-08-04
  16. [Editorial] ruvnet/rvQR -- 2026-08-04
  17. [Editorial] Decimen Optical Transfer — Fountain Codes Implementation -- 2026-08-04
  18. [Editorial] Decimen App -- 2026-08-04
  19. Nifer: 700 t/s with Qwen 3.6 35B on RTX 5090 — Cerebras-class local inference -- 2026-07-29
  20. BTL-3 27B: Open agentic coding model that fits in 8.39GB -- 2026-07-29
  21. Gemini Distillation Service — Google now offering distillation-as-a-service -- 2026-07-29
  22. Flux 3 -- 2026-07-24
  23. Grok can now generate 15-second-long videos. -- 2026-07-24
  24. I didn't give up - extGemma4-40_5B returned -- 2026-07-14
  25. J-Space Hallucination Signal Stress-Tested Across 7 Datasets on Qwen3-4B -- 2026-07-14
  26. audio.cpp: 12 audio models (Qwen3-TTS, PocketTTS, VeVo2 etc) in 1 C++/ggml runtime — TTS up to 5x faster than Python on CUDA -- 2026-06-29
  27. ONOTE: Benchmarking Omnimodal Notation Processing for Expert-level Music Intelligence -- 2026-06-29
  28. PP-OCRv6 on Hugging Face: 50-Language OCR from 1.5M to 34.5M Parameters -- 2026-06-29
  29. MolmoMotion: Language-guided 3D motion forecasting -- 2026-06-29
  30. jd-opensource/JoyAI-Echo -- 2026-06-05
  31. Improved techniques for fine-tuning flow models via adjoint matching -- 2026-06-05
  32. unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF -- 2026-06-05
  33. tencent/Hy3-preview -- 2026-06-05
  34. Running DeepSeek-V4 locally with 4x legacy RTX 2080 Ti ($2k budget setup). Custom Turing kernels, W8A8 quantization, and 255 prefill tok/s! -- 2026-05-20
  35. Ran the same models across Strix Halo, RTX 3090, and RTX 5070 because I wanted my own numbers -- 2026-05-20
  36. Intel's Crescent Island PCB Leaks, Showing a Massive Xe3P GPU, 16-Pin Connector, 160GB LPDDR5X as Intel Sidesteps the HBM Shortage -- 2026-05-20
  37. Sipeed's K3 RISC-V SBCs can run 30B-parameter LLMs 60 TOPS (INT4), Supports BF16/FP16/INT4 -- 2026-05-20
  38. club-5060ti: practical RTX 5060 Ti local LLM notes and configs -- 2026-05-20
  39. ScioMind: Cognitively Grounded Multi-Agent Social Simulation with Anchoring-Based Belief Dynamics -- 2026-05-15
  40. [MIT] RLCR: Teaching AI models to say "I'm not sure" -- 2026-05-15
  41. CDM: Continuous-Time Distribution Matching for Few-Step Diffusion Distillation -- 2026-05-15
  42. Built an open-source one-prompt-to-cinematic-reel pipeline on a single GPU — FLUX.2 + Wan2.2 + vision critic + music + 9-language narration -- 2026-05-15
  43. HumeAI/tada-3b-ml -- 2026-05-15
  44. PRX Part 3 — Training a Text-to-Image Model in 24h! -- 2026-03-12
  45. New LTX2.3 Tool for OpenWebui -- 2026-03-12
  46. Kotlin creator's new language: a formal way to talk to LLMs instead of English -- 2026-03-12
  47. Building a TB-303 from Scratch -- 2026-03-12
  48. PKU-YuanGroup/Helios: Real Real-Time Long Video Generation Model -- 2026-03-04
  49. StyleStream: Real-Time Zero-Shot Voice Style Conversion -- 2026-03-04
  50. KokoClone: Kokoro TTS, but it clones voices now -- 2026-03-04
  51. Ling-2.5-1T: 1T Parameter Open-Source Instant Model with 1M Context -- 2026-02-16
  52. Qwen3.5-397B-A17B Unsloth GGUFs — Run on Consumer Hardware -- 2026-02-16
  53. Running Qwen3-Coder-Next 80B on 8GB VRAM — 300x Speedup via Custom Expert Caching -- 2026-02-16
  54. Flame Graphs vs Tree Maps vs Sunburst (2017) -- 2025-12-31
  55. 39C3: Recreating Sandstorm -- 2025-12-31
  56. EditMGT — fast, localized image editing with Masked Generative Transformers -- 2025-12-30
  57. Francis-Rings/FlashPortrait -- 2025-12-30
  58. zai-org/GLM-TTS -- 2025-12-30
  59. Flowception: Temporally Expansive Flow Matching for Video Generation -- 2025-12-30
  60. Key Highlights of NVIDIA’s New Model: Nemotron 3 -- 2025-12-17
  61. The Open Evaluation Standard: Benchmarking NVIDIA Nemotron 3 Nano with NeMo Evaluator -- 2025-12-17
  62. ostris/Z-Image-De-Turbo -- 2025-12-12
  63. zai-org/GLM-TTS -- 2025-12-11
  64. openbmb/VoxCPM1.5 -- 2025-12-11
  65. ByteDance-Seed/Depth-Anything-3 -- 2025-12-10
  66. seominseok0429/Upsample-Anything-A-Simple-and-Hard-to-Beat-Baseline-for-Feature-Upsampling -- 2025-12-09
  67. lrzjason/QwenEdit-Anything2Real_Alpha -- 2025-12-08
  68. How Big is Your Video Again? Square vs Rectangular Pixels -- 2025-12-08
  69. shubh-io/DockMate -- 2025-12-08
  70. Comfy-Org/HunyuanVideo_1.5_repackaged -- 2025-12-08
  71. ByteDance/BindWeave -- 2025-12-08
  72. apple/starflow -- 2025-12-04
  73. princepainter/ComfyUI-PainterLongVideo -- 2025-12-04
  74. OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing -- 2025-12-03
  75. Z-Image: Powerful and highly efficient image generation model with 6B parameters -- 2025-12-03
  76. FLUX.2: Frontier Visual Intelligence -- 2025-11-28
  77. Diffusers welcomes FLUX-2 -- 2025-11-26
  78. dx8152/Relight -- 2025-11-26
  79. Question About Motherboards -- 2025-11-26
  80. Qwen/Qwen3-VL-4B-Instruct -- 2025-11-20
  81. Soul-AILab/SoulX-Podcast-1.7B -- 2025-11-20
  82. wildminder/ComfyUI-DyPE -- 2025-11-19
  83. lightx2v/Autoencoders -- 2025-11-19
  84. We ran over 600 image generations to compare AI image models -- 2025-11-13
  85. dx8152/Qwen-Image-Edit-2509-Relight -- 2025-11-13
  86. meituan-longcat/LongCat-Video -- 2025-11-05
  87. allenai/olmOCR-2-7B-1025-FP8 -- 2025-11-05
  88. deepseek-ai/DeepSeek-OCR -- 2025-11-04
  89. LiquidAI/LFM2-VL-3B -- 2025-11-04
  90. Qwen/Qwen3-VL-235B-A22B-Thinking -- 2025-11-04
  91. DeepSeek may have found a new way to improve AI’s ability to remember -- 2025-11-02
  92. Qwen/Qwen3-VL-8B-Thinking -- 2025-11-02
  93. nvidia/omnivinci -- 2025-11-02
  94. OpenImagingLab/FlashVSR -- 2025-11-02
  95. Build Your Own Force-Feedback Joystick -- 2025-11-02
  96. ZOZO's Contact Solver for physics-based simulations -- 2025-11-01
  97. valiantcat/Qwen-Image-Edit-MeiTu -- 2025-11-01
  98. ToonComposer: Streamlining Cartoon Production with Generative Post-Keyframing -- 2025-11-01
  99. fireleyfreya/AI-Art-Generator -- 2025-10-30
  100. krea/krea-realtime-video -- 2025-10-30
  101. Cerebras REAP'd GLM4.6: 25%, 30%, 40% pruned FP8 checkpoints on HF! -- 2025-10-28
  102. Qwen/Qwen3-VL-30B-A3B-Instruct-FP8 -- 2025-10-28
  103. lightx2v/Wan2.2-Distill-Loras -- 2025-10-28
  104. GPT-OSS-20b TAKE THE HELM! Further experiments in autopilot. -- 2025-10-28
  105. 5060ti chads... ram overclocking, the phantom menace -- 2025-10-28
  106. Batch inference locally on 4080 -- 2025-10-28
  107. DeepSeek just released a bombshell AI model (DeepSeek AI) so profound it may be as important as the initial release of ChatGPT-3.5/4 ------ Robots can see-------- And nobody is talking about it -- And it's Open Source - If you take this new OCR Compresion + Graphicacy = Dual-Graphicacy 2.5x improve -- 2025-10-27
  108. Pico Banana: Large-Scale Dataset for Image Editing by Apple -- 2025-10-27
  109. dvlab-research/DreamOmni2 -- 2025-10-25
  110. bytetriper/RAE -- 2025-10-25
  111. tencent/POINTS-Reader -- 2025-10-25
  112. Stitch: Training-Free Position Control in Multimodal Diffusion Transformers -- 2025-10-25
  113. Llama.cpp is looking for M5 Neural Accelerator performance testers -- 2025-10-24
  114. NVIDIA sent me a 5090 so I can demo Qwen3-VL GGUF -- 2025-10-24
  115. AMD ROCm 7.9 and dwindling GPU support -- 2025-10-24
  116. Show HN: Cuq – Formal Verification of Rust GPU Kernels -- 2025-10-24
  117. [Editorial] https://github.com/DrewThomasson/ebook2audiobook -- 2025-10-23
  118. [Editorial] https://github.com/lfnovo/open-notebook -- 2025-10-23
  119. tencent/Hunyuan3D-Omni -- 2025-10-23
  120. Doby-Xu/WithAnyone -- 2025-10-22
  121. lightx2v/Wan2.2-I2V-A14B-Moe-Distill-Lightx2v -- 2025-10-22
  122. mit-han-lab/streaming-vlm -- 2025-10-22
  123. tencent-ailab/SongPrep -- 2025-10-20
  124. opendatalab/MinerU2.5-2509-1.2B -- 2025-10-20
  125. QuantStack/Qwen-Image-Edit-2509-GGUF -- 2025-10-20
  126. LM Studio and VL models -- 2025-10-19
  127. Qwen/Qwen-Image-Edit-2509 -- 2025-10-19
  128. Alpha-VLLM/Lumina-DiMOO -- 2025-10-19
  129. linkedlist771/SoraWatermarkCleaner -- 2025-10-19
  130. Audio transcription with llama.cpp multimodal -- 2025-10-18
  131. I built a fully automated AI podcast generator that connects to ollama -- 2025-10-18
  132. Paper2Video — turn a research paper into a full presentation video (slides, speech, talking head) -- 2025-10-15
  133. Practical OCR with Nanonets OCR2‑3B -- 2025-10-15
  134. neuphonic/neutts-air -- 2025-10-15
  135. Qwen/Qwen3-VL-235B-A22B-Instruct -- 2025-10-15
  136. XiaomiMiMo/MiMo-Audio-Eval -- 2025-10-15
  137. Very interesting! OmniInsert — mask-free video insertion of any reference -- 2025-10-14
  138. facebookresearch/DepthLM_Official -- 2025-10-14
  139. Built a 1288x RTFx Parakeet Speech-to-Text server... Enjoy! -- 2025-10-13
  140. Novel OpenGL Pixel Shader Dewarping -- 2025-10-13
  141. lovis93/next-scene-qwen-image-lora-2509 -- 2025-10-13
  142. Chinny (iOS/MacOS): offline, on-device voice cloning with an optimized Chatterbox model -- 2025-10-12
  143. herimor/voxtream -- 2025-10-12
  144. microsoft/VibeVoice-Large -- 2025-10-12
  145. chetwinlow1/Ovi -- 2025-10-12
  146. Phr00t/Qwen-Image-Edit-Rapid-AIO -- 2025-10-12
  147. NVlabs/rcm -- 2025-10-11
  148. Qwen3-VL-30B-A3B-Thinking GGUF with llama.cpp patch to run it -- 2025-10-10
  149. Does quantization need training data and will it lower performance for task outside of training data? -- 2025-10-10
  150. Qwen/Qwen3-VL-30B-A3B-Instruct -- 2025-10-10
  151. Project running VLMs on a Pi 5 and NV Jetson Orin Nano -- 2025-10-05
  152. Demo: I made an open-source version of Imagine by Claude (released yesterday) -- 2025-10-05
  153. nunchaku-tech/nunchaku-qwen-image-edit-2509 -- 2025-10-05
  154. cvlab-kaist/VIRAL -- 2025-10-05
  155. Tencent-Hunyuan/Hunyuan3D-Omni -- 2025-10-04
  156. tencent/HunyuanImage-3.0 -- 2025-10-04
  157. For llama.cpp/ggml AMD MI50s are now universally faster than NVIDIA P40s -- 2025-10-03
  158. MSI EdgeXpert Compact AI Supercomputer Based on NVIDIA DGX Spark -- 2025-10-03
  159. Kairos: Immutable Distro for K8s at the Edge -- 2025-10-03
  160. Nvidia Has Been Supplying NDA'ed Docs to Red Hat for Helping NVK Driver -- 2025-10-03
  161. Mini Laptop Needs Custom Kernel -- 2025-10-03
  162. jmanhype/vggt-mps -- 2025-10-02
  163. openbmb/VoxCPM-0.5B -- 2025-10-02
  164. Comfy-Org/Qwen-Image-Edit_ComfyUI -- 2025-10-02
  165. SOTA OCR on-device with Core ML and dots.ocr -- 2025-10-02
  166. Tencent-Hunyuan/SRPO -- 2025-09-30
  167. lodestones/Chroma1-Base -- 2025-09-30
  168. MV-RAG: Retrieval Augmented Multiview Diffusion -- 2025-09-30
  169. Phantom-video/HuMo -- 2025-09-27
  170. Build Your Own 6K Camera -- 2025-09-27
  171. Training-Free Text-Guided Color Editing with Multi-Modal Diffusion Transformer -- 2025-09-27
  172. OPPOer/Qwen-Image-Pruning -- 2025-09-27
  173. We made a new AI interface that is compatible with Ollama -- 2025-09-24
  174. if-ai/ComfyUI_HunyuanVideoFoley -- 2025-09-24
  175. Show HN: Inferencer – Run and deeply control local AI models (macOS release) -- 2025-09-24
  176. tencent/HunyuanWorld-Voyager -- 2025-09-24
  177. FireRedTeam/FireRedTTS2 -- 2025-09-24
  178. Wan-AI/Wan2.2-Animate-14B -- 2025-09-22
  179. decart-ai/Lucy-Edit-Dev -- 2025-09-22
  180. OpenBMB/VoxCPM -- 2025-09-22
  181. voicepowered-ai/VibeVoice-finetuning -- 2025-09-22
  182. alibaba-pai/Wan2.2-VACE-Fun-A14B -- 2025-09-21
  183. VideoGuard: Protecting Video Content from Unauthorized Editing -- 2025-09-21
  184. zli12321/Vision-SR1 -- 2025-09-19
  185. lrzjason/Comfyui-QwenEditUtils -- 2025-09-19
  186. Mini-o3/Mini-o3 -- 2025-09-19
  187. The AI-Scraping Free-for-All Is Coming to an End -- 2025-09-18
  188. Visible Watermarking with Gradio -- 2025-09-18
  189. Tencent-Hunyuan/HunyuanImage-2.1 -- 2025-09-17
  190. xiaomi-research/q-frame -- 2025-09-17
  191. TencentCloudADP/youtu-graphrag -- 2025-09-17
  192. Renting GPUs is hilariously cheap -- 2025-09-09
  193. Tencent-Hunyuan/HunyuanWorld-Voyager -- 2025-09-09
  194. Shipping textures as PNGs is suboptimal -- 2025-09-09
  195. LiquidAI/LFM2-VL-1.6B -- 2025-09-06
  196. A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images -- 2025-09-06
  197. TencentARC/GenCompositor -- 2025-09-06
  198. TencentARC/ToonComposer -- 2025-09-04
  199. MeiGen-AI/InfiniteTalk -- 2025-09-04
  200. OWUI_File_Gen_Export v0.2.0 is out ! -- 2025-09-04
  201. MCP File Generation tool -- 2025-09-04
  202. tencent/Hunyuan-GameCraft-1.0 -- 2025-09-03
  203. InternVL 3.5 released : Best Open-Sourced Multi-Modal LLM, Ranks 3 overall -- 2025-08-31
  204. HunyuanVideo-Foley is out, an open source text-video-to-audio model -- 2025-08-31
  205. peteromallet/Flux-Kontext-InScene -- 2025-08-31
  206. TTS VibeVoice FastAPI -- 2025-08-30
  207. Microsoft VibeVoice TTS : Open-Sourced, Supports 90 minutes speech, 4 distinct speakers at a time -- 2025-08-29
  208. tencent/HunyuanVideo-Foley -- 2025-08-29
  209. bullerwins/Wan2.2-I2V-A14B-GGUF -- 2025-08-28
  210. QuantStack/Qwen-Image-Edit-GGUF -- 2025-08-28
  211. dvlab-research/MGM-Omni -- 2025-08-28
  212. RTX PRO 6000 MAX-Q Blackwell for LLM -- 2025-08-28
  213. Gemini 2.5 Flash Image -- 2025-08-27
  214. Arrexel/pattern-diffusion -- 2025-08-27
  215. An Alternative to Text-to-SQL -- 2025-08-25
  216. Best model for transcribing videos? -- 2025-08-25
  217. Compute Where It Counts: High Quality Sparsely Activated LLMs -- 2025-08-25
  218. moonshotai/Kimi-K2-Base -- 2025-08-25
  219. unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF -- 2025-08-25
  220. BlueLM-2.5-3B Technical Report -- 2025-08-25
  221. lightx2v/Qwen-Image-Lightning -- 2025-08-25
  222. Qwen-Image-Edit #6 overall on LMArena, best open model image editor -- 2025-08-24
  223. flybirdxx/ComfyUI-SDMatte -- 2025-08-24
  224. Wan-AI/Wan2.2-TI2V-5B -- 2025-08-24
  225. HiDream-ai/HiDream-E1-1 -- 2025-08-24
  226. We built a 12B model that beats Claude 4 Sonnet at video captioning while costing 17x less - fully open source -- 2025-08-15
  227. Francis-Rings/StableAvatar -- 2025-08-15
  228. SparcStation 1+ Finally Gets Attention -- 2025-08-15
  229. OmniSVG/OmniSVG -- 2025-08-15
  230. Phi-Ground Tech Report: Advancing Perception in GUI Grounding -- 2025-08-15
  231. NuMarkdown-8B-Thinking - first reasoning OCR VLM -- 2025-08-11
  232. Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling -- 2025-08-11
  233. Vision Language Model Alignment in TRL ⚡️ -- 2025-08-11
  234. Explore KittenTTS with Gradio: Easy Text-to-Speech model -- 2025-08-06
  235. [Editorial] AI personality -- 2025-08-05
  236. CUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement Learning -- 2025-08-05
  237. internlm/Intern-S1 -- 2025-08-05
  238. peteromallet/Flux-Kontext-InScene -- 2025-08-02
  239. A Dual-Screen Cyberdeck To Rule Them All -- 2025-08-02
  240. ziangcao0312/PhysX-3D -- 2025-07-29
  241. Playtron's Linux-Based GameOS Hits the Road with 1.0 -- 2025-07-26
  242. Remembering Chiptunes, the Demoscene and the Illegal Music of Keygens -- 2025-07-26
  243. albozes/shotbuddy -- 2025-07-25
  244. uttam-li/dfs -- 2025-07-25
  245. OmniSVG/OmniSVG -- 2025-07-25
  246. Freezer Monitoring: Because Ice Cream Is a Dish Best Served Cold -- 2025-07-25
  247. Fast LoRA inference for Flux with Diffusers and PEFT -- 2025-07-25
  248. boson-ai/higgs-audio -- 2025-07-24
  249. nvidia/canary-qwen-2.5b -- 2025-07-24
  250. TimeScope: How Long Can Your Video Large Multimodal Model Go? -- 2025-07-24
  251. THUDM/GLM-4.1V-Thinking -- 2025-07-23
  252. FunAudioLLM/ThinkSound -- 2025-07-23
  253. RaphaelLiu/PusaV1 -- 2025-07-23
  254. merve/smol-vision -- 2025-07-23
  255. Skywork/Skywork-R1V3-38B -- 2025-07-20
  256. ByteDance-Seed/Seed-X-PPO-7B -- 2025-07-20
  257. ChenDarYen/ComfyUI-NAG -- 2025-07-20
  258. Locality-aware Parallel Decoding for Efficient Autoregressive Image Generation -- 2025-07-20
  259. Introcuding KokoroDoki a Local, Open-Source and Real-Time TTS. -- 2025-07-19
  260. Voxtral – Frontier open source speech understanding models -- 2025-07-19
  261. AI can now translate brain scans to text -- 2025-07-19
  262. quasiblob/ComfyUI-EsesImageEffectBloom -- 2025-07-18
  263. HiDream-ai/HiDream-E1-1 -- 2025-07-18
  264. Autoregressive Image Generation with Linear Complexity: A Spatial-Aware Decay Perspective -- 2025-07-18
  265. runjiali-rl/vmem -- 2025-07-17
  266. Exploring State-Space-Model based Language Model in Music Generation -- 2025-07-16
  267. TextPixs: Glyph-Conditioned Diffusion with Character-Aware Attention and OCR-Guided Supervision -- 2025-07-15
  268. MIDI-VALLE: Improving Expressive Piano Performance Synthesis Through Neural Codec Language Modelling -- 2025-07-14
  269. Need advice on how to improve Handwritten Text Recognition of names using Vision models (for academic research purposes) -- 2025-07-14
  270. DLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive Caching -- 2025-07-08
  271. Efficient MultiModal Data Pipeline -- 2025-07-08
  272. black-forest-labs/FLUX.1-Kontext-dev-onnx -- 2025-07-07
  273. Subpixel Rendering For Impossibly Small Terminal Text -- 2025-07-04
  274. bytedance/ATI -- 2025-07-01
  275. AIDC-AI/Ovis-U1-3B -- 2025-07-01
  276. google/gemma-3n-E4B-it -- 2025-07-01
  277. baidu/ERNIE-4.5-21B-A3B-PT -- 2025-07-01
  278. THU-KEG/LongWriter-Zero-32B -- 2025-06-30
  279. Tencent-Hunyuan/Hunyuan3D-2.1 -- 2025-06-29
  280. bullerwins/FLUX.1-Kontext-dev-GGUF -- 2025-06-28
  281. google/gemma-3n-E2B-it -- 2025-06-28
  282. black-forest-labs/FLUX.1-Kontext-dev -- 2025-06-27
  283. 0.71-{\AA} resolution electron tomography enabled by deep learning aided information recovery -- 2025-06-26
  284. MeiGen-AI/MeiGen-MultiTalk -- 2025-06-26
  285. Tencent-Hunyuan/HunyuanPortrait -- 2025-06-24
  286. (0,2) hybrid models -- 2025-06-24
  287. 0-th Order Pseudo-differential Operator on the Circle -- 2025-06-24
  288. OmniGen2/OmniGen2 -- 2025-06-24
  289. gdhe17/Self-Forcing -- 2025-06-24
  290. Voyager: Real-Time Splatting City-Scale 3D Gaussians on Your Phone -- 2025-06-21
  291. lightx2v/Wan2.1-T2V-14B-StepDistill-CfgDistill -- 2025-06-20
  292. Kijai/WanVideo_comfy -- 2025-06-20
  293. tencent/Hunyuan3D-2.1 -- 2025-06-20
  294. inclusionAI/Ming-Lite-Omni -- 2025-06-15
  295. New method for creating large 3D models of urban areas is faster and cheaper -- 2025-06-15
  296. vrgamedevgirl84/Wan14BT2VFusioniX -- 2025-06-13
  297. rusjoan/streamcrypt -- 2025-06-12
  298. tang-bd/fuse-dit -- 2025-06-12
  299. Show HN: 3DGS implementation in Nvidia Warp: clean, minimal, runs on CPU and GPU -- 2025-06-12
  300. 0.75 atoms improve the clock signal of 10,000 atoms -- 2025-06-12
  301. After Deepfaking YouTube, Google's Veo 3 Could Slop-Ify Video Games Next -- 2025-06-11
  302. manycore-research/SpatialLM -- 2025-06-09
  303. XiaomiMiMo/MiMo-VL-7B-RL -- 2025-06-09
  304. fishaudio/openaudio-s1-mini -- 2025-06-09
  305. rednote-hilab/dots.llm1.inst -- 2025-06-08
  306. tencent/HunyuanPortrait -- 2025-06-08
  307. Better quantization: Yet Another Quantization Algorithm -- 2025-06-08
  308. MCP server to connect LLM agents to any database -- 2025-06-08
  309. new gemma3 abliterated models from mlabonne -- 2025-06-08
  310. Sharing my a demo of tool for easy handwritten fine-tuning dataset creation! -- 2025-06-08
  311. Yess! Open-source strikes back! This is the closest I've seen anything come to competing with @GoogleDeepMind 's Veo 3 native audio and character motion. -- 2025-06-08
  312. For task-specific agents use task-specific LLMs for routing and hand off - NOT semantic techniques. -- 2025-06-08
  313. Face Age Prediction – Achieved Human-Level Accuracy (MAE ≈ 5) -- 2025-06-08
  314. Is there any open source project leveraging genAI to run quality checks on tabular data ? -- 2025-06-08
  315. Precomputing Transparency Order in 3D -- 2025-06-06
  316. nvidia/Llama-3.1-Nemotron-Nano-VL-8B-V1 -- 2025-06-05
  317. hexgrad/Kokoro-82M -- 2025-06-05
  318. Qwen/Qwen3-Embedding-0.6B-GGUF -- 2025-06-05
  319. AMAP-ML/UniVG-R1 -- 2025-06-02
  320. showlab/OmniConsistency -- 2025-06-02
  321. tencent/HunyuanVideo-Avatar -- 2025-06-02
  322. Datadog/Toto-Open-Base-1.0 -- 2025-06-01
  323. facebook/KernelLLM -- 2025-05-31
  324. 0.08 fF, 0.72 nA dark current, 91% Quantum Efficiency, 38 Gb/s Nano-photodetector on a 45 nm CMOS Silicon-Photonic Platform -- 2025-05-30
  325. 1000 FPS HDR Video With a Spike-RGB Hybrid Camera -- 2025-05-30