Computer Vision

Image recognition, video understanding, multimodal AI

263 articles across 86 editions

Articles

  1. ByteDance Seedance 2.5: one-take creation with flexible referencing -- 2026-09-29
  2. AntLing open sourced the Ming-Image-0.1-Design family -- 2026-09-29
  3. apple/LensVLM-9B · Hugging Face -- 2026-09-29
  4. Can gzip be a language model? -- 2026-09-22
  5. circle-group/hktex -- 2026-09-22
  6. Attention is all you have -- 2026-09-22
  7. [Editorial] arXiv paper 2609.11799 -- 2026-09-14
  8. Open Source Acoustic Drone Detection -- 2026-09-14
  9. QR-Structured Thermal Triggers for Targeted Semantic Attacks on Infrared Vision-Language Models -- 2026-09-11
  10. tencent/EVIE-8B and EVIE-4.5B (High-Capacity Visual Document Retrieval) -- 2026-09-11
  11. SOTA ImageGen Locally NVIDIA Cosmos3(64B) INT4 quants CUDA/MLX -- 2026-09-11
  12. New Music Model YuE2-3B Released! -- 2026-09-11
  13. NeoMME: an efficient Multimodal-native and Multilingual Encoder -- 2026-09-03
  14. InstructMesh: Selective Refinement of Generative 3D Models for Fabrication -- 2026-09-03
  15. krea/Krea-2-Raw -- 2026-09-03
  16. SlopTV: an infinite livestream of AI slop generated from youtube chat comments, Minimax H3 on 2x5090 -- 2026-09-03
  17. Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment -- 2026-08-28
  18. Semantic Browsing: Controllable Diversity for Image Generation -- 2026-08-28
  19. [Editorial] Qwen3.8-27B-Abliterated-SFT -- 2026-08-20
  20. DiffusionGemma Technical Report -- 2026-08-12
  21. LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge -- 2026-08-12
  22. llama.cpp -- 2026-08-12
  23. [Editorial] -- 2026-08-10
  24. dboudreau00/NocTORnal -- 2026-08-10
  25. [Editorial] -- 2026-08-10
  26. [Editorial] -- 2026-08-10
  27. Flux 3 -- 2026-07-24
  28. Grok can now generate 15-second-long videos. -- 2026-07-24
  29. jd-opensource/JoyAI-Echo -- 2026-06-05
  30. Improved techniques for fine-tuning flow models via adjoint matching -- 2026-06-05
  31. unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF -- 2026-06-05
  32. tencent/Hy3-preview -- 2026-06-05
  33. Everything at Every Scale: Scale-Invariant Diffusion with Continuous Super-Resolution -- 2026-06-04
  34. Gaussian Point Splatting -- 2026-06-04
  35. DaVinci Resolve 21 -- 2026-06-04
  36. Profiling in PyTorch (Part 1): A Beginner's Guide to torch.profiler -- 2026-06-04
  37. microsoft/harrier-oss-v1-270m -- 2026-06-04
  38. CohereLabs/tiny-aya-global -- 2026-06-04
  39. HiDream-ai/HiDream-O1-Image -- 2026-06-03
  40. How we index images for RAG -- 2026-06-03
  41. Meta's AI smart glasses and data privacy concerns -- 2026-03-03
  42. [Editorial] AI Search Index -- 2026-03-03
  43. Qwen 3.5 vs Gemini 3 Pro on Screenshot-to-Code: Is the gap finally gone? -- 2026-02-19
  44. [Editorial] The evolution of vision models from CNNs to transformers -- 2026-02-19
  45. Show HN: Offline tiles and routing and geocoding in one Docker Compose stack -- 2026-01-05
  46. ostris/Z-Image-De-Turbo -- 2025-12-12
  47. mistralai/Ministral-3-8B-Instruct-2512 -- 2025-12-12
  48. Making Glasses That Detect Smartglasses -- 2025-12-11
  49. ByteDance-Seed/Depth-Anything-3 -- 2025-12-10
  50. SARLO-80: Worldwide Slant SAR Language Optic Dataset at 80 cm Resolution -- 2025-12-10
  51. Optical Context Compression Is Just (Bad) Autoencoding -- 2025-12-10
  52. seominseok0429/Upsample-Anything-A-Simple-and-Hard-to-Beat-Baseline-for-Feature-Upsampling -- 2025-12-09
  53. Vision-and-Language Navigation with Analogical Textual Descriptions in LLMs -- 2025-12-09
  54. lrzjason/QwenEdit-Anything2Real_Alpha -- 2025-12-08
  55. How Big is Your Video Again? Square vs Rectangular Pixels -- 2025-12-08
  56. ByteDance/BindWeave -- 2025-12-08
  57. 3D Gaussian and Diffusion-Based Gaze Redirection -- 2025-12-04
  58. apple/starflow -- 2025-12-04
  59. princepainter/ComfyUI-PainterLongVideo -- 2025-12-04
  60. OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing -- 2025-12-03
  61. Introducing GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization | "GeoVista is a new 7B open-source agentic model that achieves SOTA performance in geolocalization by integrating visual tools and web search into an RL loop." -- 2025-11-28
  62. facebook/sam-3d-body-dinov3 -- 2025-11-28
  63. Diffusers welcomes FLUX-2 -- 2025-11-26
  64. marinero4972/Open-o3-Video -- 2025-11-17
  65. nanonets/Nanonets-OCR2-3B -- 2025-11-17
  66. nvidia/ChronoEdit-14B-Diffusers-Upscaler-Lora -- 2025-11-17
  67. We ran over 600 image generations to compare AI image models -- 2025-11-13
  68. dx8152/Qwen-Image-Edit-2509-Relight -- 2025-11-13
  69. Last week in Multimodal AI - Local Edition -- 2025-11-12
  70. DeepSeek-OCR GGUF model runs great locally - simple and fast -- 2025-11-12
  71. Qwen3-VL works really good with Zoom-in Tool -- 2025-11-12
  72. lightonai/LightOnOCR-1B-1025 -- 2025-11-12
  73. Qwen/Qwen3-VL-2B-Thinking -- 2025-11-12
  74. Why does Image Recognition work in llama-server but not through Open WebUI? -- 2025-11-06
  75. Has anyone tested ollama on Whisplay HAT with Raspberry pi zero 2W? -- 2025-11-06
  76. allenai/olmOCR-2-7B-1025 -- 2025-11-06
  77. is there simple way like .bat to compress to q4-q8 like Unsloth, Qwen3-VL-30B-A3B-Thinking-abliterated model -- 2025-11-06
  78. meituan-longcat/LongCat-Video -- 2025-11-05
  79. allenai/olmOCR-2-7B-1025-FP8 -- 2025-11-05
  80. Worse Embedding Performance with Qwen 3 VL than with Qwen 2.5 VL? -- 2025-11-05
  81. Retrospective Sparse Attention for Efficient Long-Context Generation -- 2025-11-05
  82. deepseek-ai/DeepSeek-OCR -- 2025-11-04
  83. LiquidAI/LFM2-VL-3B -- 2025-11-04
  84. Qwen/Qwen3-VL-235B-A22B-Thinking -- 2025-11-04
  85. KTransformers Open Source New Era: Local Fine-tuning of Kimi K2 and DeepSeek V3 -- 2025-11-04
  86. F2LLM Technical Report: Matching SOTA Embedding Performance with 6 Million Open-Source Data -- 2025-11-04
  87. [Editorial] https://blog.peerllm.com/2025/11/02/announcing-v0.7.6.html -- 2025-11-04
  88. Myths Programmers Believe about CPU Caches -- 2025-11-04
  89. Pi Zero Powers A Little Indoor Rover -- 2025-11-04
  90. [Editorial] https://commsrisk.com/sms-blaster-and-imsi-catcher-news-from-lebanon-cambodia-switzerland-and-the-philippines/ -- 2025-11-03
  91. An Obscure Military Program Helps Local Cops Buy Armored Card and Spyware -- 2025-11-03
  92. DeepSeek may have found a new way to improve AI’s ability to remember -- 2025-11-02
  93. Qwen/Qwen3-VL-8B-Thinking -- 2025-11-02
  94. nvidia/omnivinci -- 2025-11-02
  95. OpenImagingLab/FlashVSR -- 2025-11-02
  96. Build Your Own Force-Feedback Joystick -- 2025-11-02
  97. DarkBitx/ICRev -- 2025-11-01
  98. dd1100/DiscordRAT -- 2025-11-01
  99. Police used Flock cameras to accuse a woman of theft, she had to prove innocence -- 2025-11-01
  100. ZOZO's Contact Solver for physics-based simulations -- 2025-11-01
  101. valiantcat/Qwen-Image-Edit-MeiTu -- 2025-11-01
  102. ToonComposer: Streamlining Cartoon Production with Generative Post-Keyframing -- 2025-11-01
  103. DeepSeek just released a bombshell AI model (DeepSeek AI) so profound it may be as important as the initial release of ChatGPT-3.5/4 ------ Robots can see-------- And nobody is talking about it -- And it's Open Source - If you take this new OCR Compresion + Graphicacy = Dual-Graphicacy 2.5x improve -- 2025-10-27
  104. Pico Banana: Large-Scale Dataset for Image Editing by Apple -- 2025-10-27
  105. What's the best embedding model for document images ? -- 2025-10-26
  106. AlphaXiv,Compare the Deepseek-OCR and Mistral-OCR OCR models -- 2025-10-26
  107. Open-Bee/Bee-8B-RL -- 2025-10-26
  108. datalab-to/chandra -- 2025-10-26
  109. Unlock the power of images with AI Sheets -- 2025-10-26
  110. dvlab-research/DreamOmni2 -- 2025-10-25
  111. bytetriper/RAE -- 2025-10-25
  112. tencent/POINTS-Reader -- 2025-10-25
  113. Stitch: Training-Free Position Control in Multimodal Diffusion Transformers -- 2025-10-25
  114. mit-han-lab/streaming-vlm -- 2025-10-22
  115. Doby-Xu/WithAnyone -- 2025-10-22
  116. lightx2v/Wan2.2-I2V-A14B-Moe-Distill-Lightx2v -- 2025-10-22
  117. tencent-ailab/SongPrep -- 2025-10-20
  118. opendatalab/MinerU2.5-2509-1.2B -- 2025-10-20
  119. QuantStack/Qwen-Image-Edit-2509-GGUF -- 2025-10-20
  120. LM Studio and VL models -- 2025-10-19
  121. Qwen/Qwen-Image-Edit-2509 -- 2025-10-19
  122. Alpha-VLLM/Lumina-DiMOO -- 2025-10-19
  123. Paper2Video — turn a research paper into a full presentation video (slides, speech, talking head) -- 2025-10-15
  124. Practical OCR with Nanonets OCR2‑3B -- 2025-10-15
  125. neuphonic/neutts-air -- 2025-10-15
  126. Qwen/Qwen3-VL-235B-A22B-Instruct -- 2025-10-15
  127. XiaomiMiMo/MiMo-Audio-Eval -- 2025-10-15
  128. Very interesting! OmniInsert — mask-free video insertion of any reference -- 2025-10-14
  129. facebookresearch/DepthLM_Official -- 2025-10-14
  130. NVlabs/rcm -- 2025-10-11
  131. Divining Air Quality With A Cheap Computer Vision Device -- 2025-10-09
  132. pixai-labs/pixai-tagger-v0.9 -- 2025-10-09
  133. TianDongL/Diffusion_pipe_in_ComfyUI -- 2025-10-09
  134. Project running VLMs on a Pi 5 and NV Jetson Orin Nano -- 2025-10-05
  135. Demo: I made an open-source version of Imagine by Claude (released yesterday) -- 2025-10-05
  136. nunchaku-tech/nunchaku-qwen-image-edit-2509 -- 2025-10-05
  137. cvlab-kaist/VIRAL -- 2025-10-05
  138. deepseek-ai/DeepSeek-V3.2-Exp -- 2025-10-03
  139. moondream/moondream3-preview -- 2025-10-03
  140. Unsupervised Hallucination Detection by Inspecting Reasoning Processes -- 2025-10-03
  141. Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training -- 2025-10-03
  142. Will fine-tuning LLaMA 3.2 11B Instruct on text-only data degrade its vision capabilities? -- 2025-10-03
  143. jmanhype/vggt-mps -- 2025-10-02
  144. openbmb/VoxCPM-0.5B -- 2025-10-02
  145. Comfy-Org/Qwen-Image-Edit_ComfyUI -- 2025-10-02
  146. SOTA OCR on-device with Core ML and dots.ocr -- 2025-10-02
  147. Stress-Testing RAG in Production: Retrieval Quality, Drift, and Hidden Costs -- 2025-09-30
  148. Tencent-Hunyuan/SRPO -- 2025-09-30
  149. lodestones/Chroma1-Base -- 2025-09-30
  150. MV-RAG: Retrieval Augmented Multiview Diffusion -- 2025-09-30
  151. Phantom-video/HuMo -- 2025-09-27
  152. Build Your Own 6K Camera -- 2025-09-27
  153. Training-Free Text-Guided Color Editing with Multi-Modal Diffusion Transformer -- 2025-09-27
  154. OPPOer/Qwen-Image-Pruning -- 2025-09-27
  155. From Research to Reality: Feasibility of Gradient Inversion Attacks in Federated Learning -- 2025-09-17
  156. Renting GPUs is hilariously cheap -- 2025-09-09
  157. Tencent-Hunyuan/HunyuanWorld-Voyager -- 2025-09-09
  158. Shipping textures as PNGs is suboptimal -- 2025-09-09
  159. inclusionAI/UI-Venus -- 2025-09-07
  160. WeChatCV/Stand-In_Preprocessor_ComfyUI -- 2025-09-07
  161. Robotic Canoe Puts Robot Arms to Work -- 2025-09-07
  162. LiquidAI/LFM2-VL-1.6B -- 2025-09-06
  163. A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images -- 2025-09-06
  164. TencentARC/GenCompositor -- 2025-09-06
  165. Doubling Your Data in Minutes: Ultra-fast Tabular Data Generation via LLM-Induced Dependency Graphs -- 2025-08-30
  166. Temporal Point-Supervised Signal Reconstruction: A Human-Annotation-Free Framework for Weak Moving Target Detection -- 2025-08-30
  167. PurinNyova/Image-Detection-Bypass-Utility -- 2025-08-25
  168. An Alternative to Text-to-SQL -- 2025-08-25
  169. Best model for transcribing videos? -- 2025-08-25
  170. Meta released DINO-V3 : SOTA for any Vision task -- 2025-08-21
  171. DatasetAgent: A Novel Multi-Agent System for Auto-Constructing Datasets from Real-World Images -- 2025-08-21
  172. WeChatCV/Stand-In -- 2025-08-21
  173. We built a 12B model that beats Claude 4 Sonnet at video captioning while costing 17x less - fully open source -- 2025-08-15
  174. Francis-Rings/StableAvatar -- 2025-08-15
  175. nvidia/canary-qwen-2.5b -- 2025-08-15
  176. OmniSVG/OmniSVG -- 2025-08-15
  177. Phi-Ground Tech Report: Advancing Perception in GUI Grounding -- 2025-08-15
  178. [UPDATE] DocStrange - Structured data extraction from images/pdfs/docs -- 2025-08-15
  179. NuMarkdown-8B-Thinking - first reasoning OCR VLM -- 2025-08-11
  180. Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling -- 2025-08-11
  181. Vision Language Model Alignment in TRL ⚡️ -- 2025-08-11
  182. How the best image generation models work from the inside ? -- 2025-08-08
  183. AIDC-AI/Ovis-U1-3B -- 2025-08-08
  184. n0xa/SecKC-MHN-Globe -- 2025-08-08
  185. LLM-Based Identification of Infostealer Infection Vectors from Screenshots: The Case of Aurora -- 2025-08-08
  186. Reason ex Machina: Jailbreaking LLMs by Squeezing Their Brains | xayan.nu -- 2025-08-08
  187. Closing the Modality Gap for Mixed Modality Search -- 2025-08-07
  188. [Editorial] Mach-O binary analysis, with a focus on malware analysis and reverse engineering. -- 2025-07-31
  189. petqoo/ROGO -- 2025-07-31
  190. wyhlovecpp/GPT-Image-Edit -- 2025-07-31
  191. Show HN: MoebiusXBIN – ASCII and text-mode art editor with custom font support -- 2025-07-31
  192. boson-ai/higgs-audio -- 2025-07-24
  193. nvidia/canary-qwen-2.5b -- 2025-07-24
  194. TimeScope: How Long Can Your Video Large Multimodal Model Go? -- 2025-07-24
  195. ICML 2025 Outstanding Paper Awards -- 2025-07-24
  196. THUDM/GLM-4.1V-Thinking -- 2025-07-23
  197. FunAudioLLM/ThinkSound -- 2025-07-23
  198. RaphaelLiu/PusaV1 -- 2025-07-23
  199. merve/smol-vision -- 2025-07-23
  200. ChenDarYen/ComfyUI-NAG -- 2025-07-20
  201. Locality-aware Parallel Decoding for Efficient Autoregressive Image Generation -- 2025-07-20
  202. quasiblob/ComfyUI-EsesImageEffectBloom -- 2025-07-18
  203. HiDream-ai/HiDream-E1-1 -- 2025-07-18
  204. Autoregressive Image Generation with Linear Complexity: A Spatial-Aware Decay Perspective -- 2025-07-18
  205. runjiali-rl/vmem -- 2025-07-17
  206. GeoArrow and GeoParquet, and the Future of Geospatial Data Analysis -- 2025-07-15
  207. TextPixs: Glyph-Conditioned Diffusion with Character-Aware Attention and OCR-Guided Supervision -- 2025-07-15
  208. Need advice on how to improve Handwritten Text Recognition of names using Vision models (for academic research purposes) -- 2025-07-14
  209. A fast 3D collision detection algorithm -- 2025-07-11
  210. 1000 African Voices: Advancing inclusive multi-speaker multi-accent speech synthesis -- 2025-07-06
  211. Hack Swaps Keys for Gang Signs, Everyone Gets In -- 2025-07-06
  212. Marker Gene Method : Identifying Stable Solutions in a Dynamic Environment -- 2025-07-06
  213. bytedance/ATI -- 2025-07-01
  214. AIDC-AI/Ovis-U1-3B -- 2025-07-01
  215. google/gemma-3n-E4B-it -- 2025-07-01
  216. baidu/ERNIE-4.5-21B-A3B-PT -- 2025-07-01
  217. bullerwins/FLUX.1-Kontext-dev-GGUF -- 2025-06-28
  218. google/gemma-3n-E2B-it -- 2025-06-28
  219. unsloth/Mistral-Small-3.2-24B-Instruct-2506-GGUF -- 2025-06-26
  220. jinaai/jina-embeddings-v4 -- 2025-06-26
  221. Intelligent-Internet/II-Medical-8B-1706 -- 2025-06-26
  222. Tencent-Hunyuan/HunyuanPortrait -- 2025-06-24
  223. (0,2) hybrid models -- 2025-06-24
  224. 0-th Order Pseudo-differential Operator on the Circle -- 2025-06-24
  225. OmniGen2/OmniGen2 -- 2025-06-24
  226. gdhe17/Self-Forcing -- 2025-06-24
  227. lightx2v/Wan2.1-T2V-14B-StepDistill-CfgDistill -- 2025-06-20
  228. Kijai/WanVideo_comfy -- 2025-06-20
  229. tencent/Hunyuan3D-2.1 -- 2025-06-20
  230. llama-server is cooking! gemma3 27b, 100K context, vision on one 24GB GPU. -- 2025-06-18
  231. I built a lightweight, private, MCP server to share context between AI tools -- 2025-06-18
  232. KVzip: Query-agnostic KV Cache Eviction — 3~4× memory reduction and 2× lower decoding latency -- 2025-06-18
  233. Ollama now supports streaming responses with tool calling -- 2025-06-18
  234. Chainlit or Open webui for production? -- 2025-06-18
  235. Ollama not releasing VRAM after running a model -- 2025-06-18
  236. Help Shape the Future of AI in India - Survey on Local vs Cloud LLM Usage (Developers/Students/AI Enthusiasts) -- 2025-06-18
  237. llmcontext: Attach you whole project in large context chats -- 2025-06-18
  238. Semantic search engine for ArXiv, biorxiv and medrxiv -- 2025-06-18
  239. Oodle 2.9.14 and Intel 13th/14th gen CPUs -- 2025-06-18
  240. A multi-turn tool-calling base model for RL agent training -- 2025-06-18
  241. GUI RAG that can do an unlimited number of documents, or at least many -- 2025-06-18
  242. showlab/OmniConsistency -- 2025-06-15
  243. echo840/MonkeyOCR -- 2025-06-15
  244. inclusionAI/Ming-Lite-Omni -- 2025-06-15
  245. New method for creating large 3D models of urban areas is faster and cheaper -- 2025-06-15
  246. showlab/D-AR -- 2025-06-14
  247. graphdeco-inria/on-the-fly-nvs -- 2025-06-14
  248. YOLO-World: Real-Time Open-Vocabulary Object Detection -- 2025-06-13
  249. 0ptical trapping with optical magnetic field and photonic Hall effect forces -- 2025-06-13
  250. (0,2) Mirror Symmetry on homogeneous Hopf surfaces -- 2025-06-13
  251. 0/1 Deep Neural Networks via Block Coordinate Descent -- 2025-06-10
  252. Hcompany/Holo1-3B -- 2025-06-04
  253. black-forest-labs/FLUX.1-schnell -- 2025-06-04
  254. PlayHT/PlayDiffusion -- 2025-06-04
  255. AMAP-ML/UniVG-R1 -- 2025-06-02
  256. showlab/OmniConsistency -- 2025-06-02
  257. tencent/HunyuanVideo-Avatar -- 2025-06-02
  258. Datadog/Toto-Open-Base-1.0 -- 2025-06-01
  259. 0.08 fF, 0.72 nA dark current, 91% Quantum Efficiency, 38 Gb/s Nano-photodetector on a 45 nm CMOS Silicon-Photonic Platform -- 2025-05-30
  260. 1000 FPS HDR Video With a Spike-RGB Hybrid Camera -- 2025-05-30
  261. 100,000 frames-per-second compressive imaging with a conventional rolling-shutter camera by random point-spread-function engineering -- 2025-05-29
  262. 1,000-Fold Enhancement of Light-Induced Magnetism in Plasmonic Au Nanoparticles -- 2025-05-29
  263. Model-Based Machine Learning (2023) -- 2025-05-28
  264. 0-MMS: Zero-Shot Multi-Motion Segmentation With A Monocular Event Camera -- 2025-05-28
  265. 0.8% Nyquist computational ghost imaging via non-experimental deep learning -- 2025-05-28