Daily Papers of 2026-01-07

  1. LTX-2: Efficient Joint Audio-Visual Foundation Model 110 upvotes, #1 of 2026-01-07
  2. InfiniDepth: Arbitrary-Resolution and Fine-Grained Depth Estimation with Neural Implicit Fields 95 upvotes, #2 of 2026-01-07
  3. MOSS Transcribe Diarize: Accurate Transcription with Speaker Diarization 52 upvotes, #3 of 2026-01-07
  4. UniCorn: Towards Self-Improving Unified Multimodal Models through Self-Generated Supervision 44 upvotes, #4 of 2026-01-07
  5. NitroGen: An Open Foundation Model for Generalist Gaming Agents 37 upvotes, #5 of 2026-01-07
  6. MindWatcher: Toward Smarter Multimodal Tool-Integrated Reasoning 36 upvotes, #6 of 2026-01-07
  7. SciEvalKit: An Open-source Evaluation Toolkit for Scientific General Intelligence 34 upvotes, #7 of 2026-01-07
  8. MiMo-V2-Flash Technical Report 31 upvotes, #8 of 2026-01-07
  9. SOP: A Scalable Online Post-Training System for Vision-Language-Action Models 26 upvotes, #9 of 2026-01-07
  10. DreamStyle: A Unified Framework for Video Stylization 22 upvotes, #10 of 2026-01-07
  11. Digital Twin AI: Opportunities and Challenges from Large Language Models to World Models 17 upvotes, #11 of 2026-01-07
  12. CogFlow: Bridging Perception and Reasoning through Knowledge Internalization for Visual Mathematical Problem Solving 17 upvotes, #11 of 2026-01-07
  13. WebGym: Scaling Training Environments for Visual Web Agents with Realistic Tasks 15 upvotes, #13 of 2026-01-07
  14. OpenRT: An Open-Source Red Teaming Framework for Multimodal LLMs 11 upvotes, #14 of 2026-01-07
  15. Unified Thinker: A General Reasoning Modular Core for Image Generation 7 upvotes, #15 of 2026-01-07
  16. Muses: Designing, Composing, Generating Nonexistent Fantasy 3D Creatures without Training 6 upvotes, #16 of 2026-01-07
  17. FFP-300K: Scaling First-Frame Propagation for Generalizable Video Editing 5 upvotes, #17 of 2026-01-07
  18. Mechanistic Interpretability of Large-Scale Counting in LLMs through a System-2 Strategy 4 upvotes, #18 of 2026-01-07
  19. Large Reasoning Models Are (Not Yet) Multilingual Latent Reasoners 4 upvotes, #18 of 2026-01-07
  20. ExposeAnyone: Personalized Audio-to-Expression Diffusion Models Are Robust Zero-Shot Face Forgery Detectors 2 upvotes, #20 of 2026-01-07
  21. Parallel Latent Reasoning for Sequential Recommendation 2 upvotes, #20 of 2026-01-07
  22. U-Net-Like Spiking Neural Networks for Single Image Dehazing 1 upvotes, #22 of 2026-01-07
  23. AceFF: A State-of-the-Art Machine Learning Potential for Small Molecules 1 upvotes, #22 of 2026-01-07
  24. X-MuTeST: A Multilingual Benchmark for Explainable Hate Speech Detection and A Novel LLM-consulted Explanation Framework 1 upvotes, #22 of 2026-01-07
  25. The Sonar Moment: Benchmarking Audio-Language Models in Audio Geo-Localization 1 upvotes, #22 of 2026-01-07
  26. Steerability of Instrumental-Convergence Tendencies in LLMs 1 upvotes, #26 of 2026-01-07
  27. Doc-PP: Document Policy Preservation Benchmark for Large Vision-Language Models 1 upvotes, #26 of 2026-01-07

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.