Daily Papers of 2026-01-07
- LTX-2: Efficient Joint Audio-Visual Foundation Model 110 upvotes, #1 of 2026-01-07
- InfiniDepth: Arbitrary-Resolution and Fine-Grained Depth Estimation with Neural Implicit Fields 95 upvotes, #2 of 2026-01-07
- MOSS Transcribe Diarize: Accurate Transcription with Speaker Diarization 52 upvotes, #3 of 2026-01-07
- UniCorn: Towards Self-Improving Unified Multimodal Models through Self-Generated Supervision 44 upvotes, #4 of 2026-01-07
- NitroGen: An Open Foundation Model for Generalist Gaming Agents 37 upvotes, #5 of 2026-01-07
- MindWatcher: Toward Smarter Multimodal Tool-Integrated Reasoning 36 upvotes, #6 of 2026-01-07
- SciEvalKit: An Open-source Evaluation Toolkit for Scientific General Intelligence 34 upvotes, #7 of 2026-01-07
- MiMo-V2-Flash Technical Report 31 upvotes, #8 of 2026-01-07
- SOP: A Scalable Online Post-Training System for Vision-Language-Action Models 26 upvotes, #9 of 2026-01-07
- DreamStyle: A Unified Framework for Video Stylization 22 upvotes, #10 of 2026-01-07
- Digital Twin AI: Opportunities and Challenges from Large Language Models to World Models 17 upvotes, #11 of 2026-01-07
- CogFlow: Bridging Perception and Reasoning through Knowledge Internalization for Visual Mathematical Problem Solving 17 upvotes, #11 of 2026-01-07
- WebGym: Scaling Training Environments for Visual Web Agents with Realistic Tasks 15 upvotes, #13 of 2026-01-07
- OpenRT: An Open-Source Red Teaming Framework for Multimodal LLMs 11 upvotes, #14 of 2026-01-07
- Unified Thinker: A General Reasoning Modular Core for Image Generation 7 upvotes, #15 of 2026-01-07
- Muses: Designing, Composing, Generating Nonexistent Fantasy 3D Creatures without Training 6 upvotes, #16 of 2026-01-07
- FFP-300K: Scaling First-Frame Propagation for Generalizable Video Editing 5 upvotes, #17 of 2026-01-07
- Mechanistic Interpretability of Large-Scale Counting in LLMs through a System-2 Strategy 4 upvotes, #18 of 2026-01-07
- Large Reasoning Models Are (Not Yet) Multilingual Latent Reasoners 4 upvotes, #18 of 2026-01-07
- ExposeAnyone: Personalized Audio-to-Expression Diffusion Models Are Robust Zero-Shot Face Forgery Detectors 2 upvotes, #20 of 2026-01-07
- Parallel Latent Reasoning for Sequential Recommendation 2 upvotes, #20 of 2026-01-07
- U-Net-Like Spiking Neural Networks for Single Image Dehazing 1 upvotes, #22 of 2026-01-07
- AceFF: A State-of-the-Art Machine Learning Potential for Small Molecules 1 upvotes, #22 of 2026-01-07
- X-MuTeST: A Multilingual Benchmark for Explainable Hate Speech Detection and A Novel LLM-consulted Explanation Framework 1 upvotes, #22 of 2026-01-07
- The Sonar Moment: Benchmarking Audio-Language Models in Audio Geo-Localization 1 upvotes, #22 of 2026-01-07
- Steerability of Instrumental-Convergence Tendencies in LLMs 1 upvotes, #26 of 2026-01-07
- Doc-PP: Document Policy Preservation Benchmark for Large Vision-Language Models 1 upvotes, #26 of 2026-01-07
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.