Daily Papers of 2025-07-08

  1. MemOS: A Memory OS for AI System 113 upvotes, #1 of 2025-07-08
  2. Should We Still Pretrain Encoders with Masked Language Modeling? 74 upvotes, #2 of 2025-07-08
  3. Agent KB: Leveraging Cross-Domain Experience for Agentic Problem Solving 66 upvotes, #3 of 2025-07-08
  4. 4DSloMo: 4D Reconstruction for High Speed Scene with Asynchronous Capture 39 upvotes, #4 of 2025-07-08
  5. DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge 37 upvotes, #5 of 2025-07-08
  6. Pre-Trained Policy Discriminators are General Reward Models 33 upvotes, #6 of 2025-07-08
  7. Easy Dataset: A Unified and Extensible Framework for Synthesizing LLM Fine-Tuning Data from Unstructured Documents 28 upvotes, #7 of 2025-07-08
  8. StreamDiT: Real-Time Streaming Text-to-Video Generation 27 upvotes, #8 of 2025-07-08
  9. RoboBrain 2.0 Technical Report 25 upvotes, #9 of 2025-07-08
  10. BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset 23 upvotes, #10 of 2025-07-08
  11. RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs 18 upvotes, #11 of 2025-07-08
  12. OmniDraft: A Cross-vocabulary, Online Adaptive Drafter for On-device Speculative Decoding 15 upvotes, #12 of 2025-07-08
  13. On the rankability of visual embeddings 15 upvotes, #12 of 2025-07-08
  14. VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents 15 upvotes, #12 of 2025-07-08
  15. Reviving Cultural Heritage: A Novel Approach for Comprehensive Historical Document Restoration 12 upvotes, #15 of 2025-07-08
  16. Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions 11 upvotes, #16 of 2025-07-08
  17. UnMix-NeRF: Spectral Unmixing Meets Neural Radiance Fields 10 upvotes, #17 of 2025-07-08
  18. Preserving Privacy, Increasing Accessibility, and Reducing Cost: An On-Device Artificial Intelligence Model for Medical Transcription and Note Generation 8 upvotes, #18 of 2025-07-08
  19. PresentAgent: Multimodal Agent for Presentation Video Generation 8 upvotes, #18 of 2025-07-08
  20. ArtifactsBench: Bridging the Visual-Interactive Gap in LLM Code Generation Evaluation 8 upvotes, #18 of 2025-07-08
  21. SeqTex: Generate Mesh Textures in Video Sequence 6 upvotes, #21 of 2025-07-08
  22. R1-RE: Cross-Domain Relationship Extraction with RLVR 6 upvotes, #21 of 2025-07-08
  23. VLAI: A RoBERTa-Based Model for Automated Vulnerability Severity Classification 5 upvotes, #23 of 2025-07-08
  24. Beyond Simple Edits: X-Planner for Complex Instruction-Based Image Editing 5 upvotes, #23 of 2025-07-08
  25. Disambiguation-Centric Finetuning Makes Enterprise Tool-Calling LLMs More Realistic and Less Risky 3 upvotes, #25 of 2025-07-08
  26. MOD-X: A Modular Open Decentralized eXchange Framework proposal for Heterogeneous Interoperable Artificial Agents 2 upvotes, #26 of 2025-07-08
  27. Evaluating LLMs on Real-World Forecasting Against Human Superforecasters 2 upvotes, #26 of 2025-07-08

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.