Daily Papers of 2024-10-04

  1. Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models 51 upvotes, #1 of 2024-10-04
  2. SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration 44 upvotes, #2 of 2024-10-04
  3. Depth Pro: Sharp Monocular Metric Depth in Less Than a Second 36 upvotes, #3 of 2024-10-04
  4. Loong: Generating Minute-level Long Videos with Autoregressive Language Models 35 upvotes, #4 of 2024-10-04
  5. Video Instruction Tuning With Synthetic Data 33 upvotes, #5 of 2024-10-04
  6. LLaVA-Critic: Learning to Evaluate Multimodal Models 31 upvotes, #6 of 2024-10-04
  7. Contrastive Localized Language-Image Pre-Training 28 upvotes, #7 of 2024-10-04
  8. Large Language Models as Markov Chains 27 upvotes, #8 of 2024-10-04
  9. Eliminating Oversaturation and Artifacts of High Guidance Scales in Diffusion Models 23 upvotes, #9 of 2024-10-04
  10. VinePPO: Unlocking RL Potential For LLM Reasoning Through Refined Credit Assignment 21 upvotes, #10 of 2024-10-04
  11. Distilling an End-to-End Voice Assistant Without Instruction Training Data 21 upvotes, #10 of 2024-10-04
  12. CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet Upcycling 17 upvotes, #12 of 2024-10-04
  13. Contextual Document Embeddings 14 upvotes, #13 of 2024-10-04
  14. Training Language Models on Synthetic Edit Sequences Improves Code Synthesis 12 upvotes, #14 of 2024-10-04
  15. L-CiteEval: Do Long-Context Models Truly Leverage Context for Responding? 10 upvotes, #15 of 2024-10-04
  16. MedVisionLlama: Leveraging Pre-Trained Large Language Model Layers to Enhance Medical Image Segmentation 9 upvotes, #16 of 2024-10-04
  17. Improving Autonomous AI Agents with Reflective Tree Search and Self-Learning 8 upvotes, #17 of 2024-10-04
  18. MVGS: Multi-view-regulated Gaussian Splatting for Novel View Synthesis 8 upvotes, #17 of 2024-10-04
  19. Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations 8 upvotes, #17 of 2024-10-04
  20. Vinoground: Scrutinizing LMMs over Dense Temporal Reasoning with Short Videos 7 upvotes, #20 of 2024-10-04
  21. Open-RAG: Enhanced Retrieval-Augmented Reasoning with Open-Source Large Language Models 6 upvotes, #21 of 2024-10-04
  22. Robin3D: Improving 3D Large Language Model via Robust Instruction Tuning 5 upvotes, #22 of 2024-10-04
  23. Layer Swapping for Zero-Shot Cross-Lingual Transfer in Large Language Models 5 upvotes, #22 of 2024-10-04
  24. Intelligence at the Edge of Chaos 5 upvotes, #22 of 2024-10-04
  25. SciPrompt: Knowledge-augmented Prompting for Fine-grained Categorization of Scientific Topics 4 upvotes, #25 of 2024-10-04
  26. Synthio: Augmenting Small-Scale Audio Classification Datasets with Synthetic Data 4 upvotes, #25 of 2024-10-04
  27. Learning the Latent Rules of a Game from Data: A Chess Story 4 upvotes, #25 of 2024-10-04

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.