Daily Papers of 2024-07-12

  1. Skywork-Math: Data Scaling Laws for Mathematical Reasoning in Large Language Models -- The Story Goes On 46 upvotes, #1 of 2024-07-12
  2. Video Diffusion Alignment via Reward Gradients 41 upvotes, #2 of 2024-07-12
  3. Multimodal Self-Instruct: Synthetic Abstract Image and Visual Reasoning Instruction Using Language Model 37 upvotes, #3 of 2024-07-12
  4. Q-GaLore: Quantized GaLore with INT4 Projection and Layer-Adaptive Low-Rank Gradients 28 upvotes, #4 of 2024-07-12
  5. MAVIS: Mathematical Visual Instruction Tuning 26 upvotes, #5 of 2024-07-12
  6. Self-Recognition in Language Models 20 upvotes, #6 of 2024-07-12
  7. MambaVision: A Hybrid Mamba-Transformer Vision Backbone 18 upvotes, #7 of 2024-07-12
  8. SEED-Story: Multimodal Long Story Generation with Large Language Model 18 upvotes, #7 of 2024-07-12
  9. Is Your Model Really A Good Math Reasoner? Evaluating Mathematical Reasoning with Checklist 18 upvotes, #7 of 2024-07-12
  10. DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception 17 upvotes, #10 of 2024-07-12
  11. Autoregressive Speech Synthesis without Vector Quantization 12 upvotes, #11 of 2024-07-12
  12. Gradient Boosting Reinforcement Learning 10 upvotes, #12 of 2024-07-12
  13. The Synergy between Data and Multi-Modal Large Language Models: A Survey from Co-Development Perspective 10 upvotes, #12 of 2024-07-12
  14. Towards Building Specialized Generalist AI with System 1 and System 2 Fusion 9 upvotes, #14 of 2024-07-12
  15. Live2Diff: Live Stream Translation via Uni-directional Attention in Video Diffusion Models 8 upvotes, #15 of 2024-07-12
  16. GTA: A Benchmark for General Tool Agents 8 upvotes, #15 of 2024-07-12
  17. Map It Anywhere (MIA): Empowering Bird's Eye View Mapping using Large-scale Public Data 8 upvotes, #15 of 2024-07-12
  18. WildGaussians: 3D Gaussian Splatting in the Wild 7 upvotes, #18 of 2024-07-12
  19. Generalizable Implicit Motion Modeling for Video Frame Interpolation 7 upvotes, #18 of 2024-07-12
  20. OmniNOCS: A unified NOCS dataset and model for 3D lifting of 2D objects 5 upvotes, #20 of 2024-07-12
  21. Scaling Up Personalized Aesthetic Assessment via Task Vector Customization 3 upvotes, #21 of 2024-07-12

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.