Daily Papers of 2024-10-18

  1. Movie Gen: A Cast of Media Foundation Models 81 upvotes, #1 of 2024-10-18
  2. MixEval-X: Any-to-Any Evaluations from Real-World Data Mixtures 72 upvotes, #2 of 2024-10-18
  3. Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens 34 upvotes, #3 of 2024-10-18
  4. Roadmap towards Superhuman Speech Understanding using Large Language Models 33 upvotes, #4 of 2024-10-18
  5. JudgeBench: A Benchmark for Evaluating LLM-based Judges 30 upvotes, #5 of 2024-10-18
  6. MobA: A Two-Level Agent System for Efficient Mobile Task Automation 30 upvotes, #5 of 2024-10-18
  7. Harnessing Webpage UIs for Text-Rich Visual Understanding 28 upvotes, #7 of 2024-10-18
  8. Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation 27 upvotes, #8 of 2024-10-18
  9. WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines 26 upvotes, #9 of 2024-10-18
  10. DreamVideo-2: Zero-Shot Subject-Driven Video Customization with Precise Motion Control 22 upvotes, #10 of 2024-10-18
  11. MoH: Multi-Head Attention as Mixture-of-Head Attention 20 upvotes, #11 of 2024-10-18
  12. MMed-RAG: Versatile Multimodal RAG System for Medical Vision Language Models 20 upvotes, #11 of 2024-10-18
  13. BenTo: Benchmark Task Reduction with In-Context Transferability 20 upvotes, #11 of 2024-10-18
  14. PopAlign: Diversifying Contrasting Patterns for a More Comprehensive Alignment 18 upvotes, #14 of 2024-10-18
  15. A Comparative Study on Reasoning Patterns of OpenAI's o1 Model 15 upvotes, #15 of 2024-10-18
  16. A Unified View of Delta Parameter Editing in Post-Trained Large-Scale Models 14 upvotes, #16 of 2024-10-18
  17. FlatQuant: Flatness Matters for LLM Quantization 12 upvotes, #17 of 2024-10-18
  18. Do LLMs Have Political Correctness? Analyzing Ethical Biases and Jailbreak Vulnerabilities in AI Systems 12 upvotes, #17 of 2024-10-18
  19. VidPanos: Generative Panoramic Videos from Casual Panning Videos 11 upvotes, #19 of 2024-10-18
  20. MedMobile: A mobile-sized language model with expert-level clinical capabilities 8 upvotes, #20 of 2024-10-18
  21. Failing Forward: Improving Generative Error Correction for ASR with Synthetic Data and Retrieval Augmentation 8 upvotes, #20 of 2024-10-18
  22. Remember, Retrieve and Generate: Understanding Infinite Visual Concepts as Your Personalized Assistant 8 upvotes, #20 of 2024-10-18
  23. MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization 7 upvotes, #23 of 2024-10-18
  24. Retrospective Learning from Interactions 7 upvotes, #23 of 2024-10-18
  25. Can MLLMs Understand the Deep Implication Behind Chinese Images? 7 upvotes, #23 of 2024-10-18
  26. γ-MoD: Exploring Mixture-of-Depth Adaptation for Multimodal Large Language Models 7 upvotes, #23 of 2024-10-18
  27. LoLDU: Low-Rank Adaptation via Lower-Diag-Upper Decomposition for Parameter-Efficient Fine-Tuning 6 upvotes, #27 of 2024-10-18
  28. Open Materials 2024 (OMat24) Inorganic Materials Dataset and Models 5 upvotes, #28 of 2024-10-18
  29. Long-LRM: Long-sequence Large Reconstruction Model for Wide-coverage Gaussian Splats 5 upvotes, #28 of 2024-10-18
  30. Toward Guidance-Free AR Visual Generation via Condition Contrastive Alignment 4 upvotes, #30 of 2024-10-18
  31. Minimum Tuning to Unlock Long Output from LLMs with High Quality Data as the Key 3 upvotes, #31 of 2024-10-18
  32. TransAgent: Transfer Vision-Language Foundation Models with Heterogeneous Agent Collaboration 3 upvotes, #31 of 2024-10-18
  33. AERO: Softmax-Only LLMs for Efficient Private Inference 3 upvotes, #31 of 2024-10-18
  34. SBI-RAG: Enhancing Math Word Problem Solving for Students through Schema-Based Instruction and Retrieval-Augmented Generation 3 upvotes, #31 of 2024-10-18

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.