Daily Papers of 2024-10-10

  1. Aria: An Open Multimodal Native Mixture-of-Experts Model 102 upvotes, #1 of 2024-10-10
  2. GLEE: A Unified Framework and Benchmark for Language-based Economic Environments 78 upvotes, #2 of 2024-10-10
  3. Personalized Visual Instruction Tuning 66 upvotes, #3 of 2024-10-10
  4. Pixtral 12B 55 upvotes, #4 of 2024-10-10
  5. Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation 42 upvotes, #5 of 2024-10-10
  6. IterComp: Iterative Composition-Aware Feedback Learning from Model Gallery for Text-to-Image Generation 41 upvotes, #6 of 2024-10-10
  7. Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate 36 upvotes, #7 of 2024-10-10
  8. Unveiling the Backbone-Optimizer Coupling Bias in Visual Representation Learning 33 upvotes, #8 of 2024-10-10
  9. Pyramidal Flow Matching for Efficient Video Generative Modeling 32 upvotes, #9 of 2024-10-10
  10. F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching 30 upvotes, #10 of 2024-10-10
  11. Falcon Mamba: The First Competitive Attention-free 7B Language Model 26 upvotes, #11 of 2024-10-10
  12. Story-Adapter: A Training-free Iterative Framework for Long Story Visualization 19 upvotes, #12 of 2024-10-10
  13. MM-Ego: Towards Building Egocentric Multimodal LLMs 19 upvotes, #12 of 2024-10-10
  14. One Initialization to Rule them All: Fine-tuning via Explained Variance Adaptation 15 upvotes, #14 of 2024-10-10
  15. T2V-Turbo-v2: Enhancing Video Generation Model Post-Training through Data, Reward, and Conditional Guidance Design 14 upvotes, #15 of 2024-10-10
  16. Self-Boosting Large Language Models with Synthetic Preference Data 14 upvotes, #15 of 2024-10-10
  17. TweedieMix: Improving Multi-Concept Fusion for Diffusion-based Image/Video Generation 13 upvotes, #17 of 2024-10-10
  18. ViBiDSampler: Enhancing Video Interpolation Using Bidirectional Diffusion Sampler 13 upvotes, #17 of 2024-10-10
  19. AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs 12 upvotes, #19 of 2024-10-10
  20. Temporal Reasoning Transfer from Text to Video 12 upvotes, #19 of 2024-10-10
  21. CursorCore: Assist Programming through Aligning Anything 12 upvotes, #19 of 2024-10-10
  22. Response Tuning: Aligning Large Language Models without Instruction 11 upvotes, #22 of 2024-10-10
  23. Trans4D: Realistic Geometry-Aware Transition for Compositional Text-to-4D Synthesis 11 upvotes, #22 of 2024-10-10
  24. Diversity-Rewarded CFG Distillation 10 upvotes, #24 of 2024-10-10
  25. BroadWay: Boost Your Text-to-Video Generation Model in a Training-free Way 10 upvotes, #24 of 2024-10-10
  26. Collective Critics for Creative Story Generation 8 upvotes, #26 of 2024-10-10
  27. Mixed-Session Conversation with Egocentric Memory 8 upvotes, #26 of 2024-10-10
  28. TRACE: Temporal Grounding Video LLM via Causal Event Modeling 8 upvotes, #26 of 2024-10-10
  29. Multimodal Situational Safety 8 upvotes, #26 of 2024-10-10
  30. ING-VP: MLLMs cannot Play Easy Vision-based Games Yet 8 upvotes, #26 of 2024-10-10
  31. Data Selection via Optimal Control for Language Models 8 upvotes, #26 of 2024-10-10
  32. Multimodal Large Language Models for Inverse Molecular Design with Retrosynthetic Planning 7 upvotes, #32 of 2024-10-10
  33. Holistic Unlearning Benchmark: A Multi-Faceted Evaluation for Text-to-Image Diffusion Model Unlearning 7 upvotes, #32 of 2024-10-10
  34. FürElise: Capturing and Physically Synthesizing Hand Motions of Piano Performance 7 upvotes, #32 of 2024-10-10
  35. LLM Self-Correction with DeCRIM: Decompose, Critique, and Refine for Enhanced Following of Instructions with Multiple Constraints 7 upvotes, #32 of 2024-10-10
  36. Hallucinating AI Hijacking Attack: Large Language Models and Malicious Code Recommenders 7 upvotes, #32 of 2024-10-10
  37. TextToon: Real-Time Text Toonify Head Avatar from Single Video 7 upvotes, #32 of 2024-10-10
  38. Retrieval-Augmented Decision Transformer: External Memory for In-context RL 6 upvotes, #38 of 2024-10-10
  39. MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering 6 upvotes, #38 of 2024-10-10
  40. MentalArena: Self-play Training of Language Models for Diagnosis and Treatment of Mental Health Disorders 5 upvotes, #40 of 2024-10-10
  41. Seeker: Enhancing Exception Handling in Code with LLM-based Multi-Agent Approach 5 upvotes, #40 of 2024-10-10
  42. Jointly Generating Multi-view Consistent PBR Textures using Collaborative Control 5 upvotes, #40 of 2024-10-10
  43. VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks 4 upvotes, #43 of 2024-10-10
  44. Do great minds think alike? Investigating Human-AI Complementarity in Question Answering with CAIMIRA 4 upvotes, #43 of 2024-10-10
  45. MEXA: Multilingual Evaluation of English-Centric LLMs via Cross-Lingual Alignment 3 upvotes, #45 of 2024-10-10
  46. TinyEmo: Scaling down Emotional Reasoning via Metric Projection 3 upvotes, #45 of 2024-10-10
  47. VHELM: A Holistic Evaluation of Vision Language Models 2 upvotes, #47 of 2024-10-10
  48. Stuffed Mamba: State Collapse and State Capacity of RNN-Based Long-Context Modeling 2 upvotes, #47 of 2024-10-10
  49. Does Spatial Cognition Emerge in Frontier Models? 1 upvotes, #49 of 2024-10-10

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.