Daily Papers of 2024-10-10
- Aria: An Open Multimodal Native Mixture-of-Experts Model 102 upvotes, #1 of 2024-10-10
- GLEE: A Unified Framework and Benchmark for Language-based Economic Environments 78 upvotes, #2 of 2024-10-10
- Personalized Visual Instruction Tuning 66 upvotes, #3 of 2024-10-10
- Pixtral 12B 55 upvotes, #4 of 2024-10-10
- Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation 42 upvotes, #5 of 2024-10-10
- IterComp: Iterative Composition-Aware Feedback Learning from Model Gallery for Text-to-Image Generation 41 upvotes, #6 of 2024-10-10
- Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate 36 upvotes, #7 of 2024-10-10
- Unveiling the Backbone-Optimizer Coupling Bias in Visual Representation Learning 33 upvotes, #8 of 2024-10-10
- Pyramidal Flow Matching for Efficient Video Generative Modeling 32 upvotes, #9 of 2024-10-10
- F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching 30 upvotes, #10 of 2024-10-10
- Falcon Mamba: The First Competitive Attention-free 7B Language Model 26 upvotes, #11 of 2024-10-10
- Story-Adapter: A Training-free Iterative Framework for Long Story Visualization 19 upvotes, #12 of 2024-10-10
- MM-Ego: Towards Building Egocentric Multimodal LLMs 19 upvotes, #12 of 2024-10-10
- One Initialization to Rule them All: Fine-tuning via Explained Variance Adaptation 15 upvotes, #14 of 2024-10-10
- T2V-Turbo-v2: Enhancing Video Generation Model Post-Training through Data, Reward, and Conditional Guidance Design 14 upvotes, #15 of 2024-10-10
- Self-Boosting Large Language Models with Synthetic Preference Data 14 upvotes, #15 of 2024-10-10
- TweedieMix: Improving Multi-Concept Fusion for Diffusion-based Image/Video Generation 13 upvotes, #17 of 2024-10-10
- ViBiDSampler: Enhancing Video Interpolation Using Bidirectional Diffusion Sampler 13 upvotes, #17 of 2024-10-10
- AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs 12 upvotes, #19 of 2024-10-10
- Temporal Reasoning Transfer from Text to Video 12 upvotes, #19 of 2024-10-10
- CursorCore: Assist Programming through Aligning Anything 12 upvotes, #19 of 2024-10-10
- Response Tuning: Aligning Large Language Models without Instruction 11 upvotes, #22 of 2024-10-10
- Trans4D: Realistic Geometry-Aware Transition for Compositional Text-to-4D Synthesis 11 upvotes, #22 of 2024-10-10
- Diversity-Rewarded CFG Distillation 10 upvotes, #24 of 2024-10-10
- BroadWay: Boost Your Text-to-Video Generation Model in a Training-free Way 10 upvotes, #24 of 2024-10-10
- Collective Critics for Creative Story Generation 8 upvotes, #26 of 2024-10-10
- Mixed-Session Conversation with Egocentric Memory 8 upvotes, #26 of 2024-10-10
- TRACE: Temporal Grounding Video LLM via Causal Event Modeling 8 upvotes, #26 of 2024-10-10
- Multimodal Situational Safety 8 upvotes, #26 of 2024-10-10
- ING-VP: MLLMs cannot Play Easy Vision-based Games Yet 8 upvotes, #26 of 2024-10-10
- Data Selection via Optimal Control for Language Models 8 upvotes, #26 of 2024-10-10
- Multimodal Large Language Models for Inverse Molecular Design with Retrosynthetic Planning 7 upvotes, #32 of 2024-10-10
- Holistic Unlearning Benchmark: A Multi-Faceted Evaluation for Text-to-Image Diffusion Model Unlearning 7 upvotes, #32 of 2024-10-10
- FürElise: Capturing and Physically Synthesizing Hand Motions of Piano Performance 7 upvotes, #32 of 2024-10-10
- LLM Self-Correction with DeCRIM: Decompose, Critique, and Refine for Enhanced Following of Instructions with Multiple Constraints 7 upvotes, #32 of 2024-10-10
- Hallucinating AI Hijacking Attack: Large Language Models and Malicious Code Recommenders 7 upvotes, #32 of 2024-10-10
- TextToon: Real-Time Text Toonify Head Avatar from Single Video 7 upvotes, #32 of 2024-10-10
- Retrieval-Augmented Decision Transformer: External Memory for In-context RL 6 upvotes, #38 of 2024-10-10
- MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering 6 upvotes, #38 of 2024-10-10
- MentalArena: Self-play Training of Language Models for Diagnosis and Treatment of Mental Health Disorders 5 upvotes, #40 of 2024-10-10
- Seeker: Enhancing Exception Handling in Code with LLM-based Multi-Agent Approach 5 upvotes, #40 of 2024-10-10
- Jointly Generating Multi-view Consistent PBR Textures using Collaborative Control 5 upvotes, #40 of 2024-10-10
- VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks 4 upvotes, #43 of 2024-10-10
- Do great minds think alike? Investigating Human-AI Complementarity in Question Answering with CAIMIRA 4 upvotes, #43 of 2024-10-10
- MEXA: Multilingual Evaluation of English-Centric LLMs via Cross-Lingual Alignment 3 upvotes, #45 of 2024-10-10
- TinyEmo: Scaling down Emotional Reasoning via Metric Projection 3 upvotes, #45 of 2024-10-10
- VHELM: A Holistic Evaluation of Vision Language Models 2 upvotes, #47 of 2024-10-10
- Stuffed Mamba: State Collapse and State Capacity of RNN-Based Long-Context Modeling 2 upvotes, #47 of 2024-10-10
- Does Spatial Cognition Emerge in Frontier Models? 1 upvotes, #49 of 2024-10-10
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.