Daily Papers of 2025-02-20
- Qwen2.5-VL Technical Report 146 upvotes, #1 of 2025-02-20
- On the Trustworthiness of Generative Foundation Models: Guideline, Assessment, and Perspective 44 upvotes, #2 of 2025-02-20
- RAD: Training an End-to-End Driving Policy via Large-Scale 3DGS-based Reinforcement Learning 36 upvotes, #3 of 2025-02-20
- SongGen: A Single Stage Auto-regressive Transformer for Text-to-Song Generation 35 upvotes, #4 of 2025-02-20
- MMTEB: Massive Multilingual Text Embedding Benchmark 31 upvotes, #5 of 2025-02-20
- MoM: Linear Sequence Modeling with Mixture-of-Memories 31 upvotes, #5 of 2025-02-20
- Small Models Struggle to Learn from Strong Reasoners 27 upvotes, #7 of 2025-02-20
- Craw4LLM: Efficient Web Crawling for LLM Pretraining 27 upvotes, #7 of 2025-02-20
- LongPO: Long Context Self-Evolution of Large Language Models through Short-to-Long Preference Optimization 25 upvotes, #9 of 2025-02-20
- Autellix: An Efficient Serving Engine for LLM Agents as General Programs 18 upvotes, #10 of 2025-02-20
- Thinking Preference Optimization 15 upvotes, #11 of 2025-02-20
- SearchRAG: Can Search Engines Be Helpful for LLM-based Medical Question Answering? 12 upvotes, #12 of 2025-02-20
- Presumed Cultural Identity: How Names Shape LLM Responses 10 upvotes, #13 of 2025-02-20
- Train Small, Infer Large: Memory-Efficient LoRA Training for Large Language Models 9 upvotes, #14 of 2025-02-20
- Why Safeguarded Ships Run Aground? Aligned Large Language Models' Safety Mechanisms Tend to Be Anchored in The Template Region 9 upvotes, #14 of 2025-02-20
- InfiR : Crafting Effective Small Language Models and Multimodal Small Language Models in Reasoning 8 upvotes, #16 of 2025-02-20
- NExT-Mol: 3D Diffusion Meets 1D Language Modeling for 3D Molecule Generation 7 upvotes, #17 of 2025-02-20
- AIDE: AI-Driven Exploration in the Space of Code 7 upvotes, #17 of 2025-02-20
- AdaptiveStep: Automatically Dividing Reasoning Step through Model Confidence 7 upvotes, #17 of 2025-02-20
- REALTALK: A 21-Day Real-World Dataset for Long-Term Conversation 6 upvotes, #20 of 2025-02-20
- ActionPiece: Contextually Tokenizing Action Sequences for Generative Recommendation 5 upvotes, #21 of 2025-02-20
- From Tools to Teammates: Evaluating LLMs in Multi-Session Coding Interactions 5 upvotes, #21 of 2025-02-20
- TESS 2: A Large-Scale Generalist Diffusion Language Model 5 upvotes, #21 of 2025-02-20
- REFIND: Retrieval-Augmented Factuality Hallucination Detection in Large Language Models 4 upvotes, #24 of 2025-02-20
- Judging the Judges: A Collection of LLM-Generated Relevance Judgements 4 upvotes, #24 of 2025-02-20
- High-Fidelity Novel View Synthesis via Splatting-Guided Diffusion 3 upvotes, #26 of 2025-02-20
- MVL-SIB: A Massively Multilingual Vision-Language Benchmark for Cross-Modal Topical Matching 3 upvotes, #26 of 2025-02-20
- GIMMICK -- Globally Inclusive Multimodal Multitask Cultural Knowledge Benchmarking 3 upvotes, #26 of 2025-02-20
- Reducing Hallucinations in Language Model-based SPARQL Query Generation Using Post-Generation Memory Retrieval 2 upvotes, #29 of 2025-02-20
- Noise May Contain Transferable Knowledge: Understanding Semi-supervised Heterogeneous Domain Adaptation from an Empirical Perspective 2 upvotes, #29 of 2025-02-20
- Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above 2 upvotes, #29 of 2025-02-20
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.