Daily Papers of 2024-06-10
- Mixture-of-Agents Enhances Large Language Model Capabilities 41 upvotes, #1 of 2024-06-10
- CRAG -- Comprehensive RAG Benchmark 30 upvotes, #2 of 2024-06-10
- WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild 23 upvotes, #3 of 2024-06-10
- GenAI Arena: An Open Evaluation Platform for Generative Models 18 upvotes, #4 of 2024-06-10
- Large Language Model Confidence Estimation via Black-Box Access 17 upvotes, #5 of 2024-06-10
- NATURAL PLAN: Benchmarking LLMs on Natural Language Planning 9 upvotes, #6 of 2024-06-10
- Proofread: Fixes All Errors with One Tap 9 upvotes, #6 of 2024-06-10
- Why Has Predicting Downstream Capabilities of Frontier AI Models with Scale Remained Elusive? 5 upvotes, #8 of 2024-06-10
- Boosting Large-scale Parallel Training Efficiency with C4: A Communication-Driven Approach 4 upvotes, #9 of 2024-06-10
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.