Daily Papers of 2024-11-05

  1. AndroidLab: Training and Systematic Benchmarking of Android Autonomous Agents 46 upvotes, #1 of 2024-11-05
  2. "Give Me BF16 or Give Me Death"? Accuracy-Performance Trade-Offs in LLM Quantization 44 upvotes, #2 of 2024-11-05
  3. WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning 34 upvotes, #3 of 2024-11-05
  4. How Far is Video Generation from World Model: A Physical Law Perspective 29 upvotes, #4 of 2024-11-05
  5. Survey of Cultural Awareness in Language Models: Text and Beyond 23 upvotes, #5 of 2024-11-05
  6. MVPaint: Synchronized Multi-View Diffusion for Painting Anything 3D 23 upvotes, #5 of 2024-11-05
  7. Training-free Regional Prompting for Diffusion Transformers 23 upvotes, #5 of 2024-11-05
  8. Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencent 22 upvotes, #8 of 2024-11-05
  9. GenXD: Generating Any 3D and 4D Scenes 18 upvotes, #9 of 2024-11-05
  10. Adaptive Caching for Faster Video Generation with Diffusion Transformers 18 upvotes, #9 of 2024-11-05
  11. DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models 15 upvotes, #11 of 2024-11-05
  12. AutoVFX: Physically Realistic Video Editing from Natural Language Instructions 14 upvotes, #12 of 2024-11-05
  13. DynaSaur: Large Language Agents Beyond Predefined Actions 13 upvotes, #13 of 2024-11-05
  14. PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance 11 upvotes, #14 of 2024-11-05
  15. Sparsing Law: Towards Large Language Models with Greater Activation Sparsity 10 upvotes, #15 of 2024-11-05
  16. IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI 8 upvotes, #16 of 2024-11-05
  17. LIBMoE: A Library for comprehensive benchmarking Mixture of Experts in Large Language Models 8 upvotes, #16 of 2024-11-05
  18. SALSA: Soup-based Alignment Learning for Stronger Adaptation in RLHF 8 upvotes, #16 of 2024-11-05
  19. Decoding Dark Matter: Specialized Sparse Autoencoders for Interpreting Rare Concepts in Foundation Models 6 upvotes, #19 of 2024-11-05
  20. Constrained Diffusion Implicit Models 5 upvotes, #20 of 2024-11-05
  21. Multi-expert Prompting Improves Reliability, Safety, and Usefulness of Large Language Models 5 upvotes, #20 of 2024-11-05
  22. LoRA-Contextualizing Adaptation of Large Multimodal Models for Long Document Understanding 4 upvotes, #22 of 2024-11-05
  23. Swan and ArabicMTEB: Dialect-Aware, Arabic-Centric, Cross-Lingual, and Cross-Cultural Embedding Models and Benchmarks 3 upvotes, #23 of 2024-11-05

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.