Daily Papers of 2024-07-19

  1. Scaling Laws with Vocabulary: Larger Models Deserve Larger Vocabularies 46 upvotes, #1 of 2024-07-19
  2. Scaling Retrieval-Based Language Models with a Trillion-Token Datastore 27 upvotes, #2 of 2024-07-19
  3. Shape of Motion: 4D Reconstruction from a Single Video 16 upvotes, #3 of 2024-07-19
  4. Scaling Granite Code Models to 128K Context 14 upvotes, #4 of 2024-07-19
  5. Streetscapes: Large-scale Consistent Street View Generation Using Autoregressive Video Diffusion 14 upvotes, #4 of 2024-07-19
  6. Understanding Reference Policies in Direct Preference Optimization 13 upvotes, #6 of 2024-07-19
  7. Benchmarking Trustworthiness of Multimodal Large Language Models: A Comprehensive Study 10 upvotes, #7 of 2024-07-19
  8. Attention Overflow: Language Model Input Blur during Long-Context Missing Items Recommendation 7 upvotes, #8 of 2024-07-19
  9. CodeV: Empowering LLMs for Verilog Generation through Multi-Level Summarization 6 upvotes, #9 of 2024-07-19
  10. BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive Retrieval 6 upvotes, #9 of 2024-07-19
  11. CLAY: A Controllable Large-scale Generative Model for Creating High-quality 3D Assets 5 upvotes, #11 of 2024-07-19
  12. Retrieval-Enhanced Machine Learning: Synthesis and Opportunities 4 upvotes, #12 of 2024-07-19
  13. A Comparative Study on Automatic Coding of Medical Letters with Explainability 4 upvotes, #12 of 2024-07-19
  14. Benchmark Agreement Testing Done Right: A Guide for LLM Benchmark Evaluation 3 upvotes, #14 of 2024-07-19
  15. PM-LLM-Benchmark: Evaluating Large Language Models on Process Mining Tasks 2 upvotes, #15 of 2024-07-19

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.