Daily Papers of 2024-07-10

  1. Vision language models are blind 71 upvotes, #1 of 2024-07-10
  2. AgentInstruct: Toward Generative Teaching with Agentic Flows 34 upvotes, #2 of 2024-07-10
  3. Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision 24 upvotes, #3 of 2024-07-10
  4. Internet of Agents: Weaving a Web of Heterogeneous Agents for Collaborative Intelligence 23 upvotes, #4 of 2024-07-10
  5. RodinHD: High-Fidelity 3D Avatar Generation with Diffusion Models 19 upvotes, #5 of 2024-07-10
  6. Adapting LLMs to Hebrew: Unveiling DictaLM 2.0 with Enhanced Vocabulary and Instruction Capabilities 19 upvotes, #5 of 2024-07-10
  7. MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions 14 upvotes, #7 of 2024-07-10
  8. Lookback Lens: Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention Maps 10 upvotes, #8 of 2024-07-10
  9. Knowledge Composition using Task Vectors with Learned Anisotropic Scaling 9 upvotes, #9 of 2024-07-10
  10. TheoremLlama: Transforming General-Purpose LLMs into Lean4 Experts 9 upvotes, #9 of 2024-07-10
  11. BM25S: Orders of magnitude faster lexical search via eager sparse scoring 9 upvotes, #9 of 2024-07-10
  12. Graph-Based Captioning: Enhancing Visual Descriptions by Interconnecting Region Captions 9 upvotes, #9 of 2024-07-10
  13. VIMI: Grounding Video Generation through Multi-modal Instruction 8 upvotes, #13 of 2024-07-10
  14. From Loops to Oops: Fallback Behaviors of Language Models Under Uncertainty 7 upvotes, #14 of 2024-07-10
  15. How do you know that? Teaching Generative Language Models to Reference Answers to Biomedical Questions 4 upvotes, #15 of 2024-07-10
  16. LETS-C: Leveraging Language Embedding for Time Series Classification 2 upvotes, #16 of 2024-07-10

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.