Daily Papers of 2024-07-08

  1. Unveiling Encoder-Free Vision-Language Models 45 upvotes, #1 of 2024-07-08
  2. FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs 33 upvotes, #2 of 2024-07-08
  3. AriGraph: Learning Knowledge Graph World Models with Episodic Memory for LLM Agents 25 upvotes, #3 of 2024-07-08
  4. Learning to (Learn at Test Time): RNNs with Expressive Hidden States 21 upvotes, #4 of 2024-07-08
  5. ChartGemma: Visual Instruction-tuning for Chart Reasoning in the Wild 19 upvotes, #5 of 2024-07-08
  6. RULE: Reliable Multimodal RAG for Factuality in Medical Vision Language Models 18 upvotes, #6 of 2024-07-08
  7. Stark: Social Long-Term Multi-Modal Conversation with Persona Commonsense Knowledge 15 upvotes, #7 of 2024-07-08
  8. DotaMath: Decomposition of Thought with Code Assistance and Self-correction for Mathematical Reasoning 13 upvotes, #8 of 2024-07-08
  9. LLM-jp: A Cross-organizational Project for the Research and Development of Fully Open Japanese LLMs 12 upvotes, #9 of 2024-07-08
  10. Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams 11 upvotes, #10 of 2024-07-08
  11. On scalable oversight with weak LLMs judging strong LLMs 11 upvotes, #10 of 2024-07-08
  12. Safe Unlearning: A Surprisingly Effective and Generalizable Solution to Defend Against Jailbreak Attacks 9 upvotes, #12 of 2024-07-08
  13. HEMM: Holistic Evaluation of Multimodal Foundation Models 8 upvotes, #13 of 2024-07-08
  14. CRiM-GS: Continuous Rigid Motion-Aware Gaussian Splatting from Motion Blur Images 7 upvotes, #14 of 2024-07-08
  15. Granular Privacy Control for Geolocation with Vision Language Models 3 upvotes, #15 of 2024-07-08

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.