Daily Papers of 2024-08-12

  1. VITA: Towards Open-Source Interactive Omni Multimodal LLM 40 upvotes, #1 of 2024-08-12
  2. Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2 31 upvotes, #2 of 2024-08-12
  3. mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models 27 upvotes, #3 of 2024-08-12
  4. UniBench: Visual Reasoning Requires Rethinking Vision-Language Beyond Scaling 18 upvotes, #4 of 2024-08-12
  5. ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities 13 upvotes, #5 of 2024-08-12
  6. BRAT: Bonus oRthogonAl Token for Architecture Agnostic Textual Inversion 5 upvotes, #6 of 2024-08-12
  7. MulliVC: Multi-lingual Voice Conversion With Cycle Consistency 4 upvotes, #7 of 2024-08-12
  8. MooER: LLM-based Speech Recognition and Translation Models from Moore Threads 4 upvotes, #7 of 2024-08-12
  9. Generating novel experimental hypotheses from language models: A case study on cross-dative generalization 3 upvotes, #9 of 2024-08-12
  10. Kalman-Inspired Feature Propagation for Video Face Super-Resolution 3 upvotes, #9 of 2024-08-12

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.