Daily Papers of 2024-08-12
- VITA: Towards Open-Source Interactive Omni Multimodal LLM 40 upvotes, #1 of 2024-08-12
- Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2 31 upvotes, #2 of 2024-08-12
- mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models 27 upvotes, #3 of 2024-08-12
- UniBench: Visual Reasoning Requires Rethinking Vision-Language Beyond Scaling 18 upvotes, #4 of 2024-08-12
- ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities 13 upvotes, #5 of 2024-08-12
- BRAT: Bonus oRthogonAl Token for Architecture Agnostic Textual Inversion 5 upvotes, #6 of 2024-08-12
- MulliVC: Multi-lingual Voice Conversion With Cycle Consistency 4 upvotes, #7 of 2024-08-12
- MooER: LLM-based Speech Recognition and Translation Models from Moore Threads 4 upvotes, #7 of 2024-08-12
- Generating novel experimental hypotheses from language models: A case study on cross-dative generalization 3 upvotes, #9 of 2024-08-12
- Kalman-Inspired Feature Propagation for Video Face Super-Resolution 3 upvotes, #9 of 2024-08-12
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.