Daily Papers of 2024-06-17
- XLand-100B: A Large-Scale Multi-Task Dataset for In-Context Reinforcement Learning 81 upvotes, #1 of 2024-06-17
- Make It Count: Text-to-Image Generation with an Accurate Number of Objects 69 upvotes, #2 of 2024-06-17
- ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code Generation 52 upvotes, #3 of 2024-06-17
- Needle In A Multimodal Haystack 51 upvotes, #4 of 2024-06-17
- BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack 46 upvotes, #5 of 2024-06-17
- OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text 28 upvotes, #6 of 2024-06-17
- SEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian Languages 24 upvotes, #7 of 2024-06-17
- GUI Odyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile Devices 21 upvotes, #8 of 2024-06-17
- Glyph-ByT5-v2: A Strong Aesthetic Baseline for Accurate Multilingual Visual Text Rendering 20 upvotes, #9 of 2024-06-17
- GEB-1.3B: Open Lightweight Large Language Model 17 upvotes, #10 of 2024-06-17
- Training-free Camera Control for Video Generation 11 upvotes, #11 of 2024-06-17
- Designing a Dashboard for Transparency and Control of Conversational AI 9 upvotes, #12 of 2024-06-17
- Rethinking Human Evaluation Protocol for Text-to-Video Models: Enhancing Reliability,Reproducibility, and Practicality 8 upvotes, #13 of 2024-06-17
- Be like a Goldfish, Don't Memorize! Mitigating Memorization in Generative LLMs 8 upvotes, #13 of 2024-06-17
- VideoGUI: A Benchmark for GUI Automation from Instructional Videos 8 upvotes, #13 of 2024-06-17
- RVT-2: Learning Precise Manipulation from Few Demonstrations 7 upvotes, #16 of 2024-06-17
- Vivid-ZOO: Multi-View Video Generation with Diffusion Model 7 upvotes, #16 of 2024-06-17
- AV-GS: Learning Material and Geometry Aware Priors for Novel View Acoustic Synthesis 6 upvotes, #18 of 2024-06-17
- GaussianSR: 3D Gaussian Super-Resolution with 2D Diffusion Priors 6 upvotes, #18 of 2024-06-17
- MaskLID: Code-Switching Language Identification through Iterative Masking 5 upvotes, #20 of 2024-06-17
- Decoding the Diversity: A Review of the Indic AI Research Landscape 5 upvotes, #20 of 2024-06-17
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.