Daily Papers of 2024-11-08
- OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models 99 upvotes, #1 of 2024-11-08
- ReCapture: Generative Video Camera Controls for User-Provided Videos using Masked Video Fine-Tuning 64 upvotes, #2 of 2024-11-08
- BitNet a4.8: 4-bit Activations for 1-bit LLMs 61 upvotes, #3 of 2024-11-08
- Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models 46 upvotes, #4 of 2024-11-08
- DimensionX: Create Any 3D and 4D Scenes from a Single Image with Controllable Video Diffusion 42 upvotes, #5 of 2024-11-08
- M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding 25 upvotes, #6 of 2024-11-08
- TIP-I2V: A Million-Scale Real Text and Image Prompt Dataset for Image-to-Video Generation 23 upvotes, #7 of 2024-11-08
- Thanos: Enhancing Conversational Agents with Skill-of-Mind-Infused Large Language Model 20 upvotes, #8 of 2024-11-08
- VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos 20 upvotes, #8 of 2024-11-08
- Needle Threading: Can LLMs Follow Threads through Near-Million-Scale Haystacks? 20 upvotes, #8 of 2024-11-08
- Analyzing The Language of Visual Tokens 19 upvotes, #11 of 2024-11-08
- DynaMem: Online Dynamic Spatio-Semantic Memory for Open World Mobile Manipulation 16 upvotes, #12 of 2024-11-08
- RetrieveGPT: Merging Prompts and Mathematical Models for Enhanced Code-Mixed Information Retrieval 15 upvotes, #13 of 2024-11-08
- SVDQunat: Absorbing Outliers by Low-Rank Components for 4-Bit Diffusion Models 15 upvotes, #13 of 2024-11-08
- M3SciQA: A Multi-Modal Multi-Document Scientific QA Benchmark for Evaluating Foundation Models 14 upvotes, #15 of 2024-11-08
- GazeGen: Gaze-Driven User Interaction for Visual Content Generation 14 upvotes, #15 of 2024-11-08
- SG-I2V: Self-Guided Trajectory Control in Image-to-Video Generation 13 upvotes, #17 of 2024-11-08
- Diff-2-in-1: Bridging Generation and Dense Perception with Diffusion Models 13 upvotes, #17 of 2024-11-08
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.