Daily Papers of 2024-11-08

  1. OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models 99 upvotes, #1 of 2024-11-08
  2. ReCapture: Generative Video Camera Controls for User-Provided Videos using Masked Video Fine-Tuning 64 upvotes, #2 of 2024-11-08
  3. BitNet a4.8: 4-bit Activations for 1-bit LLMs 61 upvotes, #3 of 2024-11-08
  4. Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models 46 upvotes, #4 of 2024-11-08
  5. DimensionX: Create Any 3D and 4D Scenes from a Single Image with Controllable Video Diffusion 42 upvotes, #5 of 2024-11-08
  6. M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding 25 upvotes, #6 of 2024-11-08
  7. TIP-I2V: A Million-Scale Real Text and Image Prompt Dataset for Image-to-Video Generation 23 upvotes, #7 of 2024-11-08
  8. Thanos: Enhancing Conversational Agents with Skill-of-Mind-Infused Large Language Model 20 upvotes, #8 of 2024-11-08
  9. VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos 20 upvotes, #8 of 2024-11-08
  10. Needle Threading: Can LLMs Follow Threads through Near-Million-Scale Haystacks? 20 upvotes, #8 of 2024-11-08
  11. Analyzing The Language of Visual Tokens 19 upvotes, #11 of 2024-11-08
  12. DynaMem: Online Dynamic Spatio-Semantic Memory for Open World Mobile Manipulation 16 upvotes, #12 of 2024-11-08
  13. RetrieveGPT: Merging Prompts and Mathematical Models for Enhanced Code-Mixed Information Retrieval 15 upvotes, #13 of 2024-11-08
  14. SVDQunat: Absorbing Outliers by Low-Rank Components for 4-Bit Diffusion Models 15 upvotes, #13 of 2024-11-08
  15. M3SciQA: A Multi-Modal Multi-Document Scientific QA Benchmark for Evaluating Foundation Models 14 upvotes, #15 of 2024-11-08
  16. GazeGen: Gaze-Driven User Interaction for Visual Content Generation 14 upvotes, #15 of 2024-11-08
  17. SG-I2V: Self-Guided Trajectory Control in Image-to-Video Generation 13 upvotes, #17 of 2024-11-08
  18. Diff-2-in-1: Bridging Generation and Dense Perception with Diffusion Models 13 upvotes, #17 of 2024-11-08

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.