Daily Papers of 2024-08-26

  1. Building and better understanding vision-language models: insights and future directions 99 upvotes, #1 of 2024-08-26
  2. MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? 25 upvotes, #2 of 2024-08-26
  3. LayerPano3D: Layered 3D Panorama for Hyper-Immersive Scene Generation 23 upvotes, #3 of 2024-08-26
  4. Multi-Layer Transformers Gradient Can be Approximated in Almost Linear Time 19 upvotes, #4 of 2024-08-26
  5. Memory-Efficient LLM Training with Online Subspace Descent 10 upvotes, #5 of 2024-08-26
  6. CustomCrafter: Customized Video Generation with Preserving Motion and Concept Composition Abilities 9 upvotes, #6 of 2024-08-26
  7. T3M: Text Guided 3D Human Motion Synthesis from Speech 8 upvotes, #7 of 2024-08-26
  8. A Web-Based Solution for Federated Learning with LLM-Based Automation 7 upvotes, #8 of 2024-08-26
  9. HiRED: Attention-Guided Token Dropping for Efficient Inference of High-Resolution Vision-Language Models in Resource-Constrained Environments 6 upvotes, #9 of 2024-08-26
  10. CODE: Confident Ordinary Differential Editing 3 upvotes, #10 of 2024-08-26
  11. FLoD: Integrating Flexible Level of Detail into 3D Gaussian Splatting for Customizable Rendering 3 upvotes, #10 of 2024-08-26
  12. RoundTable: Leveraging Dynamic Schema and Contextual Autocomplete for Enhanced Query Precision in Tabular Question Answering 2 upvotes, #12 of 2024-08-26

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.