Daily Papers of 2024-12-13

  1. InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions 89 upvotes, #1 of 2024-12-13
  2. Phi-4 Technical Report 87 upvotes, #2 of 2024-12-13
  3. Euclid: Supercharging Multimodal LLMs with Synthetic High-Fidelity Visual Descriptions 50 upvotes, #3 of 2024-12-13
  4. Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition 43 upvotes, #4 of 2024-12-13
  5. Multimodal Latent Language Modeling with Next-Token Diffusion 38 upvotes, #5 of 2024-12-13
  6. AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web Tutorials 24 upvotes, #6 of 2024-12-13
  7. EasyRef: Omni-Generalized Group Image Reference for Diffusion Models via Multimodal LLM 21 upvotes, #7 of 2024-12-13
  8. SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training 20 upvotes, #8 of 2024-12-13
  9. JuStRank: Benchmarking LLM Judges for System Ranking 18 upvotes, #9 of 2024-12-13
  10. Neural LightRig: Unlocking Accurate Object Normal and Material Estimation with Multi-Light Diffusion 17 upvotes, #10 of 2024-12-13
  11. PIG: Physics-Informed Gaussians as Adaptive Parametric Mesh Representations 16 upvotes, #11 of 2024-12-13
  12. VisionArena: 230K Real World User-VLM Conversations with Preference Labels 11 upvotes, #12 of 2024-12-13
  13. Learned Compression for Compressed Learning 11 upvotes, #12 of 2024-12-13
  14. Arbitrary-steps Image Super-resolution via Diffusion Inversion 10 upvotes, #14 of 2024-12-13
  15. OLA-VLM: Elevating Visual Perception in Multimodal LLMs with Auxiliary Embedding Distillation 10 upvotes, #14 of 2024-12-13
  16. RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios 9 upvotes, #16 of 2024-12-13
  17. Normalizing Flows are Capable Generative Models 8 upvotes, #17 of 2024-12-13
  18. Word Sense Linking: Disambiguating Outside the Sandbox 8 upvotes, #17 of 2024-12-13
  19. FreeSplatter: Pose-free Gaussian Splatting for Sparse-view 3D Reconstruction 7 upvotes, #19 of 2024-12-13
  20. LoRACLR: Contrastive Adaptation for Customization of Diffusion Models 7 upvotes, #19 of 2024-12-13
  21. ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities 6 upvotes, #21 of 2024-12-13
  22. DisPose: Disentangling Pose Guidance for Controllable Human Image Animation 6 upvotes, #21 of 2024-12-13
  23. The Impact of Copyrighted Material on Large Language Models: A Norwegian Perspective 5 upvotes, #23 of 2024-12-13
  24. Gaze-LLE: Gaze Target Estimation via Large-Scale Learned Encoders 5 upvotes, #23 of 2024-12-13
  25. SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of Experts 4 upvotes, #25 of 2024-12-13
  26. Shiksha: A Technical Domain focused Translation Dataset and Model for Indian Languages 4 upvotes, #25 of 2024-12-13

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.