Daily Papers of 2024-10-17

  1. VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI 45 upvotes, #1 of 2024-10-17
  2. HumanEval-V: Evaluating Visual Understanding and Reasoning Abilities of Large Multimodal Models Through Coding Tasks 40 upvotes, #2 of 2024-10-17
  3. The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio 29 upvotes, #3 of 2024-10-17
  4. Revealing the Barriers of Language Agents in Planning 23 upvotes, #4 of 2024-10-17
  5. DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception 20 upvotes, #5 of 2024-10-17
  6. Exploring Model Kinship for Merging Large Language Models 19 upvotes, #6 of 2024-10-17
  7. Large Language Model Evaluation via Matrix Nuclear-Norm 18 upvotes, #7 of 2024-10-17
  8. Improving Long-Text Alignment for Text-to-Image Diffusion Models 13 upvotes, #8 of 2024-10-17
  9. ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs 13 upvotes, #8 of 2024-10-17
  10. DyVo: Dynamic Vocabularies for Learned Sparse Retrieval with Entities 12 upvotes, #10 of 2024-10-17
  11. Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements 12 upvotes, #10 of 2024-10-17
  12. ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification and KV Cache Compression 11 upvotes, #12 of 2024-10-17
  13. Simplifying, Stabilizing and Scaling Continuous-Time Consistency Models 10 upvotes, #13 of 2024-10-17
  14. ChroKnowledge: Unveiling Chronological Knowledge of Language Models in Multiple Domains 7 upvotes, #14 of 2024-10-17
  15. Neural Metamorphosis 6 upvotes, #15 of 2024-10-17
  16. Stabilize the Latent Space for Image Autoregressive Modeling: A Unified Perspective 6 upvotes, #15 of 2024-10-17
  17. Tracking Universal Features Through Fine-Tuning and Model Merging 5 upvotes, #17 of 2024-10-17
  18. OMCAT: Omni Context Aware Transformer 4 upvotes, #18 of 2024-10-17
  19. WorldMedQA-V: a multilingual, multimodal medical examination dataset for multimodal language models evaluation 4 upvotes, #18 of 2024-10-17
  20. FLARE: Faithful Logic-Aided Reasoning and Exploration 3 upvotes, #20 of 2024-10-17
  21. Insights from the Inverse: Reconstructing LLM Training Goals Through Inverse RL 3 upvotes, #20 of 2024-10-17
  22. Taming Overconfidence in LLMs: Reward Calibration in RLHF 2 upvotes, #22 of 2024-10-17
  23. From Commands to Prompts: LLM-based Semantic File System for AIOS 1 upvotes, #23 of 2024-10-17

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.