Daily Papers of 2024-07-02

  1. We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning? 69 upvotes, #1 of 2024-07-02
  2. ROS-LLM: A ROS framework for embodied AI with task feedback and structured reasoning 55 upvotes, #2 of 2024-07-02
  3. MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation 35 upvotes, #3 of 2024-07-02
  4. LiteSearch: Efficacious Tree Search for LLM 34 upvotes, #4 of 2024-07-02
  5. ColPali: Efficient Document Retrieval with Vision Language Models 27 upvotes, #5 of 2024-07-02
  6. Wavelets Are All You Need for Autoregressive Image Generation 26 upvotes, #6 of 2024-07-02
  7. RegMix: Data Mixture as Regression for Language Model Pre-training 24 upvotes, #7 of 2024-07-02
  8. DiffIR2VR-Zero: Zero-Shot Video Restoration with Diffusion-based Image Restoration Models 22 upvotes, #8 of 2024-07-02
  9. Is It Really Long Context if All You Need Is Retrieval? Towards Genuinely Difficult Long Context NLP 21 upvotes, #9 of 2024-07-02
  10. Step-Controlled DPO: Leveraging Stepwise Error for Enhanced Mathematical Reasoning 21 upvotes, #9 of 2024-07-02
  11. InstantStyle-Plus: Style Transfer with Content-Preserving in Text-to-Image Generation 20 upvotes, #11 of 2024-07-02
  12. E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS 18 upvotes, #12 of 2024-07-02
  13. RealTalk: Real-time and Realistic Audio-driven Face Generation with 3D Facial Prior-guided Identity Alignment Network 15 upvotes, #13 of 2024-07-02
  14. MIRAI: Evaluating LLM Agents for Event Forecasting 14 upvotes, #14 of 2024-07-02
  15. OmniJARVIS: Unified Vision-Language-Action Tokenization Enables Open-World Instruction Following Agents 12 upvotes, #15 of 2024-07-02
  16. Auto Cherry-Picker: Learning from High-quality Generative Data Driven by Language 9 upvotes, #16 of 2024-07-02
  17. SVG: 3D Stereoscopic Video Generation via Denoising Frame Matrix 9 upvotes, #16 of 2024-07-02
  18. Chain-of-Knowledge: Integrating Knowledge Reasoning into Large Language Models by Learning from Knowledge Graphs 8 upvotes, #18 of 2024-07-02
  19. Towards Robust Speech Representation Learning for Thousands of Languages 8 upvotes, #18 of 2024-07-02
  20. T-MAC: CPU Renaissance via Table Lookup for Low-Bit LLM Deployment on Edge 7 upvotes, #20 of 2024-07-02
  21. Show Less, Instruct More: Enriching Prompts with Definitions and Guidelines for Zero-Shot NER 6 upvotes, #21 of 2024-07-02
  22. UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI 5 upvotes, #22 of 2024-07-02
  23. DogeRM: Equipping Reward Models with Domain Knowledge through Model Merging 5 upvotes, #22 of 2024-07-02
  24. Accurate Prediction of Ligand-Protein Interaction Affinities with Fine-Tuned Small Language Models 4 upvotes, #24 of 2024-07-02
  25. The SIFo Benchmark: Investigating the Sequential Instruction Following Ability of Large Language Models 3 upvotes, #25 of 2024-07-02
  26. Token Erasure as a Footprint of Implicit Vocabulary Items in LLMs 3 upvotes, #25 of 2024-07-02
  27. ProgressGym: Alignment with a Millennium of Moral Progress 2 upvotes, #27 of 2024-07-02

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.