Daily Papers of 2025-02-24

  1. LLM-Microscope: Uncovering the Hidden Role of Punctuation in Context Memory of Transformers 155 upvotes, #1 of 2025-02-24
  2. SurveyX: Academic Survey Automation via Large Language Models 90 upvotes, #2 of 2025-02-24
  3. Mol-LLaMA: Towards General Understanding of Molecules in Large Molecular Language Model 42 upvotes, #3 of 2025-02-24
  4. MaskGWM: A Generalizable Driving World Model with Video Mask Reconstruction 36 upvotes, #4 of 2025-02-24
  5. PhotoDoodle: Learning Artistic Image Editing from Few-Shot Pairwise Data 36 upvotes, #4 of 2025-02-24
  6. VLM^2-Bench: A Closer Look at How Well VLMs Implicitly Link Explicit Matching Visual Cues 29 upvotes, #6 of 2025-02-24
  7. SIFT: Grounding LLM Reasoning in Contexts via Stickers 29 upvotes, #6 of 2025-02-24
  8. LightThinker: Thinking Step-by-Step Compression 25 upvotes, #8 of 2025-02-24
  9. Is Safety Standard Same for Everyone? User-Specific Safety Evaluation of Large Language Models 14 upvotes, #9 of 2025-02-24
  10. StructFlowBench: A Structured Flow Benchmark for Multi-turn Instruction Following 13 upvotes, #10 of 2025-02-24
  11. MoBA: Mixture of Block Attention for Long-Context LLMs 12 upvotes, #11 of 2025-02-24
  12. Towards Fully-Automated Materials Discovery via Large-Scale Synthesis Dataset and Expert-Level LLM-as-a-Judge 10 upvotes, #12 of 2025-02-24
  13. MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models 9 upvotes, #13 of 2025-02-24
  14. Think Inside the JSON: Reinforcement Strategy for Strict LLM Schema Adherence 9 upvotes, #13 of 2025-02-24
  15. Evaluating Multimodal Generative AI with Korean Educational Standards 9 upvotes, #13 of 2025-02-24
  16. FantasyID: Face Knowledge Enhanced ID-Preserving Video Generation 8 upvotes, #16 of 2025-02-24
  17. The Relationship Between Reasoning and Performance in Large Language Models -- o3 (mini) Thinks Harder, Not Longer 7 upvotes, #17 of 2025-02-24
  18. ReQFlow: Rectified Quaternion Flow for Efficient and High-Quality Protein Backbone Generation 6 upvotes, #18 of 2025-02-24
  19. KITAB-Bench: A Comprehensive Multi-Domain Benchmark for Arabic OCR and Document Understanding 6 upvotes, #18 of 2025-02-24
  20. InterFeedback: Unveiling Interactive Intelligence of Large Multimodal Models via Human Feedback 6 upvotes, #18 of 2025-02-24
  21. EgoSpeak: Learning When to Speak for Egocentric Conversational Agents in the Wild 5 upvotes, #21 of 2025-02-24
  22. One-step Diffusion Models with f-Divergence Distribution Matching 5 upvotes, #21 of 2025-02-24
  23. Tree-of-Debate: Multi-Persona Debate Trees Elicit Critical Thinking for Scientific Comparative Analysis 4 upvotes, #23 of 2025-02-24
  24. Superintelligent Agents Pose Catastrophic Risks: Can Scientist AI Offer a Safer Path? 4 upvotes, #23 of 2025-02-24
  25. mStyleDistance: Multilingual Style Embeddings and their Evaluation 3 upvotes, #25 of 2025-02-24
  26. WHAC: World-grounded Humans and Cameras 2 upvotes, #26 of 2025-02-24
  27. Benchmarking LLMs for Political Science: A United Nations Perspective 2 upvotes, #26 of 2025-02-24
  28. CrossOver: 3D Scene Cross-Modal Alignment 2 upvotes, #26 of 2025-02-24
  29. Rare Disease Differential Diagnosis with Large Language Models at Scale: From Abdominal Actinomycosis to Wilson's Disease 2 upvotes, #26 of 2025-02-24
  30. JL1-CD: A New Benchmark for Remote Sensing Change Detection and a Robust Multi-Teacher Knowledge Distillation Framework 1 upvotes, #30 of 2025-02-24
  31. PLDR-LLMs Learn A Generalizable Tensor Operator That Can Replace Its Own Deep Neural Net At Inference 1 upvotes, #30 of 2025-02-24
  32. Learning to Discover Regulatory Elements for Gene Expression Prediction 1 upvotes, #30 of 2025-02-24
  33. UPCORE: Utility-Preserving Coreset Selection for Balanced Unlearning 1 upvotes, #30 of 2025-02-24
  34. Beyond No: Quantifying AI Over-Refusal and Emotional Attachment Boundaries 0 upvotes, #34 of 2025-02-24

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.