Daily Papers of 2026-01-20

  1. ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development 63 upvotes, #1 of 2026-01-20
  2. Multiplex Thinking: Reasoning via Token-wise Branch-and-Merge 38 upvotes, #2 of 2026-01-20
  3. NAACL: Noise-AwAre Verbal Confidence Calibration for LLMs in RAG Systems 29 upvotes, #3 of 2026-01-20
  4. Medical SAM3: A Foundation Model for Universal Prompt-Driven Medical Image Segmentation 15 upvotes, #4 of 2026-01-20
  5. PubMed-OCR: PMC Open Access OCR Annotations 11 upvotes, #5 of 2026-01-20
  6. The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models 10 upvotes, #6 of 2026-01-20
  7. CoDance: An Unbind-Rebind Paradigm for Robust Multi-Subject Animation 8 upvotes, #7 of 2026-01-20
  8. YaPO: Learnable Sparse Activation Steering Vectors for Domain Adaptation 7 upvotes, #8 of 2026-01-20
  9. SIN-Bench: Tracing Native Evidence Chains in Long-Context Multimodal Scientific Interleaved Literature 7 upvotes, #8 of 2026-01-20
  10. Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs 7 upvotes, #8 of 2026-01-20
  11. CLARE: Continual Learning for Vision-Language-Action Models via Autonomous Adapter Routing and Expansion 4 upvotes, #11 of 2026-01-20

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.