Daily Papers of 2025-01-29

  1. SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training 100 upvotes, #1 of 2025-01-29
  2. Optimizing Large Language Model Training Using FP4 Quantization 32 upvotes, #2 of 2025-01-29
  3. Over-Tokenized Transformer: Vocabulary is Generally Worth Scaling 22 upvotes, #3 of 2025-01-29
  4. DiffSplat: Repurposing Image Diffusion Models for Scalable Gaussian Splat Generation 21 upvotes, #4 of 2025-01-29
  5. Open Problems in Mechanistic Interpretability 16 upvotes, #5 of 2025-01-29
  6. Low-Rank Adapters Meet Neural Architecture Search for LLM Compression 7 upvotes, #6 of 2025-01-29
  7. IndicMMLU-Pro: Benchmarking Indic Large Language Models on Multi-Task Language Understanding 6 upvotes, #7 of 2025-01-29
  8. TAID: Temporally Adaptive Interpolated Distillation for Efficient Knowledge Transfer in Language Models 4 upvotes, #8 of 2025-01-29
  9. Histoires Morales: A French Dataset for Assessing Moral Alignment 3 upvotes, #9 of 2025-01-29
  10. DeepFlow: Serverless Large Language Model Serving at Scale 2 upvotes, #10 of 2025-01-29

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.