Salesforce AI Research

Salesforce AI Research on Hugging Face Daily Papers: 28 papers, 1 in the top 3 of their day, 0 paper of the day.

  1. Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents 51 upvotes, #5 of 2026-09-24
  2. Flattening Every Memory Peak in Long-Context Mixture-of-Experts Training 19 upvotes, #16 of 2026-09-17
  3. EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness? 29 upvotes, #20 of 2026-09-09
  4. RISE: Recursive Improvement via Self-Extrapolating Policy Distillation 16 upvotes, #13 of 2026-09-07
  5. Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning 177 upvotes, #4 of 2026-09-04
  6. DarwinX: Evolving Agent Harnesses Through Natural Selection 110 upvotes, #2 of 2026-08-14
  7. StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents 62 upvotes, #6 of 2026-07-28
  8. Evidence-Backed Video Question Answering 10 upvotes, #19 of 2026-07-14
  9. Learning from Language Feedback via Variational Policy Distillation 10 upvotes, #22 of 2026-05-21
  10. The Illusion of Certainty: Decoupling Capability and Calibration in On-Policy Distillation 14 upvotes, #13 of 2026-04-21
  11. Beyond the Assistant Turn: User Turn Generation as a Probe of Interaction Awareness in Language Models 5 upvotes, #20 of 2026-04-13
  12. GPA: Learning GUI Process Automation from Demonstrations 16 upvotes, #17 of 2026-04-03
  13. Least-Loaded Expert Parallelism: Load Balancing An Imbalanced Mixture-of-Experts 5 upvotes, #27 of 2026-01-27
  14. Agentic Confidence Calibration 5 upvotes, #21 of 2026-01-23
  15. Agentic Uncertainty Quantification 8 upvotes, #19 of 2026-01-23
  16. Future Optical Flow Prediction Improves Robot Control & Video Generation 19 upvotes, #9 of 2026-01-19
  17. LoCoBench-Agent: An Interactive Benchmark for LLM Agents in Long-Context Software Engineering 2 upvotes, #24 of 2025-11-18
  18. MMPersuade: A Dataset and Evaluation Framework for Multimodal Persuasion 6 upvotes, #28 of 2025-10-29
  19. Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains 2 upvotes, #28 of 2025-10-21
  20. Enterprise Deep Research: Steerable Multi-Agent Deep Research for Enterprise Analytics 7 upvotes, #18 of 2025-10-21
  21. LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild 11 upvotes, #22 of 2025-10-17
  22. Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math 4 upvotes, #26 of 2025-10-16
  23. Webscale-RL: Automated Data Pipeline for Scaling RL Data to Pretraining Levels 31 upvotes, #9 of 2025-10-13
  24. UNIDOC-BENCH: A Unified Benchmark for Document-Centric Multimodal RAG 15 upvotes, #25 of 2025-10-10
  25. CoDA: Coding LM via Diffusion Adaptation 39 upvotes, #6 of 2025-10-08
  26. MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers 39 upvotes, #5 of 2025-08-21
  27. GTA1: GUI Test-time Scaling Agent 24 upvotes, #9 of 2025-07-09
  28. Demystifying Domain-adaptive Post-training for Financial LLMs 10 upvotes, #13 of 2025-01-13

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.