LongCat

LongCat on Hugging Face Daily Papers: 32 papers, 4 in the top 3 of their day, 1 paper of the day.

  1. Better Supervision Is Nearby: Neighborhood On-Policy Self-Distillation 16 upvotes, #46 of 2026-10-02
  2. LongCat-DeepResearch Technical Report 66 upvotes, #18 of 2026-09-30
  3. DiagEvo: Diagnosis-Guided Self-Evolution via Hierarchical Error Memory 22 upvotes, #11 of 2026-09-02
  4. UniSpace: Unified Visual Representation and Scalable Multimodal Modeling 15 upvotes, #7 of 2026-08-24
  5. Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development 55 upvotes, #4 of 2026-08-17
  6. SAF-OPD: Stable Advantage Fusion for On-Policy Distillation 34 upvotes, #9 of 2026-08-03
  7. CAST: Game Solvers as Turn-Level Teachers for LLM Agents 41 upvotes, #7 of 2026-07-30
  8. GUI-CIDER: Mid-training GUI Agents via Causal Internalization and Density-aware Exemplar Reselection 23 upvotes, #17 of 2026-05-28
  9. VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions 19 upvotes, #17 of 2026-05-27
  10. Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments 16 upvotes, #18 of 2026-05-27
  11. WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation 101 upvotes, #2 of 2026-05-26
  12. HeavySkill: Heavy Thinking as the Inner Skill in Agentic Harness 22 upvotes, #6 of 2026-05-06
  13. AJ-Bench: Benchmarking Agent-as-a-Judge for Environment-Aware Evaluation 15 upvotes, #11 of 2026-04-22
  14. LARY: A Latent Action Representation Yielding Benchmark for Generalizable Vision-to-Action Alignment 20 upvotes, #10 of 2026-04-15
  15. General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks 9 upvotes, #24 of 2026-04-14
  16. Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation 75 upvotes, #4 of 2026-04-14
  17. LongCat-Next: Lexicalizing Modalities as Discrete Tokens 137 upvotes, #3 of 2026-04-01
  18. LongCat-Flash-Prover: Advancing Native Formal Reasoning via Agentic Tool-Integrated Reinforcement Learning 75 upvotes, #5 of 2026-03-24
  19. V_{0.5}: Generalist Value Model as a Prior for Sparse RL Rollouts 10 upvotes, #15 of 2026-03-12
  20. ScaleEnv: Scaling Environment Synthesis from Scratch for Generalist Interactive Tool-Use Agent Training 13 upvotes, #23 of 2026-02-11
  21. CoBA-RL: Capability-Oriented Budget Allocation for Reinforcement Learning in LLMs 33 upvotes, #11 of 2026-02-04
  22. DIFFA-2: A Practical Diffusion Large Language Model for General Audio Understanding 10 upvotes, #18 of 2026-02-02
  23. MemOCR: Layout-Aware Visual Memory for Efficient Long-Horizon Reasoning 20 upvotes, #10 of 2026-02-02
  24. Scaling Embeddings Outperforms Scaling Experts in Language Models 97 upvotes, #3 of 2026-01-30
  25. LongCat-Flash-Thinking-2601 Technical Report 171 upvotes, #1 of 2026-01-26
  26. Unlocking Implicit Experience: Synthesizing Tool-Use Trajectories from Text 38 upvotes, #4 of 2026-01-19
  27. LongCat-Image Technical Report 17 upvotes, #13 of 2025-12-09
  28. LongCat-Flash-Omni Technical Report 21 upvotes, #13 of 2025-11-04
  29. AMO-Bench: Large Language Models Still Struggle in High School Math Competitions 33 upvotes, #8 of 2025-10-31
  30. LongCat-Video Technical Report 20 upvotes, #14 of 2025-10-28
  31. R-Horizon: How Far Can Your Large Reasoning Model Really Go in Breadth and Depth? 25 upvotes, #11 of 2025-10-13
  32. VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications 17 upvotes, #17 of 2025-10-01

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.