LongCat
LongCat on Hugging Face Daily Papers: 32 papers, 4 in the top 3 of their day, 1 paper of the day.
- Better Supervision Is Nearby: Neighborhood On-Policy Self-Distillation 16 upvotes, #46 of 2026-10-02
- LongCat-DeepResearch Technical Report 66 upvotes, #18 of 2026-09-30
- DiagEvo: Diagnosis-Guided Self-Evolution via Hierarchical Error Memory 22 upvotes, #11 of 2026-09-02
- UniSpace: Unified Visual Representation and Scalable Multimodal Modeling 15 upvotes, #7 of 2026-08-24
- Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development 55 upvotes, #4 of 2026-08-17
- SAF-OPD: Stable Advantage Fusion for On-Policy Distillation 34 upvotes, #9 of 2026-08-03
- CAST: Game Solvers as Turn-Level Teachers for LLM Agents 41 upvotes, #7 of 2026-07-30
- GUI-CIDER: Mid-training GUI Agents via Causal Internalization and Density-aware Exemplar Reselection 23 upvotes, #17 of 2026-05-28
- VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions 19 upvotes, #17 of 2026-05-27
- Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments 16 upvotes, #18 of 2026-05-27
- WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation 101 upvotes, #2 of 2026-05-26
- HeavySkill: Heavy Thinking as the Inner Skill in Agentic Harness 22 upvotes, #6 of 2026-05-06
- AJ-Bench: Benchmarking Agent-as-a-Judge for Environment-Aware Evaluation 15 upvotes, #11 of 2026-04-22
- LARY: A Latent Action Representation Yielding Benchmark for Generalizable Vision-to-Action Alignment 20 upvotes, #10 of 2026-04-15
- General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks 9 upvotes, #24 of 2026-04-14
- Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation 75 upvotes, #4 of 2026-04-14
- LongCat-Next: Lexicalizing Modalities as Discrete Tokens 137 upvotes, #3 of 2026-04-01
- LongCat-Flash-Prover: Advancing Native Formal Reasoning via Agentic Tool-Integrated Reinforcement Learning 75 upvotes, #5 of 2026-03-24
- V_{0.5}: Generalist Value Model as a Prior for Sparse RL Rollouts 10 upvotes, #15 of 2026-03-12
- ScaleEnv: Scaling Environment Synthesis from Scratch for Generalist Interactive Tool-Use Agent Training 13 upvotes, #23 of 2026-02-11
- CoBA-RL: Capability-Oriented Budget Allocation for Reinforcement Learning in LLMs 33 upvotes, #11 of 2026-02-04
- DIFFA-2: A Practical Diffusion Large Language Model for General Audio Understanding 10 upvotes, #18 of 2026-02-02
- MemOCR: Layout-Aware Visual Memory for Efficient Long-Horizon Reasoning 20 upvotes, #10 of 2026-02-02
- Scaling Embeddings Outperforms Scaling Experts in Language Models 97 upvotes, #3 of 2026-01-30
- LongCat-Flash-Thinking-2601 Technical Report 171 upvotes, #1 of 2026-01-26
- Unlocking Implicit Experience: Synthesizing Tool-Use Trajectories from Text 38 upvotes, #4 of 2026-01-19
- LongCat-Image Technical Report 17 upvotes, #13 of 2025-12-09
- LongCat-Flash-Omni Technical Report 21 upvotes, #13 of 2025-11-04
- AMO-Bench: Large Language Models Still Struggle in High School Math Competitions 33 upvotes, #8 of 2025-10-31
- LongCat-Video Technical Report 20 upvotes, #14 of 2025-10-28
- R-Horizon: How Far Can Your Large Reasoning Model Really Go in Breadth and Depth? 25 upvotes, #11 of 2025-10-13
- VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications 17 upvotes, #17 of 2025-10-01
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.