Ningyu Zhang

Ningyu Zhang on Hugging Face Daily Papers: 68 papers, 9 in the top 3 of their day, 1,972 upvotes.

  1. MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use 33 upvotes, #5 of 2026-08-21
  2. MobileMem: Learning from a Year of Mobile Experiences 25 upvotes, #9 of 2026-08-17
  3. Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence 84 upvotes, #4 of 2026-08-13
  4. OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents 35 upvotes, #7 of 2026-08-06
  5. LightMem-Ego: Your AI Memory for Everyday Life 47 upvotes, #5 of 2026-07-14
  6. TokenPilot: Cache-Efficient Context Management for LLM Agents 16 upvotes, #13 of 2026-06-16
  7. LabVLA: Grounding Vision-Language-Action Models in Scientific Laboratories 53 upvotes, #9 of 2026-06-12
  8. LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis 21 upvotes, #20 of 2026-06-01
  9. How LoRA Remembers? A Parametric Memory Law for LLM Finetuning 41 upvotes, #8 of 2026-05-29
  10. MemTrace: Tracing and Attributing Errors in Large Language Model Memory Systems 39 upvotes, #9 of 2026-05-28
  11. Rethinking Memory as Continuously Evolving Connectivity 34 upvotes, #12 of 2026-05-28
  12. SciAtlas: A Large-Scale Knowledge Graph for Automated Scientific Research 58 upvotes, #4 of 2026-05-25
  13. OceanPile: A Large-Scale Multimodal Ocean Corpus for Foundation Models 15 upvotes, #7 of 2026-05-05
  14. Rewarding the Scientific Process: Process-Level Reward Modeling for Agentic Data Analysis 14 upvotes, #9 of 2026-04-28
  15. Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Language 21 upvotes, #9 of 2026-04-22
  16. LightThinker++: From Reasoning Compression to Memory Management 33 upvotes, #12 of 2026-04-07
  17. SkillNet: Create, Evaluate, and Connect AI Skills 79 upvotes, #2 of 2026-03-06
  18. How Controllable Are Large Language Models? A Unified Evaluation across Behavioral Granularities 22 upvotes, #8 of 2026-03-04
  19. InnoEval: On Research Idea Evaluation as a Knowledge-Grounded, Multi-Perspective Reasoning Problem 16 upvotes, #8 of 2026-02-17
  20. From Data to Behavior: Predicting Unintended Model Behaviors Before Training 15 upvotes, #25 of 2026-02-05
  21. Why Steering Works: Toward a Unified View of Language Model Parameter Dynamics 13 upvotes, #32 of 2026-02-03
  22. Aligning Agentic World Models via Knowledgeable Experience Learning 15 upvotes, #15 of 2026-01-21
  23. Illusions of Confidence? Diagnosing LLM Truthfulness via Neighborhood Consistency 18 upvotes, #12 of 2026-01-12
  24. Can We Predict Before Executing Machine Learning Agents? 25 upvotes, #9 of 2026-01-12
  25. InnoGym: Benchmarking the Innovation Potential of AI Agents 33 upvotes, #10 of 2025-12-03
  26. LightMem: Lightweight and Efficient Memory-Augmented Generation 105 upvotes, #2 of 2025-10-22
  27. Executable Knowledge Graphs for Replicating AI Research 12 upvotes, #15 of 2025-10-21
  28. OceanGym: A Benchmark Environment for Underwater Embodied Agents 33 upvotes, #8 of 2025-10-01
  29. Scaling Generalist Data-Analytic Agents 16 upvotes, #25 of 2025-09-30
  30. Towards Personalized Deep Research: Benchmarks and Evaluations 27 upvotes, #13 of 2025-09-30
  31. Memp: Exploring Agent Procedural Memory 27 upvotes, #3 of 2025-08-11
  32. MemOS: A Memory OS for AI System 113 upvotes, #1 of 2025-07-08
  33. ReCode: Updating Code API Knowledge with Reinforcement Learning 7 upvotes, #12 of 2025-06-26
  34. Why Do Open-Source LLMs Struggle with Data Analysis? A Systematic Empirical Study 8 upvotes, #18 of 2025-06-25
  35. KnowRL: Exploring Knowledgeable Reinforcement Learning for Factuality 6 upvotes, #19 of 2025-06-25
  36. AutoMind: Adaptive Knowledgeable Agent for Automated Data Science 18 upvotes, #17 of 2025-06-13
  37. ChineseHarm-Bench: A Chinese Harmful Content Detection Benchmark 12 upvotes, #18 of 2025-06-13
  38. Spatial Knowledge Graph-Guided Multimodal Synthesis 3 upvotes, #57 of 2025-05-28
  39. Beyond Prompt Engineering: Robust Behavior Control in LLMs via Steering Target Atoms 14 upvotes, #29 of 2025-05-28
  40. Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training 9 upvotes, #23 of 2025-05-21
  41. Knowledge Augmented Complex Problem Solving with Large Language Models: A Survey 8 upvotes, #12 of 2025-05-08
  42. A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment 13 upvotes, #11 of 2025-04-24
  43. EasyEdit2: An Easy-to-use Steering Framework for Editing Large Language Models 21 upvotes, #12 of 2025-04-22
  44. SynWorld: Virtual Scenario Synthesis for Agentic Action Knowledge Refinement 17 upvotes, #5 of 2025-04-07
  45. Agentic Knowledgeable Self-awareness 27 upvotes, #3 of 2025-04-07
  46. ZJUKLAB at SemEval-2025 Task 4: Unlearning via Model Merging 6 upvotes, #18 of 2025-03-28
  47. ADS-Edit: A Multimodal Knowledge Editing Dataset for Autonomous Driving Systems 6 upvotes, #17 of 2025-03-27
  48. LookAhead Tuning: Safer Language Models via Partial Answer Previews 5 upvotes, #18 of 2025-03-26
  49. CaKE: Circuit-aware Editing Enables Generalizable Knowledge Learners 15 upvotes, #18 of 2025-03-21
  50. BiasEdit: Debiasing Stereotyped Language Models via Model Editing 6 upvotes, #24 of 2025-03-12
  51. LightThinker: Thinking Step-by-Step Compression 25 upvotes, #8 of 2025-02-24
  52. ReLearn: Unlearning via Learning for Large Language Models 28 upvotes, #4 of 2025-02-18
  53. How Do LLMs Acquire New Knowledge? A Knowledge Circuits Perspective on Continual Pre-Training 21 upvotes, #6 of 2025-02-18
  54. OmniThink: Expanding Knowledge Boundaries in Machine Writing through Thinking 46 upvotes, #2 of 2025-01-17
  55. A Multi-Modal AI Copilot for Single-Cell Analysis with Instruction Following 24 upvotes, #7 of 2025-01-15
  56. OneKE: A Dockerized Schema-Guided LLM Agent-based Knowledge Extraction System 16 upvotes, #9 of 2024-12-31
  57. Exploring Model Kinship for Merging Large Language Models 19 upvotes, #6 of 2024-10-17
  58. MLLM can see? Dynamic Correction Decoding for Hallucination Mitigation 24 upvotes, #5 of 2024-10-16
  59. Benchmarking Agentic Workflow Generation 22 upvotes, #7 of 2024-10-11
  60. OneGen: Efficient One-Pass Unified Generation and Retrieval for LLMs 26 upvotes, #3 of 2024-09-10
  61. Benchmarking Chinese Knowledge Rectification in Large Language Models 13 upvotes, #7 of 2024-09-10
  62. Knowledge Mechanisms in Large Language Models: A Survey and Perspective 31 upvotes, #4 of 2024-07-23
  63. To Forget or Not? Towards Practical Knowledge Unlearning for Large Language Models 13 upvotes, #8 of 2024-07-03
  64. Symbolic Learning Enables Self-Evolving Agents 9 upvotes, #9 of 2024-06-27
  65. ChatCell: Facilitating Single-Cell Analysis with Natural Language 13 upvotes, #7 of 2024-02-14
  66. Weaver: Foundation Models for Creative Writing 46 upvotes, #1 of 2024-01-31
  67. A Comprehensive Study of Knowledge Editing for Large Language Models 19 upvotes, #6 of 2024-01-03
  68. Agents: An Open-source Framework for Autonomous Language Agents 43 upvotes, #2 of 2023-09-15

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.