Ningyu Zhang
Ningyu Zhang on Hugging Face Daily Papers: 68 papers, 9 in the top 3 of their day, 1,972 upvotes.
- MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use 33 upvotes, #5 of 2026-08-21
- MobileMem: Learning from a Year of Mobile Experiences 25 upvotes, #9 of 2026-08-17
- Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence 84 upvotes, #4 of 2026-08-13
- OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents 35 upvotes, #7 of 2026-08-06
- LightMem-Ego: Your AI Memory for Everyday Life 47 upvotes, #5 of 2026-07-14
- TokenPilot: Cache-Efficient Context Management for LLM Agents 16 upvotes, #13 of 2026-06-16
- LabVLA: Grounding Vision-Language-Action Models in Scientific Laboratories 53 upvotes, #9 of 2026-06-12
- LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis 21 upvotes, #20 of 2026-06-01
- How LoRA Remembers? A Parametric Memory Law for LLM Finetuning 41 upvotes, #8 of 2026-05-29
- MemTrace: Tracing and Attributing Errors in Large Language Model Memory Systems 39 upvotes, #9 of 2026-05-28
- Rethinking Memory as Continuously Evolving Connectivity 34 upvotes, #12 of 2026-05-28
- SciAtlas: A Large-Scale Knowledge Graph for Automated Scientific Research 58 upvotes, #4 of 2026-05-25
- OceanPile: A Large-Scale Multimodal Ocean Corpus for Foundation Models 15 upvotes, #7 of 2026-05-05
- Rewarding the Scientific Process: Process-Level Reward Modeling for Agentic Data Analysis 14 upvotes, #9 of 2026-04-28
- Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Language 21 upvotes, #9 of 2026-04-22
- LightThinker++: From Reasoning Compression to Memory Management 33 upvotes, #12 of 2026-04-07
- SkillNet: Create, Evaluate, and Connect AI Skills 79 upvotes, #2 of 2026-03-06
- How Controllable Are Large Language Models? A Unified Evaluation across Behavioral Granularities 22 upvotes, #8 of 2026-03-04
- InnoEval: On Research Idea Evaluation as a Knowledge-Grounded, Multi-Perspective Reasoning Problem 16 upvotes, #8 of 2026-02-17
- From Data to Behavior: Predicting Unintended Model Behaviors Before Training 15 upvotes, #25 of 2026-02-05
- Why Steering Works: Toward a Unified View of Language Model Parameter Dynamics 13 upvotes, #32 of 2026-02-03
- Aligning Agentic World Models via Knowledgeable Experience Learning 15 upvotes, #15 of 2026-01-21
- Illusions of Confidence? Diagnosing LLM Truthfulness via Neighborhood Consistency 18 upvotes, #12 of 2026-01-12
- Can We Predict Before Executing Machine Learning Agents? 25 upvotes, #9 of 2026-01-12
- InnoGym: Benchmarking the Innovation Potential of AI Agents 33 upvotes, #10 of 2025-12-03
- LightMem: Lightweight and Efficient Memory-Augmented Generation 105 upvotes, #2 of 2025-10-22
- Executable Knowledge Graphs for Replicating AI Research 12 upvotes, #15 of 2025-10-21
- OceanGym: A Benchmark Environment for Underwater Embodied Agents 33 upvotes, #8 of 2025-10-01
- Scaling Generalist Data-Analytic Agents 16 upvotes, #25 of 2025-09-30
- Towards Personalized Deep Research: Benchmarks and Evaluations 27 upvotes, #13 of 2025-09-30
- Memp: Exploring Agent Procedural Memory 27 upvotes, #3 of 2025-08-11
- MemOS: A Memory OS for AI System 113 upvotes, #1 of 2025-07-08
- ReCode: Updating Code API Knowledge with Reinforcement Learning 7 upvotes, #12 of 2025-06-26
- Why Do Open-Source LLMs Struggle with Data Analysis? A Systematic Empirical Study 8 upvotes, #18 of 2025-06-25
- KnowRL: Exploring Knowledgeable Reinforcement Learning for Factuality 6 upvotes, #19 of 2025-06-25
- AutoMind: Adaptive Knowledgeable Agent for Automated Data Science 18 upvotes, #17 of 2025-06-13
- ChineseHarm-Bench: A Chinese Harmful Content Detection Benchmark 12 upvotes, #18 of 2025-06-13
- Spatial Knowledge Graph-Guided Multimodal Synthesis 3 upvotes, #57 of 2025-05-28
- Beyond Prompt Engineering: Robust Behavior Control in LLMs via Steering Target Atoms 14 upvotes, #29 of 2025-05-28
- Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training 9 upvotes, #23 of 2025-05-21
- Knowledge Augmented Complex Problem Solving with Large Language Models: A Survey 8 upvotes, #12 of 2025-05-08
- A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment 13 upvotes, #11 of 2025-04-24
- EasyEdit2: An Easy-to-use Steering Framework for Editing Large Language Models 21 upvotes, #12 of 2025-04-22
- SynWorld: Virtual Scenario Synthesis for Agentic Action Knowledge Refinement 17 upvotes, #5 of 2025-04-07
- Agentic Knowledgeable Self-awareness 27 upvotes, #3 of 2025-04-07
- ZJUKLAB at SemEval-2025 Task 4: Unlearning via Model Merging 6 upvotes, #18 of 2025-03-28
- ADS-Edit: A Multimodal Knowledge Editing Dataset for Autonomous Driving Systems 6 upvotes, #17 of 2025-03-27
- LookAhead Tuning: Safer Language Models via Partial Answer Previews 5 upvotes, #18 of 2025-03-26
- CaKE: Circuit-aware Editing Enables Generalizable Knowledge Learners 15 upvotes, #18 of 2025-03-21
- BiasEdit: Debiasing Stereotyped Language Models via Model Editing 6 upvotes, #24 of 2025-03-12
- LightThinker: Thinking Step-by-Step Compression 25 upvotes, #8 of 2025-02-24
- ReLearn: Unlearning via Learning for Large Language Models 28 upvotes, #4 of 2025-02-18
- How Do LLMs Acquire New Knowledge? A Knowledge Circuits Perspective on Continual Pre-Training 21 upvotes, #6 of 2025-02-18
- OmniThink: Expanding Knowledge Boundaries in Machine Writing through Thinking 46 upvotes, #2 of 2025-01-17
- A Multi-Modal AI Copilot for Single-Cell Analysis with Instruction Following 24 upvotes, #7 of 2025-01-15
- OneKE: A Dockerized Schema-Guided LLM Agent-based Knowledge Extraction System 16 upvotes, #9 of 2024-12-31
- Exploring Model Kinship for Merging Large Language Models 19 upvotes, #6 of 2024-10-17
- MLLM can see? Dynamic Correction Decoding for Hallucination Mitigation 24 upvotes, #5 of 2024-10-16
- Benchmarking Agentic Workflow Generation 22 upvotes, #7 of 2024-10-11
- OneGen: Efficient One-Pass Unified Generation and Retrieval for LLMs 26 upvotes, #3 of 2024-09-10
- Benchmarking Chinese Knowledge Rectification in Large Language Models 13 upvotes, #7 of 2024-09-10
- Knowledge Mechanisms in Large Language Models: A Survey and Perspective 31 upvotes, #4 of 2024-07-23
- To Forget or Not? Towards Practical Knowledge Unlearning for Large Language Models 13 upvotes, #8 of 2024-07-03
- Symbolic Learning Enables Self-Evolving Agents 9 upvotes, #9 of 2024-06-27
- ChatCell: Facilitating Single-Cell Analysis with Natural Language 13 upvotes, #7 of 2024-02-14
- Weaver: Foundation Models for Creative Writing 46 upvotes, #1 of 2024-01-31
- A Comprehensive Study of Knowledge Editing for Large Language Models 19 upvotes, #6 of 2024-01-03
- Agents: An Open-source Framework for Autonomous Language Agents 43 upvotes, #2 of 2023-09-15
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.