KnowMe-Bench: Benchmarking Person Understanding for Lifelong Digital Companions

wu, Zhisheng Chen, Ziyan Weng, Shuhe Wangv2, Chenglong Li, zhangshuo, Sen Hu, Silin Wu, Qizhen Lan, Huacan Wang, RongHao Chen

KnowMe-Bench: Benchmarking Person Understanding for Lifelong Digital Companions: 35 upvotes on Hugging Face Daily Papers, #4 of 24 papers on 2026-01-14. Day-by-day upvote history. It lost 24 votes when the Hub removed votes in bulk.

Existing long-horizon memory benchmarks mostly use multi-turn dialogues or synthetic user histories, which makes retrieval performance an imperfect proxy for person understanding. We present \BenchName, a publicly releasable benchmark built from long-form autobiographical narratives, where actions, context, and inner thoughts provide dense evidence for inferring stable motivations and decision principles. \BenchName~reconstructs each narrative into a flashback-aware, time-anchored stream and evaluates models with evidence-linked questions spanning factual recall, subjective state attribution, and principle-level reasoning. Across diverse narrative sources, retrieval-augmented systems mainly improve factual accuracy, while errors persist on temporally grounded explanations and higher-level inferences, highlighting the need for memory mechanisms beyond retrieval. Our data is in KnowMeBench{https://github.com/QuantaAlpha/KnowMeBench}.

Paper page on Hugging Face · arXiv

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.