AI Safety & Interpretability Lab

AI Safety & Interpretability Lab on Hugging Face Daily Papers: 3 papers, 0 in the top 3 of their day, 0 paper of the day.

  1. Selecting The Most Informative Tokens in Natural Language Autoencoders 16 upvotes, #49 of 2026-09-30
  2. Emergent Languages in Populations of Language Model Agents: From Token Efficiency to Oversight Evasion 11 upvotes, #29 of 2026-06-01
  3. Confidence and Calibration of Activation Oracles for Reliable Interpretation of Language Model Internals 12 upvotes, #24 of 2026-05-27

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.