Carnegie Mellon University

Carnegie Mellon University on Hugging Face Daily Papers: 59 papers, 2 in the top 3 of their day, 2 paper of the day.

  1. Predictive Credit: Measuring What Scientific Explanations Add to Experimental Forecasts 90 upvotes, #5 of 2026-10-02
  2. BiasReducer: Adaptive Bias Mitigation for Reward Models 24 upvotes, #28 of 2026-10-01
  3. CUA-SWE: When Computer-Use Agents Meet Visual Software Engineering 18 upvotes, #32 of 2026-10-01
  4. TrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations 13 upvotes, #15 of 2026-09-28
  5. Training Object Permanence in World Models 238 upvotes, #1 of 2026-09-25
  6. WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents 14 upvotes, #15 of 2026-09-24
  7. JEV-as-a-Judge: Accept When Confident, Escalate When Unsure 42 upvotes, #8 of 2026-09-23
  8. Decoy Direction Optimization: A Post-Hoc Defense Against LLM Abliteration 4 upvotes, #22 of 2026-09-16
  9. Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation 170 upvotes, #1 of 2026-09-14
  10. Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents 20 upvotes, #11 of 2026-09-10
  11. CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs 25 upvotes, #23 of 2026-09-09
  12. MOLE: Detecting Insider Threats in AI Agents 19 upvotes, #34 of 2026-09-09
  13. Same Agent, Different Answers: A Repeat-Aware Audit of Corpus-Induced Answer Churn in Retrieval-Augmented QA 4 upvotes, #28 of 2026-08-25
  14. The Problem Is the Problem: Towards Scalable Mathematical Discovery 4 upvotes, #29 of 2026-08-19
  15. Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models 3 upvotes, #27 of 2026-08-17
  16. WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks 8 upvotes, #22 of 2026-08-11
  17. O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning 11 upvotes, #9 of 2026-07-27
  18. RocketSmith: Agentic Additive Manufacturing of High-Powered Rockets 6 upvotes, #34 of 2026-06-30
  19. Causal Discovery in the Era of Agents 7 upvotes, #29 of 2026-06-23
  20. iOSWorld: A Benchmark for Personally Intelligent Phone Agents 3 upvotes, #29 of 2026-06-18
  21. MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents 6 upvotes, #20 of 2026-06-18
  22. Hardening Agent Benchmarks with Adversarial Hacker-Fixer Loops 1 upvotes, #44 of 2026-06-09
  23. K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts 55 upvotes, #6 of 2026-06-02
  24. Same Question, Different Source, Different Answer: Auditing Source-Dependence in Medical Multi-Source RAG 1 upvotes, #58 of 2026-06-02
  25. The Chain Holds, the Answer Folds: Trace-Answer Dissociation in Reasoning Models Under Adversarial Pressure 1 upvotes, #58 of 2026-06-02
  26. PANDO: Efficient Multimodal AI Agents via Online Skill Distillation 11 upvotes, #33 of 2026-05-29
  27. CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM 9 upvotes, #41 of 2026-05-29
  28. Multi-view Consistent 3D Gaussian Head Avatars 'without' Multi-view Generation 7 upvotes, #48 of 2026-05-29
  29. Verus-SpecGym: An Agentic Environment for Evaluating Specification Autoformalization 6 upvotes, #43 of 2026-05-28
  30. LatentUMM: Dual Latent Alignment for Unified Multimodal Models 8 upvotes, #21 of 2026-05-25
  31. On the limits and opportunities of AI reviewers: Reviewing the reviews of Nature-family papers with 45 expert scientists 11 upvotes, #20 of 2026-05-21
  32. Base Models Look Human To AI Detectors 2 upvotes, #49 of 2026-05-20
  33. MixSD: Mixed Contextual Self-Distillation for Knowledge Injection 7 upvotes, #30 of 2026-05-19
  34. From Generalist to Specialist Representation 2 upvotes, #42 of 2026-05-14
  35. Auto Research with Specialist Agents Develops Effective and Non-Trivial Training Recipes 15 upvotes, #13 of 2026-05-08
  36. Building a Precise Video Language with Human-AI Oversight 10 upvotes, #7 of 2026-04-27
  37. Diverse Dictionary Learning 3 upvotes, #24 of 2026-04-23
  38. SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks 14 upvotes, #12 of 2026-04-22
  39. Symbolic Guardrails for Domain-Specific Agents: Stronger Safety and Security Guarantees Without Sacrificing Utility 2 upvotes, #35 of 2026-04-21
  40. TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training 6 upvotes, #28 of 2026-04-14
  41. AgentSocialBench: Evaluating Privacy Risks in Human-Centered Agentic Social Networks 10 upvotes, #12 of 2026-04-06
  42. Revision or Re-Solving? Decomposing Second-Pass Gains in Multi-LLM Pipelines 7 upvotes, #20 of 2026-04-02
  43. The Model Says Walk: How Surface Heuristics Override Implicit Constraints in LLM Reasoning 12 upvotes, #22 of 2026-04-01
  44. When Documents Disagree: Measuring Institutional Variation in Transplant Guidance with Retrieval-Augmented Language Models 5 upvotes, #31 of 2026-04-01
  45. In-the-Wild Camouflage Attack on Vehicle Detectors through Controllable Image Editing 2 upvotes, #36 of 2026-03-24
  46. AdditiveLLM2: A Multi-modal Large Language Model for Additive Manufacturing 2 upvotes, #36 of 2026-03-24
  47. SurvHTE-Bench: A Benchmark for Heterogeneous Treatment Effect Estimation in Survival Analysis 2 upvotes, #37 of 2026-03-13
  48. EditCtrl: Disentangled Local and Global Control for Real-Time Generative Video Editing 1 upvotes, #25 of 2026-02-17
  49. GameDevBench: Evaluating Agentic Capabilities Through Game Development 15 upvotes, #16 of 2026-02-12
  50. Reliable and Responsible Foundation Models: A Comprehensive Survey 8 upvotes, #28 of 2026-02-10
  51. CoPE: Clipped RoPE as A Scalable Free Lunch for Long Context LLMs 7 upvotes, #30 of 2026-02-06
  52. AgentArk: Distilling Multi-Agent Intelligence into a Single LLM Agent 8 upvotes, #31 of 2026-02-05
  53. Value-Based Pre-Training with Downstream Feedback 1 upvotes, #40 of 2026-02-02
  54. GPCR-Filter: a deep learning framework for efficient and precise GPCR modulator discovery 2 upvotes, #20 of 2026-01-28
  55. Agentic Search in the Wild: Intents and Trajectory Dynamics from 14M+ Real Search Requests 2 upvotes, #32 of 2026-01-27
  56. CRISP: Contact-Guided Real2Sim from Monocular Video with Planar Scene Primitives 6 upvotes, #26 of 2025-12-17
  57. RefineBench: Evaluating Refinement Capability of Language Models via Checklists 12 upvotes, #15 of 2025-12-01
  58. Adapting Vehicle Detectors for Aerial Imagery to Unseen Domains with Weak Supervision 10 upvotes, #7 of 2025-07-31
  59. LLM-3D Print: Large Language Models To Monitor and Control 3D Printing 2 upvotes, #12 of 2024-08-28

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.