Ameya Prabhu
Ameya Prabhu on Hugging Face Daily Papers: 20 papers, 3 in the top 3 of their day, 472 upvotes.
- Stealing Reasoning Traces from Proprietary LLM APIs 107 upvotes, #5 of 2026-08-11
- DataComp-VLM: Improved Open Datasets for Vision-Language Models 51 upvotes, #3 of 2026-07-06
- QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents 12 upvotes, #23 of 2026-07-01
- FutureSim: Replaying World Events to Evaluate Adaptive Agents 7 upvotes, #35 of 2026-05-15
- Personalizing Text-to-Image Generation to Individual Taste 8 upvotes, #28 of 2026-04-10
- Scaling Open-Ended Reasoning to Predict the Future 15 upvotes, #11 of 2026-01-01
- Adaptive Attacks on Trusted Monitors Subvert AI Control Protocols 5 upvotes, #29 of 2025-10-13
- Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLM 10 upvotes, #15 of 2025-09-23
- Answer Matching Outperforms Multiple Choice for Language Model Evaluation 8 upvotes, #11 of 2025-07-03
- A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility 18 upvotes, #7 of 2025-04-10
- Are We Done with Object-Centric Learning? 5 upvotes, #15 of 2025-04-10
- Can Language Models Falsify? Evaluating Algorithmic Reasoning with Counterexample Creation 18 upvotes, #11 of 2025-02-27
- Project Alexandria: Towards Freeing Scientific Knowledge from Copyright Burdens via LLMs 19 upvotes, #10 of 2025-02-27
- Great Models Think Alike and this Undermines AI Oversight 25 upvotes, #5 of 2025-02-07
- Humanity's Last Exam 50 upvotes, #1 of 2025-01-27
- ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities 6 upvotes, #21 of 2024-12-13
- Data Contamination Report from the 2024 CONDA Shared Task 8 upvotes, #6 of 2024-08-01
- No "Zero-Shot" Without Exponential Data: Pretraining Concept Frequency Determines Multimodal Model Performance 24 upvotes, #2 of 2024-04-08
- Inverse Scaling: When Bigger Isn't Better 10 upvotes, #7 of 2023-06-19
- Online Continual Learning Without the Storage Constraint 4 upvotes, #4 of 2023-05-17
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.