Rahul Gupta
Rahul Gupta on Hugging Face Daily Papers: 11 papers, 1 in the top 3 of their day, 70 upvotes.
- SkillHarm: Lifecycle-Aware Skill-Based Attacks via Automated Construction 10 upvotes, #21 of 2026-06-10
- Emergent Strategic Reasoning Risks in AI: A Taxonomy-Driven Evaluation Framework 5 upvotes, #16 of 2026-04-27
- When Actions Go Off-Task: Detecting and Correcting Misaligned Actions in Computer-Use Agents 2 upvotes, #39 of 2026-02-12
- D-REX: A Benchmark for Detecting Deceptive Reasoning in Large Language Models 3 upvotes, #26 of 2025-09-23
- Evaluating the Critical Risks of Amazon's Nova Premier under the Frontier Model Safety Framework 4 upvotes, #16 of 2025-07-10
- Towards Safety Reasoning in LLMs: AI-agentic Deliberation for Policy-embedded CoT Data Creation 17 upvotes, #21 of 2025-05-30
- Discovering Knowledge Deficiencies of Language Models on Massive Knowledge Base 6 upvotes, #26 of 2025-04-02
- Embodied Red Teaming for Auditing Robotic Foundation Models 1 upvotes, #25 of 2025-02-11
- Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language Models 3 upvotes, #22 of 2024-10-11
- FLIRT: Feedback Loop In-context Red Teaming 14 upvotes, #3 of 2023-08-09
- Controlling the Extraction of Memorized Data from Large Language Models via Prompt-Tuning 2 upvotes, #9 of 2023-05-22
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.