Rahul Gupta

Rahul Gupta on Hugging Face Daily Papers: 11 papers, 1 in the top 3 of their day, 70 upvotes.

  1. SkillHarm: Lifecycle-Aware Skill-Based Attacks via Automated Construction 10 upvotes, #21 of 2026-06-10
  2. Emergent Strategic Reasoning Risks in AI: A Taxonomy-Driven Evaluation Framework 5 upvotes, #16 of 2026-04-27
  3. When Actions Go Off-Task: Detecting and Correcting Misaligned Actions in Computer-Use Agents 2 upvotes, #39 of 2026-02-12
  4. D-REX: A Benchmark for Detecting Deceptive Reasoning in Large Language Models 3 upvotes, #26 of 2025-09-23
  5. Evaluating the Critical Risks of Amazon's Nova Premier under the Frontier Model Safety Framework 4 upvotes, #16 of 2025-07-10
  6. Towards Safety Reasoning in LLMs: AI-agentic Deliberation for Policy-embedded CoT Data Creation 17 upvotes, #21 of 2025-05-30
  7. Discovering Knowledge Deficiencies of Language Models on Massive Knowledge Base 6 upvotes, #26 of 2025-04-02
  8. Embodied Red Teaming for Auditing Robotic Foundation Models 1 upvotes, #25 of 2025-02-11
  9. Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language Models 3 upvotes, #22 of 2024-10-11
  10. FLIRT: Feedback Loop In-context Red Teaming 14 upvotes, #3 of 2023-08-09
  11. Controlling the Extraction of Memorized Data from Large Language Models via Prompt-Tuning 2 upvotes, #9 of 2023-05-22

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.