Aashiq Muhamed
Aashiq Muhamed on Hugging Face Daily Papers: 10 papers, 0 in the top 3 of their day, 65 upvotes.
- Decoy Direction Optimization: A Post-Hoc Defense Against LLM Abliteration 4 upvotes, #22 of 2026-09-16
- Pick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks 8 upvotes, #24 of 2026-09-15
- MOLE: Detecting Insider Threats in AI Agents 19 upvotes, #34 of 2026-09-09
- Beyond Understanding: Evaluating the Pragmatic Gap in LLMs' Cultural Processing of Figurative Language 2 upvotes, #41 of 2025-10-29
- RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models 1 upvotes, #47 of 2025-10-17
- Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs 5 upvotes, #48 of 2025-05-27
- CoRAG: Collaborative Retrieval-Augmented Generation 10 upvotes, #8 of 2025-04-14
- SAEs Can Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs 4 upvotes, #20 of 2025-04-14
- Decoding Dark Matter: Specialized Sparse Autoencoders for Interpreting Rare Concepts in Foundation Models 6 upvotes, #19 of 2024-11-05
- Grass: Compute Efficient Low-Memory LLM Training with Structured Sparse Gradients 5 upvotes, #16 of 2024-06-26
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.