Franck Dernoncourt

Franck Dernoncourt on Hugging Face Daily Papers: 82 papers, 4 in the top 3 of their day, 882 upvotes.

  1. Joint and Cross-Modal Video-Audio Generation and Editing: A Unified Formulation and Design Taxonomy 7 upvotes, #70 of 2026-10-02
  2. Controlled Decoding Attacks on Black-Box LLMs 6 upvotes, #74 of 2026-10-02
  3. FlexRouter: Learning Complementary Model Sets for Flexible LLM Routing 8 upvotes, #67 of 2026-10-02
  4. Personalized Image Generation with Reasoning and Reflection 7 upvotes, #70 of 2026-10-02
  5. Prompt2Skill: Unsupervised Skill Optimization From Natural Language Instructions 9 upvotes, #65 of 2026-10-02
  6. DISCO: Distributed Long Context Scaling with Grounding-Reasoning Disaggregation 3 upvotes, #78 of 2026-09-29
  7. Beyond Top-k Skill Retrieval: Diversity-Aware Skill Routing for LLM Agents 7 upvotes, #14 of 2026-09-14
  8. Online Learning with LLM Experts from Limited Feedback 6 upvotes, #15 of 2026-09-14
  9. A Glance Is All You Need: Single-Pass Fine-Grained Image Captioning with SimLoss 18 upvotes, #15 of 2026-09-03
  10. CRISP: Cliff-awaRe Input-adaptive Sparse Prefilling with Structural-Mass-Motivated Routing 8 upvotes, #21 of 2026-09-03
  11. Personalized Auto-Research: Towards a True AI Co-Scientist 5 upvotes, #26 of 2026-08-19
  12. Unifying Graph Neural Networks Through a Common Layer Equation 5 upvotes, #26 of 2026-08-19
  13. TSDS-Toolbox: A Toolbox for Measuring Time-Series Dataset Similarity 11 upvotes, #15 of 2026-08-12
  14. JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles 8 upvotes, #22 of 2026-08-12
  15. Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents 12 upvotes, #13 of 2026-08-12
  16. GRASP: GRanularity-Aware Search Policy for Agentic RAG 8 upvotes, #25 of 2026-07-17
  17. MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering 13 upvotes, #8 of 2026-07-06
  18. RL-Index: Reinforcement Learning for Retrieval Index Reasoning 6 upvotes, #17 of 2026-06-25
  19. CAVEWOMAN: How Large Language Models Behave Under Linguistic Input and Output Compression 5 upvotes, #20 of 2026-06-25
  20. Multimodal Music Recommendation System using LLMs 1 upvotes, #43 of 2026-06-05
  21. Orthrus: Memory-Efficient Parallel Token Generation via Dual-View Diffusion 12 upvotes, #16 of 2026-05-14
  22. FORTIS: Benchmarking Over-Privilege in Agent Skills 3 upvotes, #43 of 2026-05-12
  23. FASH-iCNN: Making Editorial Fashion Identity Inspectable Through Multimodal CNN Probing 2 upvotes, #15 of 2026-04-30
  24. A Survey on LLM-based Conversational User Simulation 7 upvotes, #9 of 2026-04-30
  25. Trust but Verify: Introducing DAVinCI -- A Framework for Dual Attribution and Verification in Claim Inference for Language Models 2 upvotes, #17 of 2026-04-24
  26. Explainable Disentangled Representation Learning for Generalizable Authorship Attribution in the Era of Generative AI 2 upvotes, #17 of 2026-04-24
  27. Anticipatory Planning for Multimodal AI Agents 2 upvotes, #33 of 2026-03-18
  28. ViT-AdaLA: Adapting Vision Transformers with Linear Attention 2 upvotes, #33 of 2026-03-18
  29. Test-Time Strategies for More Efficient and Accurate Agentic RAG 1 upvotes, #42 of 2026-03-18
  30. Can Large Language Models Keep Up? Benchmarking Online Adaptation to Continual Knowledge Streams 17 upvotes, #9 of 2026-03-12
  31. Agentic Planning with Reasoning for Image Styling via Offline RL 3 upvotes, #30 of 2026-03-10
  32. InfinityStory: Unlimited Video Generation with World Consistency and Character-Aware Shot Transitions 8 upvotes, #13 of 2026-03-05
  33. Blind to the Human Touch: Overlap Bias in LLM-Based Summary Evaluation 1 upvotes, #25 of 2026-02-17
  34. Benchmarking Knowledge-Extraction Attack and Defense on Retrieval-Augmented Generation 1 upvotes, #25 of 2026-02-17
  35. Agent Banana: High-Fidelity Image Editing with Agentic Thinking and Tooling 27 upvotes, #10 of 2026-02-11
  36. Segment Length Matters: A Study of Segment Lengths on Audio Fingerprinting Performance 1 upvotes, #39 of 2026-01-30
  37. PRISM: Learning Design Knowledge from Data for Stylistic Design Improvement 1 upvotes, #39 of 2026-01-30
  38. StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos 7 upvotes, #26 of 2025-12-02
  39. MLLM as a UI Judge: Benchmarking Multimodal LLMs for Predicting Human Perception of User Interfaces 4 upvotes, #28 of 2025-10-15
  40. Learning to Route LLMs from Bandit Feedback: One Policy, Many Trade-offs 3 upvotes, #42 of 2025-10-10
  41. Knowledge Homophily in Large Language Models 2 upvotes, #42 of 2025-10-01
  42. The Photographer Eye: Teaching Multimodal Large Language Models to See and Critique like Photographers 4 upvotes, #56 of 2025-09-30
  43. mSCoRe: a Multilingual and Scalable Benchmark for Skill-based Commonsense Reasoning 1 upvotes, #17 of 2025-08-21
  44. Lizard: An Efficient Linearization Framework for Large Language Models 16 upvotes, #8 of 2025-07-17
  45. A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality 22 upvotes, #10 of 2025-07-11
  46. MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos 3 upvotes, #28 of 2025-06-17
  47. Forecasting Time Series with LLMs via Patch-Based Prompting and Decomposition 3 upvotes, #28 of 2025-06-17
  48. LaMP-Cap: Personalized Figure Caption Generation With Multimodal Figure Profiles 2 upvotes, #36 of 2025-06-13
  49. Quantitative LLM Judges 5 upvotes, #30 of 2025-06-05
  50. Follow the Flow: Fine-grained Flowchart Attribution with Neurosymbolic Agents 4 upvotes, #33 of 2025-06-05
  51. A Graph Perspective to Probe Structural Patterns of Knowledge in Large Language Models 4 upvotes, #49 of 2025-05-30
  52. ChartLens: Fine-grained Visual Attribution in Charts 4 upvotes, #49 of 2025-05-30
  53. Understanding Generative AI Capabilities in Everyday Image Editing Tasks 23 upvotes, #12 of 2025-05-23
  54. Document Attribution: Examining Citation Relationships using Large Language Models 3 upvotes, #21 of 2025-05-13
  55. InfoVids: Reimagining the Viewer Experience with Alternative Visualization-Presenter Relationships 5 upvotes, #15 of 2025-05-07
  56. CORG: Generating Answers from Complex, Interrelated Contexts 8 upvotes, #6 of 2025-05-05
  57. Skill Discovery for Software Scripting Automation via Offline Simulations with LLMs 6 upvotes, #10 of 2025-05-02
  58. Towards Visual Text Grounding of Multimodal Large Language Model 13 upvotes, #11 of 2025-04-11
  59. Efficient Model Selection for Time Series Forecasting via LLMs 16 upvotes, #14 of 2025-04-04
  60. Mixture of Structural-and-Textual Retrieval over Text-rich Graph Knowledge Bases 6 upvotes, #11 of 2025-03-06
  61. Exploring Rewriting Approaches for Different Conversational Tasks 5 upvotes, #13 of 2025-03-06
  62. NoLiMa: Long-Context Evaluation Beyond Literal Matching 14 upvotes, #14 of 2025-02-13
  63. Towards Trustworthy Retrieval Augmented Generation for Large Language Models: A Survey 8 upvotes, #18 of 2025-02-13
  64. PlotGen: Multi-Agent LLM-based Scientific Data Visualization via Multimodal Feedback 5 upvotes, #20 of 2025-02-07
  65. ChartCitor: Multi-Agent Framework for Fine-Grained Chart Visual Attribution 7 upvotes, #18 of 2025-02-07
  66. Personalized Graph-Based Retrieval for Large Language Models 26 upvotes, #5 of 2025-01-07
  67. LUSIFER: Language Universal Space Integration for Enhanced Multilingual Embeddings with Large Language Models 12 upvotes, #7 of 2025-01-06
  68. Multi-LLM Text Summarization 5 upvotes, #10 of 2024-12-23
  69. GUI Agents: A Survey 22 upvotes, #5 of 2024-12-19
  70. VisDoM: Multi-Document QA with Visually Rich Elements Using Multimodal Retrieval-Augmented Generation 14 upvotes, #6 of 2024-12-18
  71. Personalized Multimodal Large Language Models: A Survey 12 upvotes, #16 of 2024-12-06
  72. Comprehensive and Practical Evaluation of Retrieval-Augmented Generation Systems for Medical Question Answering 6 upvotes, #14 of 2024-11-19
  73. SlimLM: An Efficient Small Language Model for On-Device Document Assistance 12 upvotes, #7 of 2024-11-19
  74. LoRA-Contextualizing Adaptation of Large Multimodal Models for Long Document Understanding 4 upvotes, #22 of 2024-11-05
  75. DynaSaur: Large Language Agents Beyond Predefined Actions 13 upvotes, #13 of 2024-11-05
  76. GRS-QA -- Graph Reasoning-Structured Question Answering Dataset 6 upvotes, #16 of 2024-11-04
  77. Survey of User Interface Design and Interaction Techniques in Generative AI Applications 11 upvotes, #7 of 2024-11-04
  78. Personalization of Large Language Models: A Survey 30 upvotes, #2 of 2024-11-04
  79. A Survey of Small Language Models 35 upvotes, #3 of 2024-10-29
  80. Taipan: Efficient and Expressive State Space Language Models with Selective Attention 14 upvotes, #10 of 2024-10-25
  81. CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages 87 upvotes, #1 of 2023-09-19
  82. PDFTriage: Question Answering over Long, Structured Documents 55 upvotes, #3 of 2023-09-19

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.