Franck Dernoncourt
Franck Dernoncourt on Hugging Face Daily Papers: 82 papers, 4 in the top 3 of their day, 882 upvotes.
- Joint and Cross-Modal Video-Audio Generation and Editing: A Unified Formulation and Design Taxonomy 7 upvotes, #70 of 2026-10-02
- Controlled Decoding Attacks on Black-Box LLMs 6 upvotes, #74 of 2026-10-02
- FlexRouter: Learning Complementary Model Sets for Flexible LLM Routing 8 upvotes, #67 of 2026-10-02
- Personalized Image Generation with Reasoning and Reflection 7 upvotes, #70 of 2026-10-02
- Prompt2Skill: Unsupervised Skill Optimization From Natural Language Instructions 9 upvotes, #65 of 2026-10-02
- DISCO: Distributed Long Context Scaling with Grounding-Reasoning Disaggregation 3 upvotes, #78 of 2026-09-29
- Beyond Top-k Skill Retrieval: Diversity-Aware Skill Routing for LLM Agents 7 upvotes, #14 of 2026-09-14
- Online Learning with LLM Experts from Limited Feedback 6 upvotes, #15 of 2026-09-14
- A Glance Is All You Need: Single-Pass Fine-Grained Image Captioning with SimLoss 18 upvotes, #15 of 2026-09-03
- CRISP: Cliff-awaRe Input-adaptive Sparse Prefilling with Structural-Mass-Motivated Routing 8 upvotes, #21 of 2026-09-03
- Personalized Auto-Research: Towards a True AI Co-Scientist 5 upvotes, #26 of 2026-08-19
- Unifying Graph Neural Networks Through a Common Layer Equation 5 upvotes, #26 of 2026-08-19
- TSDS-Toolbox: A Toolbox for Measuring Time-Series Dataset Similarity 11 upvotes, #15 of 2026-08-12
- JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles 8 upvotes, #22 of 2026-08-12
- Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents 12 upvotes, #13 of 2026-08-12
- GRASP: GRanularity-Aware Search Policy for Agentic RAG 8 upvotes, #25 of 2026-07-17
- MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering 13 upvotes, #8 of 2026-07-06
- RL-Index: Reinforcement Learning for Retrieval Index Reasoning 6 upvotes, #17 of 2026-06-25
- CAVEWOMAN: How Large Language Models Behave Under Linguistic Input and Output Compression 5 upvotes, #20 of 2026-06-25
- Multimodal Music Recommendation System using LLMs 1 upvotes, #43 of 2026-06-05
- Orthrus: Memory-Efficient Parallel Token Generation via Dual-View Diffusion 12 upvotes, #16 of 2026-05-14
- FORTIS: Benchmarking Over-Privilege in Agent Skills 3 upvotes, #43 of 2026-05-12
- FASH-iCNN: Making Editorial Fashion Identity Inspectable Through Multimodal CNN Probing 2 upvotes, #15 of 2026-04-30
- A Survey on LLM-based Conversational User Simulation 7 upvotes, #9 of 2026-04-30
- Trust but Verify: Introducing DAVinCI -- A Framework for Dual Attribution and Verification in Claim Inference for Language Models 2 upvotes, #17 of 2026-04-24
- Explainable Disentangled Representation Learning for Generalizable Authorship Attribution in the Era of Generative AI 2 upvotes, #17 of 2026-04-24
- Anticipatory Planning for Multimodal AI Agents 2 upvotes, #33 of 2026-03-18
- ViT-AdaLA: Adapting Vision Transformers with Linear Attention 2 upvotes, #33 of 2026-03-18
- Test-Time Strategies for More Efficient and Accurate Agentic RAG 1 upvotes, #42 of 2026-03-18
- Can Large Language Models Keep Up? Benchmarking Online Adaptation to Continual Knowledge Streams 17 upvotes, #9 of 2026-03-12
- Agentic Planning with Reasoning for Image Styling via Offline RL 3 upvotes, #30 of 2026-03-10
- InfinityStory: Unlimited Video Generation with World Consistency and Character-Aware Shot Transitions 8 upvotes, #13 of 2026-03-05
- Blind to the Human Touch: Overlap Bias in LLM-Based Summary Evaluation 1 upvotes, #25 of 2026-02-17
- Benchmarking Knowledge-Extraction Attack and Defense on Retrieval-Augmented Generation 1 upvotes, #25 of 2026-02-17
- Agent Banana: High-Fidelity Image Editing with Agentic Thinking and Tooling 27 upvotes, #10 of 2026-02-11
- Segment Length Matters: A Study of Segment Lengths on Audio Fingerprinting Performance 1 upvotes, #39 of 2026-01-30
- PRISM: Learning Design Knowledge from Data for Stylistic Design Improvement 1 upvotes, #39 of 2026-01-30
- StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos 7 upvotes, #26 of 2025-12-02
- MLLM as a UI Judge: Benchmarking Multimodal LLMs for Predicting Human Perception of User Interfaces 4 upvotes, #28 of 2025-10-15
- Learning to Route LLMs from Bandit Feedback: One Policy, Many Trade-offs 3 upvotes, #42 of 2025-10-10
- Knowledge Homophily in Large Language Models 2 upvotes, #42 of 2025-10-01
- The Photographer Eye: Teaching Multimodal Large Language Models to See and Critique like Photographers 4 upvotes, #56 of 2025-09-30
- mSCoRe: a Multilingual and Scalable Benchmark for Skill-based Commonsense Reasoning 1 upvotes, #17 of 2025-08-21
- Lizard: An Efficient Linearization Framework for Large Language Models 16 upvotes, #8 of 2025-07-17
- A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality 22 upvotes, #10 of 2025-07-11
- MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos 3 upvotes, #28 of 2025-06-17
- Forecasting Time Series with LLMs via Patch-Based Prompting and Decomposition 3 upvotes, #28 of 2025-06-17
- LaMP-Cap: Personalized Figure Caption Generation With Multimodal Figure Profiles 2 upvotes, #36 of 2025-06-13
- Quantitative LLM Judges 5 upvotes, #30 of 2025-06-05
- Follow the Flow: Fine-grained Flowchart Attribution with Neurosymbolic Agents 4 upvotes, #33 of 2025-06-05
- A Graph Perspective to Probe Structural Patterns of Knowledge in Large Language Models 4 upvotes, #49 of 2025-05-30
- ChartLens: Fine-grained Visual Attribution in Charts 4 upvotes, #49 of 2025-05-30
- Understanding Generative AI Capabilities in Everyday Image Editing Tasks 23 upvotes, #12 of 2025-05-23
- Document Attribution: Examining Citation Relationships using Large Language Models 3 upvotes, #21 of 2025-05-13
- InfoVids: Reimagining the Viewer Experience with Alternative Visualization-Presenter Relationships 5 upvotes, #15 of 2025-05-07
- CORG: Generating Answers from Complex, Interrelated Contexts 8 upvotes, #6 of 2025-05-05
- Skill Discovery for Software Scripting Automation via Offline Simulations with LLMs 6 upvotes, #10 of 2025-05-02
- Towards Visual Text Grounding of Multimodal Large Language Model 13 upvotes, #11 of 2025-04-11
- Efficient Model Selection for Time Series Forecasting via LLMs 16 upvotes, #14 of 2025-04-04
- Mixture of Structural-and-Textual Retrieval over Text-rich Graph Knowledge Bases 6 upvotes, #11 of 2025-03-06
- Exploring Rewriting Approaches for Different Conversational Tasks 5 upvotes, #13 of 2025-03-06
- NoLiMa: Long-Context Evaluation Beyond Literal Matching 14 upvotes, #14 of 2025-02-13
- Towards Trustworthy Retrieval Augmented Generation for Large Language Models: A Survey 8 upvotes, #18 of 2025-02-13
- PlotGen: Multi-Agent LLM-based Scientific Data Visualization via Multimodal Feedback 5 upvotes, #20 of 2025-02-07
- ChartCitor: Multi-Agent Framework for Fine-Grained Chart Visual Attribution 7 upvotes, #18 of 2025-02-07
- Personalized Graph-Based Retrieval for Large Language Models 26 upvotes, #5 of 2025-01-07
- LUSIFER: Language Universal Space Integration for Enhanced Multilingual Embeddings with Large Language Models 12 upvotes, #7 of 2025-01-06
- Multi-LLM Text Summarization 5 upvotes, #10 of 2024-12-23
- GUI Agents: A Survey 22 upvotes, #5 of 2024-12-19
- VisDoM: Multi-Document QA with Visually Rich Elements Using Multimodal Retrieval-Augmented Generation 14 upvotes, #6 of 2024-12-18
- Personalized Multimodal Large Language Models: A Survey 12 upvotes, #16 of 2024-12-06
- Comprehensive and Practical Evaluation of Retrieval-Augmented Generation Systems for Medical Question Answering 6 upvotes, #14 of 2024-11-19
- SlimLM: An Efficient Small Language Model for On-Device Document Assistance 12 upvotes, #7 of 2024-11-19
- LoRA-Contextualizing Adaptation of Large Multimodal Models for Long Document Understanding 4 upvotes, #22 of 2024-11-05
- DynaSaur: Large Language Agents Beyond Predefined Actions 13 upvotes, #13 of 2024-11-05
- GRS-QA -- Graph Reasoning-Structured Question Answering Dataset 6 upvotes, #16 of 2024-11-04
- Survey of User Interface Design and Interaction Techniques in Generative AI Applications 11 upvotes, #7 of 2024-11-04
- Personalization of Large Language Models: A Survey 30 upvotes, #2 of 2024-11-04
- A Survey of Small Language Models 35 upvotes, #3 of 2024-10-29
- Taipan: Efficient and Expressive State Space Language Models with Selective Attention 14 upvotes, #10 of 2024-10-25
- CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages 87 upvotes, #1 of 2023-09-19
- PDFTriage: Question Answering over Long, Structured Documents 55 upvotes, #3 of 2023-09-19
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.