Mohamed Bin Zayed University of Artificial Intelligence
Mohamed Bin Zayed University of Artificial Intelligence on Hugging Face Daily Papers: 43 papers, 1 in the top 3 of their day, 0 paper of the day.
- HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents 18 upvotes, #11 of 2026-10-05
- Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation 20 upvotes, #41 of 2026-10-02
- Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation 26 upvotes, #37 of 2026-10-02
- Program-Verified Self-Evolution for Vision-Language Models 14 upvotes, #48 of 2026-09-29
- Hard Vision, Easy Vision: What GPT-6 Astra Reveals Across Computer Vision 22 upvotes, #34 of 2026-09-29
- Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs 36 upvotes, #10 of 2026-09-23
- Geometry of Values: Task Vector Composition for Ethical Preference Alignment in Language Models 5 upvotes, #27 of 2026-09-21
- Training-Free Speech-Centric Omni Understanding with Frozen VLMs 7 upvotes, #23 of 2026-09-07
- Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding 19 upvotes, #14 of 2026-08-31
- Can Dialects Be Steered Like Languages? Sparse Neurons and Distributed Directions in Arabic LLMs 4 upvotes, #20 of 2026-07-10
- A Gravitational Interpretation of Fine-Tuning Reversion 4 upvotes, #43 of 2026-06-30
- CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization 14 upvotes, #15 of 2026-05-20
- DocAtlas: Multilingual Document Understanding Across 80+ Languages 4 upvotes, #40 of 2026-05-20
- SafeDiffusion-R1: Online Reward Steering for Safe Diffusion Post-Training 5 upvotes, #34 of 2026-05-19
- Efficient Image Synthesis with Sphere Latent Encoder 8 upvotes, #24 of 2026-05-18
- Reliable Chain-of-Thought via Prefix Consistency 1 upvotes, #56 of 2026-05-13
- Can Muon Fine-tune Adam-Pretrained Models? 6 upvotes, #34 of 2026-05-12
- Instruction-Guided Poetry Generation in Arabic and Its Dialects 3 upvotes, #21 of 2026-05-01
- When Background Matters: Breaking Medical Vision Language Models by Transferable Attack 3 upvotes, #31 of 2026-04-21
- Counting to Four is still a Chore for VLMs 2 upvotes, #40 of 2026-04-14
- Paper Circle: An Open-source Multi-agent Research Discovery and Analysis Framework 30 upvotes, #13 of 2026-04-08
- LinguDistill: Recovering Linguistic Ability in Vision- Language Models via Selective Cross-Modal Distillation 8 upvotes, #31 of 2026-04-03
- CarePilot: A Multi-Agent Framework for Long-Horizon Computer Task Automation in Healthcare 10 upvotes, #13 of 2026-03-26
- From Masks to Pixels and Meaning: A New Taxonomy, Benchmark, and Metrics for VLM Image Tampering 1 upvotes, #36 of 2026-03-23
- SWE-Skills-Bench: Do Agent Skills Actually Help in Real-World Software Engineering? 18 upvotes, #18 of 2026-03-18
- MediX-R1: Open Ended Medical Reinforcement Learning 22 upvotes, #8 of 2026-02-27
- Mobile-O: Unified Multimodal Understanding and Generation on Mobile Device 23 upvotes, #6 of 2026-02-24
- Sink-Aware Pruning for Diffusion Language Models 3 upvotes, #10 of 2026-02-23
- A Benchmark and Agentic Framework for Omni-Modal Reasoning and Tool Use in Long Videos 4 upvotes, #21 of 2025-12-22
- Robust and Calibrated Detection of Authentic Multimedia Content 15 upvotes, #14 of 2025-12-18
- Do LLMs "Feel"? Emotion Circuits Discovery and Control 4 upvotes, #22 of 2025-10-20
- Attention Is All You Need for KV Cache in Diffusion LLMs 35 upvotes, #10 of 2025-10-17
- Character Mixing for Video Generation 5 upvotes, #22 of 2025-10-07
- Guaranteed Guess: A Language Modeling Approach for CISC-to-RISC Transpilation with Testing Guarantees 11 upvotes, #15 of 2025-06-18
- VideoMolmo: Spatio-Temporal Grounding Meets Pointing 10 upvotes, #16 of 2025-06-18
- SVRPBench: A Realistic Benchmark for Stochastic Vehicle Routing Problem 14 upvotes, #21 of 2025-05-29
- CASS: Nvidia to AMD Transpilation with Data, Models, and Benchmark 1 upvotes, #66 of 2025-05-27
- KITAB-Bench: A Comprehensive Multi-Domain Benchmark for Arabic OCR and Document Understanding 6 upvotes, #18 of 2025-02-24
- AIN: The Arabic INclusive Large Multimodal Model 15 upvotes, #12 of 2025-02-04
- LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs 57 upvotes, #3 of 2025-01-13
- From CISC to RISC: language-model guided assembly transpilation 11 upvotes, #16 of 2024-11-26
- VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos 20 upvotes, #8 of 2024-11-08
- CAMEL-Bench: A Comprehensive Arabic LMM Benchmark 8 upvotes, #17 of 2024-10-25
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.