Daily Papers of 2026-03-27
- Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale 125 upvotes, #1 of 2026-03-27
- PixelSmile: Toward Fine-Grained Facial Expression Editing 116 upvotes, #2 of 2026-03-27
- Calibri: Enhancing Diffusion Transformers via Parameter-Efficient Calibration 65 upvotes, #3 of 2026-03-27
- Voxtral TTS 56 upvotes, #4 of 2026-03-27
- RealRestorer: Towards Generalizable Real-World Image Restoration with Large-Scale Image Editing Models 55 upvotes, #5 of 2026-03-27
- MSA: Memory Sparse Attention for Efficient End-to-End Memory Model Scaling to 100M Tokens 45 upvotes, #6 of 2026-03-27
- MACRO: Advancing Multi-Reference Image Generation with Structured Long-Context Data 32 upvotes, #7 of 2026-03-27
- SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks 27 upvotes, #8 of 2026-03-27
- AVControl: Efficient Framework for Training Audio-Visual Controls 25 upvotes, #9 of 2026-03-27
- VFIG: Vectorizing Complex Figures in SVG with Vision-Language Models 18 upvotes, #10 of 2026-03-27
- Representation Alignment for Just Image Transformers is not Easier than You Think 13 upvotes, #11 of 2026-03-27
- Less Gaussians, Texture More: 4K Feed-Forward Textured Splatting 13 upvotes, #11 of 2026-03-27
- FinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under the Model Context Protocol 12 upvotes, #13 of 2026-03-27
- MuRF: Unlocking the Multi-Scale Potential of Vision Foundation Models 12 upvotes, #13 of 2026-03-27
- MemMA: Coordinating the Memory Cycle through Multi-Agent Reasoning and In-Situ Self-Evolution 10 upvotes, #15 of 2026-03-27
- AVO: Agentic Variation Operators for Autonomous Evolutionary Search 10 upvotes, #15 of 2026-03-27
- Reaching Beyond the Mode: RL for Distributional Reasoning in Language Models 9 upvotes, #17 of 2026-03-27
- Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes 8 upvotes, #18 of 2026-03-27
- BioVITA: Biological Dataset, Model, and Benchmark for Visual-Textual-Acoustic Alignment 6 upvotes, #19 of 2026-03-27
- S2D2: Fast Decoding for Diffusion LLMs via Training-Free Self-Speculation 6 upvotes, #19 of 2026-03-27
- Vega: Learning to Drive with Natural Language Instructions 6 upvotes, #19 of 2026-03-27
- IQuest-Coder-V1 Technical Report 5 upvotes, #22 of 2026-03-27
- WAFT-Stereo: Warping-Alone Field Transforms for Stereo Matching 5 upvotes, #22 of 2026-03-27
- Pixel-level Scene Understanding in One Token: Visual States Need What-is-Where Composition 4 upvotes, #24 of 2026-03-27
- Nudging Hidden States: Training-Free Model Steering for Chain-of-Thought Reasoning in Large Audio-Language Models 4 upvotes, #24 of 2026-03-27
- Can MLLMs Read Students' Minds? Unpacking Multimodal Error Analysis in Handwritten Math 4 upvotes, #24 of 2026-03-27
- Electrostatic Photoluminescence Tuning in All-Solid-State Perovskite Transistors 4 upvotes, #24 of 2026-03-27
- Extending Precipitation Nowcasting Horizons via Spectral Fusion of Radar Observations and Foundation Model Priors 3 upvotes, #28 of 2026-03-27
- PMT: Plain Mask Transformer for Image and Video Segmentation with Frozen Vision Encoders 2 upvotes, #29 of 2026-03-27
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.