Xin Eric Wang
Xin Eric Wang on Hugging Face Daily Papers: 33 papers, 1 in the top 3 of their day, 458 upvotes.
- Auditing Agent Harness Safety 54 upvotes, #6 of 2026-05-18
- Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling 24 upvotes, #10 of 2026-05-01
- On the Reliability of Computer Use Agents 11 upvotes, #17 of 2026-04-21
- Proactive Agent Research Environment: Simulating Active Users to Evaluate Proactive Assistants 13 upvotes, #15 of 2026-04-02
- Learning Situated Awareness in the Real World 6 upvotes, #16 of 2026-02-19
- Group-Evolving Agents: Open-Ended Self-Improvement via Experience Sharing 8 upvotes, #23 of 2026-02-09
- SafeGround: Know When to Trust GUI Grounding Models via Uncertainty Calibration 4 upvotes, #38 of 2026-02-04
- Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space 2 upvotes, #32 of 2025-12-19
- Presenting a Paper is an Art: Self-Improvement Aesthetic Agents for Academic Presentations 13 upvotes, #13 of 2025-10-08
- The Unreasonable Effectiveness of Scaling Agents for Computer Use 22 upvotes, #13 of 2025-10-03
- VLM4D: Towards Spatiotemporal Awareness in Vision Language Models 6 upvotes, #11 of 2025-08-11
- "PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models 14 upvotes, #16 of 2025-07-22
- Hidden in Plain Sight: Probing Implicit Reasoning in Multimodal Language Models 3 upvotes, #37 of 2025-06-10
- Agents of Change: Self-Evolving LLM Agents for Strategic Planning 6 upvotes, #27 of 2025-06-10
- More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models 14 upvotes, #17 of 2025-06-02
- SafeKey: Amplifying Aha-Moment Insights for Safety Reasoning 7 upvotes, #33 of 2025-05-23
- GRIT: Teaching MLLMs to Think with Images 12 upvotes, #22 of 2025-05-23
- Constructing a 3D Town from a Single Image 23 upvotes, #9 of 2025-05-22
- Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space 13 upvotes, #19 of 2025-05-22
- Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents 19 upvotes, #14 of 2025-04-02
- Multimodal Inconsistency Reasoning (MMIR): A New Benchmark for Multimodal Reasoning Models 15 upvotes, #12 of 2025-02-25
- The Hidden Risks of Large Reasoning Models: A Safety Assessment of R1 5 upvotes, #26 of 2025-02-19
- Agent S: An Open Agentic Framework that Uses Computers Like a Human 24 upvotes, #5 of 2024-10-11
- Multimodal Situational Safety 8 upvotes, #26 of 2024-10-10
- NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models 3 upvotes, #15 of 2024-07-18
- Read Anywhere Pointed: Layout-aware GUI Screen Reading with Tree-of-Lens Grounding 9 upvotes, #13 of 2024-06-28
- VIA: A Spatiotemporal Video Adaptation Framework for Global and Local Video Editing 4 upvotes, #24 of 2024-06-19
- Toffee: Efficient Million-Scale Dataset Construction for Subject-Driven Text-to-Image Generation 4 upvotes, #25 of 2024-06-14
- MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos 23 upvotes, #10 of 2024-06-13
- SwapAnything: Enabling Arbitrary Object Swapping in Personalized Visual Editing 21 upvotes, #3 of 2024-04-09
- Controllable Text-to-Image Generation with GPT-4 4 upvotes, #9 of 2023-05-31
- Photoswap: Personalized Subject Swapping in Images 4 upvotes, #8 of 2023-05-30
- Discriminative Diffusion Models as Few-shot Vision and Language Learners 4 upvotes, #6 of 2023-05-19
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.