wangjunjie
wangjunjie on Hugging Face Daily Papers: 16 papers, 4 in the top 3 of their day, 861 upvotes.
- E2A-Bench: Benchmarking Evidence-to-Action Reliability in Financial Chart Reasoning 3 upvotes, #30 of 2026-09-15
- Learning Where Outcomes Change:Credit-Addressable Reasoning for Multimodal Geometry 9 upvotes, #18 of 2026-09-02
- AutoResearch: Insight In, Hallucination Out 10 upvotes, #19 of 2026-08-25
- LongWoF-Bench: Evaluating EvoMap Genes for Verifiable Long-Workflow Tasks 7 upvotes, #21 of 2026-08-25
- Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO 26 upvotes, #11 of 2026-06-15
- WebRISE: Requirement-Induced State Evaluation for MLLM-Generated Web Artifacts 8 upvotes, #23 of 2026-06-04
- SIN-Bench: Tracing Native Evidence Chains in Long-Context Multimodal Scientific Interleaved Literature 7 upvotes, #8 of 2026-01-20
- ProFit: Leveraging High-Value Signals in SFT via Probability-Guided Token Selection 15 upvotes, #13 of 2026-01-19
- A.S.E: A Repository-Level Benchmark for Evaluating Security in AI-Generated Code 214 upvotes, #1 of 2025-09-01
- VeriGUI: Verifiable Long-Chain GUI Dataset 137 upvotes, #2 of 2025-08-07
- AnyCap Project: A Unified Framework, Dataset, and Benchmark for Controllable Omni-modal Captioning 37 upvotes, #6 of 2025-07-18
- YuE: Scaling Open Foundation Models for Long-Form Music Generation 57 upvotes, #3 of 2025-03-12
- PIN: A Knowledge-Intensive Dataset for Paired and Interleaved Multimodal Documents 20 upvotes, #10 of 2024-06-21
- ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code Generation 52 upvotes, #3 of 2024-06-17
- StructLM: Towards Building Generalist Models for Structured Knowledge Grounding 27 upvotes, #7 of 2024-02-27
- CMMMU: A Chinese Massive Multi-discipline Multimodal Understanding Benchmark 27 upvotes, #4 of 2024-01-23
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.