Daily Papers of 2024-01-23
- Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text 45 upvotes, #1 of 2024-01-23
- Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs 30 upvotes, #2 of 2024-01-23
- SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities 30 upvotes, #2 of 2024-01-23
- CMMMU: A Chinese Massive Multi-discipline Multimodal Understanding Benchmark 27 upvotes, #4 of 2024-01-23
- Scalable High-Resolution Pixel-Space Image Synthesis with Hourglass Diffusion Transformers 23 upvotes, #5 of 2024-01-23
- CheXagent: Towards a Foundation Model for Chest X-Ray Interpretation 22 upvotes, #6 of 2024-01-23
- DITTO: Diffusion Inference-Time T-Optimization for Music Generation 21 upvotes, #7 of 2024-01-23
- WARM: On the Benefits of Weight Averaged Reward Models 19 upvotes, #8 of 2024-01-23
- Make-A-Shape: a Ten-Million-scale 3D Shape Model 17 upvotes, #9 of 2024-01-23
- EmerDiff: Emerging Pixel-level Semantic Knowledge in Diffusion Models 17 upvotes, #9 of 2024-01-23
- StreamVoice: Streamable Context-Aware Language Modeling for Real-time Zero-Shot Voice Conversion 11 upvotes, #11 of 2024-01-23
- OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics 10 upvotes, #12 of 2024-01-23
- UltrAvatar: A Realistic Animatable 3D Avatar Diffusion Model with Authenticity Guided Textures 7 upvotes, #13 of 2024-01-23
- Single-View 3D Human Digitalization with Large Reconstruction Models 6 upvotes, #14 of 2024-01-23
- Scaling Face Interaction Graph Networks to Real World Scenes 3 upvotes, #15 of 2024-01-23
- Fast Registration of Photorealistic Avatars for VR Facial Animation 2 upvotes, #16 of 2024-01-23
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.