Jonas Geiping
Jonas Geiping on Hugging Face Daily Papers: 30 papers, 7 in the top 3 of their day, 572 upvotes.
- FutureSim: Replaying World Events to Evaluate Adaptive Agents 7 upvotes, #35 of 2026-05-15
- Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs 17 upvotes, #17 of 2026-05-13
- NESSiE: The Necessary Safety Benchmark -- Identifying Errors that should not Exist 3 upvotes, #17 of 2026-02-20
- Scaling Open-Ended Reasoning to Predict the Future 15 upvotes, #11 of 2026-01-01
- Training AI Co-Scientists Using Rubric Rewards 17 upvotes, #13 of 2025-12-30
- Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence 15 upvotes, #12 of 2025-11-11
- Rewiring Experts on the Fly:Continuous Rerouting for Better Online Adaptation in Mixture-of-Expert models 3 upvotes, #24 of 2025-10-20
- Efficient Parallel Samplers for Recurrent-Depth Models and Their Connection to Diffusion Language Models 6 upvotes, #31 of 2025-10-17
- Adaptive Attacks on Trusted Monitors Subvert AI Control Protocols 5 upvotes, #29 of 2025-10-13
- Training Dynamics Impact Post-Training Quantization Robustness 2 upvotes, #34 of 2025-10-08
- Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLM 10 upvotes, #15 of 2025-09-23
- The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs 33 upvotes, #1 of 2025-09-15
- MedSAMix: A Training-Free Model Merging Approach for Medical Image Segmentation 2 upvotes, #23 of 2025-08-20
- Answer Matching Outperforms Multiple Choice for Language Model Evaluation 8 upvotes, #11 of 2025-07-03
- GPTailor: Large Language Model Pruning Through Layer Cutting and Stitching 7 upvotes, #12 of 2025-06-26
- Pitfalls in Evaluating Language Model Forecasters 3 upvotes, #43 of 2025-06-03
- Capability-Based Scaling Laws for LLM Red-Teaming 3 upvotes, #57 of 2025-05-28
- Can Language Models Falsify? Evaluating Algorithmic Reasoning with Counterexample Creation 18 upvotes, #11 of 2025-02-27
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach 105 upvotes, #1 of 2025-02-10
- Great Models Think Alike and this Undermines AI Oversight 25 upvotes, #5 of 2025-02-07
- Be like a Goldfish, Don't Memorize! Mitigating Memorization in Generative LLMs 8 upvotes, #13 of 2024-06-17
- Transformers Can Do Arithmetic with the Right Embeddings 49 upvotes, #2 of 2024-05-28
- Measuring Style Similarity in Diffusion Models 13 upvotes, #6 of 2024-04-02
- Coercing LLMs to do and reveal (almost) anything 13 upvotes, #7 of 2024-02-22
- Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text 45 upvotes, #1 of 2024-01-23
- Object Recognition as Next Token Prediction 13 upvotes, #10 of 2023-12-05
- Bring Your Own Data! Self-Supervised Evaluation for Large Language Models 16 upvotes, #4 of 2023-06-26
- On the Reliability of Watermarks for Large Language Models 6 upvotes, #2 of 2023-06-08
- Understanding and Mitigating Copying in Diffusion Models 3 upvotes, #3 of 2023-06-01
- Tree-Ring Watermarks: Fingerprints for Diffusion Images that are Invisible and Robust 9 upvotes, #1 of 2023-06-01
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.