Jonas Geiping

Jonas Geiping on Hugging Face Daily Papers: 30 papers, 7 in the top 3 of their day, 572 upvotes.

  1. FutureSim: Replaying World Events to Evaluate Adaptive Agents 7 upvotes, #35 of 2026-05-15
  2. Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs 17 upvotes, #17 of 2026-05-13
  3. NESSiE: The Necessary Safety Benchmark -- Identifying Errors that should not Exist 3 upvotes, #17 of 2026-02-20
  4. Scaling Open-Ended Reasoning to Predict the Future 15 upvotes, #11 of 2026-01-01
  5. Training AI Co-Scientists Using Rubric Rewards 17 upvotes, #13 of 2025-12-30
  6. Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence 15 upvotes, #12 of 2025-11-11
  7. Rewiring Experts on the Fly:Continuous Rerouting for Better Online Adaptation in Mixture-of-Expert models 3 upvotes, #24 of 2025-10-20
  8. Efficient Parallel Samplers for Recurrent-Depth Models and Their Connection to Diffusion Language Models 6 upvotes, #31 of 2025-10-17
  9. Adaptive Attacks on Trusted Monitors Subvert AI Control Protocols 5 upvotes, #29 of 2025-10-13
  10. Training Dynamics Impact Post-Training Quantization Robustness 2 upvotes, #34 of 2025-10-08
  11. Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLM 10 upvotes, #15 of 2025-09-23
  12. The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs 33 upvotes, #1 of 2025-09-15
  13. MedSAMix: A Training-Free Model Merging Approach for Medical Image Segmentation 2 upvotes, #23 of 2025-08-20
  14. Answer Matching Outperforms Multiple Choice for Language Model Evaluation 8 upvotes, #11 of 2025-07-03
  15. GPTailor: Large Language Model Pruning Through Layer Cutting and Stitching 7 upvotes, #12 of 2025-06-26
  16. Pitfalls in Evaluating Language Model Forecasters 3 upvotes, #43 of 2025-06-03
  17. Capability-Based Scaling Laws for LLM Red-Teaming 3 upvotes, #57 of 2025-05-28
  18. Can Language Models Falsify? Evaluating Algorithmic Reasoning with Counterexample Creation 18 upvotes, #11 of 2025-02-27
  19. Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach 105 upvotes, #1 of 2025-02-10
  20. Great Models Think Alike and this Undermines AI Oversight 25 upvotes, #5 of 2025-02-07
  21. Be like a Goldfish, Don't Memorize! Mitigating Memorization in Generative LLMs 8 upvotes, #13 of 2024-06-17
  22. Transformers Can Do Arithmetic with the Right Embeddings 49 upvotes, #2 of 2024-05-28
  23. Measuring Style Similarity in Diffusion Models 13 upvotes, #6 of 2024-04-02
  24. Coercing LLMs to do and reveal (almost) anything 13 upvotes, #7 of 2024-02-22
  25. Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text 45 upvotes, #1 of 2024-01-23
  26. Object Recognition as Next Token Prediction 13 upvotes, #10 of 2023-12-05
  27. Bring Your Own Data! Self-Supervised Evaluation for Large Language Models 16 upvotes, #4 of 2023-06-26
  28. On the Reliability of Watermarks for Large Language Models 6 upvotes, #2 of 2023-06-08
  29. Understanding and Mitigating Copying in Diffusion Models 3 upvotes, #3 of 2023-06-01
  30. Tree-Ring Watermarks: Fingerprints for Diffusion Images that are Invisible and Robust 9 upvotes, #1 of 2023-06-01

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.