Daily Papers of 2024-06-24

  1. LongRAG: Enhancing Retrieval-Augmented Generation with Long-context LLMs 54 upvotes, #1 of 2024-06-24
  2. BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions 41 upvotes, #2 of 2024-06-24
  3. Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges 34 upvotes, #3 of 2024-06-24
  4. Complexity of Symbolic Representation in Working Memory of Transformer Correlates with the Complexity of a Task 20 upvotes, #4 of 2024-06-24
  5. Towards Retrieval Augmented Generation over Large Video Libraries 18 upvotes, #5 of 2024-06-24
  6. Stylebreeder: Exploring and Democratizing Artistic Styles through Text-to-Image Models 16 upvotes, #6 of 2024-06-24
  7. Evaluating RAG-Fusion with RAGElo: an Automated Elo-based Framework 15 upvotes, #7 of 2024-06-24
  8. MantisScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation 13 upvotes, #8 of 2024-06-24
  9. EvTexture: Event-driven Texture Enhancement for Video Super-Resolution 12 upvotes, #9 of 2024-06-24
  10. Jailbreaking as a Reward Misspecification Problem 12 upvotes, #9 of 2024-06-24
  11. Reward Steering with Evolutionary Heuristics for Decoding-time Alignment 11 upvotes, #11 of 2024-06-24
  12. Two Giraffes in a Dirt Field: Using Game Play to Investigate Situation Modelling in Large Multimodal Models 10 upvotes, #12 of 2024-06-24
  13. Cognitive Map for Language Models: Optimal Planning via Verbally Representing the World Model 10 upvotes, #12 of 2024-06-24
  14. Data Contamination Can Cross Language Barriers 8 upvotes, #14 of 2024-06-24
  15. DELLA-Merging: Reducing Interference in Model Merging through Magnitude-Based Sampling 7 upvotes, #15 of 2024-06-24
  16. Ruby Teaming: Improving Quality Diversity Search with Memory for Automated Red Teaming 6 upvotes, #16 of 2024-06-24
  17. Learning Molecular Representation in a Cell 6 upvotes, #16 of 2024-06-24
  18. 4K4DGen: Panoramic 4D Generation at 4K Resolution 6 upvotes, #16 of 2024-06-24
  19. A Tale of Trust and Accuracy: Base vs. Instruct LLMs in RAG Systems 6 upvotes, #16 of 2024-06-24
  20. Style-NeRF2NeRF: 3D Style Transfer From Style-Aligned Multi-View Images 5 upvotes, #20 of 2024-06-24
  21. NAVSIM: Data-Driven Non-Reactive Autonomous Vehicle Simulation and Benchmarking 5 upvotes, #20 of 2024-06-24
  22. Multimodal Structured Generation: CVPR's 2nd MMFM Challenge Technical Report 4 upvotes, #22 of 2024-06-24
  23. ICAL: Continual Learning of Multimodal Agents by Transforming Trajectories into Actionable Insights 4 upvotes, #22 of 2024-06-24
  24. RE-AdaptIR: Improving Information Retrieval through Reverse Engineered Adaptation 4 upvotes, #22 of 2024-06-24
  25. Low-Resource Machine Translation through the Lens of Personalized Federated Learning 3 upvotes, #25 of 2024-06-24
  26. How Well Do LLMs Represent Values Across Cultures? Empirical Analysis of LLM Responses Based on Hofstede Cultural Dimensions 3 upvotes, #25 of 2024-06-24
  27. ToVo: Toxicity Taxonomy via Voting 3 upvotes, #25 of 2024-06-24

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.