Daily Papers of 2024-06-24
- LongRAG: Enhancing Retrieval-Augmented Generation with Long-context LLMs 54 upvotes, #1 of 2024-06-24
- BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions 41 upvotes, #2 of 2024-06-24
- Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges 34 upvotes, #3 of 2024-06-24
- Complexity of Symbolic Representation in Working Memory of Transformer Correlates with the Complexity of a Task 20 upvotes, #4 of 2024-06-24
- Towards Retrieval Augmented Generation over Large Video Libraries 18 upvotes, #5 of 2024-06-24
- Stylebreeder: Exploring and Democratizing Artistic Styles through Text-to-Image Models 16 upvotes, #6 of 2024-06-24
- Evaluating RAG-Fusion with RAGElo: an Automated Elo-based Framework 15 upvotes, #7 of 2024-06-24
- MantisScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation 13 upvotes, #8 of 2024-06-24
- EvTexture: Event-driven Texture Enhancement for Video Super-Resolution 12 upvotes, #9 of 2024-06-24
- Jailbreaking as a Reward Misspecification Problem 12 upvotes, #9 of 2024-06-24
- Reward Steering with Evolutionary Heuristics for Decoding-time Alignment 11 upvotes, #11 of 2024-06-24
- Two Giraffes in a Dirt Field: Using Game Play to Investigate Situation Modelling in Large Multimodal Models 10 upvotes, #12 of 2024-06-24
- Cognitive Map for Language Models: Optimal Planning via Verbally Representing the World Model 10 upvotes, #12 of 2024-06-24
- Data Contamination Can Cross Language Barriers 8 upvotes, #14 of 2024-06-24
- DELLA-Merging: Reducing Interference in Model Merging through Magnitude-Based Sampling 7 upvotes, #15 of 2024-06-24
- Ruby Teaming: Improving Quality Diversity Search with Memory for Automated Red Teaming 6 upvotes, #16 of 2024-06-24
- Learning Molecular Representation in a Cell 6 upvotes, #16 of 2024-06-24
- 4K4DGen: Panoramic 4D Generation at 4K Resolution 6 upvotes, #16 of 2024-06-24
- A Tale of Trust and Accuracy: Base vs. Instruct LLMs in RAG Systems 6 upvotes, #16 of 2024-06-24
- Style-NeRF2NeRF: 3D Style Transfer From Style-Aligned Multi-View Images 5 upvotes, #20 of 2024-06-24
- NAVSIM: Data-Driven Non-Reactive Autonomous Vehicle Simulation and Benchmarking 5 upvotes, #20 of 2024-06-24
- Multimodal Structured Generation: CVPR's 2nd MMFM Challenge Technical Report 4 upvotes, #22 of 2024-06-24
- ICAL: Continual Learning of Multimodal Agents by Transforming Trajectories into Actionable Insights 4 upvotes, #22 of 2024-06-24
- RE-AdaptIR: Improving Information Retrieval through Reverse Engineered Adaptation 4 upvotes, #22 of 2024-06-24
- Low-Resource Machine Translation through the Lens of Personalized Federated Learning 3 upvotes, #25 of 2024-06-24
- How Well Do LLMs Represent Values Across Cultures? Empirical Analysis of LLM Responses Based on Hofstede Cultural Dimensions 3 upvotes, #25 of 2024-06-24
- ToVo: Toxicity Taxonomy via Voting 3 upvotes, #25 of 2024-06-24
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.