Daily Papers of 2024-03-07
- GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection 150 upvotes, #1 of 2024-03-07
- ShortGPT: Layers in Large Language Models are More Redundant Than You Expect 56 upvotes, #2 of 2024-03-07
- SaulLM-7B: A pioneering Large Language Model for Law 54 upvotes, #3 of 2024-03-07
- Learning to Decode Collaboratively with Multiple Language Models 17 upvotes, #4 of 2024-03-07
- Enhancing Vision-Language Pre-training with Rich Supervisions 11 upvotes, #5 of 2024-03-07
- Stop Regressing: Training Value Functions via Classification for Scalable Deep RL 10 upvotes, #6 of 2024-03-07
- 3D Diffusion Policy 10 upvotes, #6 of 2024-03-07
- Backtracing: Retrieving the Cause of the Query 9 upvotes, #8 of 2024-03-07
- Caduceus: Bi-Directional Equivariant Long-Range DNA Sequence Modeling 8 upvotes, #9 of 2024-03-07
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.