Rameswar Panda
Rameswar Panda on Hugging Face Daily Papers: 5 papers, 1 in the top 3 of their day, 116 upvotes.
- Power Scheduler: A Batch Size and Token Number Agnostic Learning Rate Scheduler 18 upvotes, #8 of 2024-08-27
- Scaling Granite Code Models to 128K Context 14 upvotes, #4 of 2024-07-19
- Self-MoE: Towards Compositional Large Language Models with Self-Specialized Experts 12 upvotes, #4 of 2024-06-20
- Reducing Transformer Key-Value Cache Size with Cross-Layer Attention 23 upvotes, #3 of 2024-05-22
- Data Engineering for Scaling Language Models to 128K Context 24 upvotes, #7 of 2024-02-16
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.