Furu Wei
Furu Wei on Hugging Face Daily Papers: 49 papers, 21 in the top 3 of their day, 2,557 upvotes.
- VIBEVOICE-ASR Technical Report 19 upvotes, #11 of 2026-01-27
- Multiplex Thinking: Reasoning via Token-wise Branch-and-Merge 38 upvotes, #2 of 2026-01-20
- Thinking Augmented Pre-training 22 upvotes, #9 of 2025-09-26
- Rectified Sparse Attention 10 upvotes, #19 of 2025-06-05
- BitNet v2: Native 4-bit Activations with Hadamard Transformation for 1-bit LLMs 41 upvotes, #3 of 2025-04-28
- BitNet b1.58 2B4T Technical Report 66 upvotes, #1 of 2025-04-17
- mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data 13 upvotes, #15 of 2025-02-14
- Chain-of-Retrieval Augmented Generation 41 upvotes, #2 of 2025-01-27
- GeAR: Generation Augmented Retrieval 21 upvotes, #8 of 2025-01-09
- Multimodal Latent Language Modeling with Next-Token Diffusion 38 upvotes, #5 of 2024-12-13
- MH-MoE:Multi-Head Mixture-of-Experts 22 upvotes, #7 of 2024-11-26
- BitNet a4.8: 4-bit Activations for 1-bit LLMs 61 upvotes, #3 of 2024-11-08
- Data Selection via Optimal Control for Language Models 8 upvotes, #26 of 2024-10-10
- Self-Boosting Large Language Models with Synthetic Preference Data 14 upvotes, #15 of 2024-10-10
- Differential Transformer 148 upvotes, #1 of 2024-10-08
- Autoregressive Speech Synthesis without Vector Quantization 12 upvotes, #11 of 2024-07-12
- Direct Preference Knowledge Distillation for Large Language Models 21 upvotes, #3 of 2024-07-01
- Instruction Pre-Training: Language Models are Supervised Multitask Learners 74 upvotes, #2 of 2024-06-21
- VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers 11 upvotes, #5 of 2024-06-11
- MoRA: High-Rank Updating for Parameter-Efficient Fine-Tuning 39 upvotes, #2 of 2024-05-21
- Multi-Head Mixture-of-Experts 45 upvotes, #2 of 2024-04-24
- MathScale: Scaling Instruction Tuning for Mathematical Reasoning 13 upvotes, #6 of 2024-03-06
- Towards Optimal Learning of Language Models 18 upvotes, #10 of 2024-02-28
- The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits 630 upvotes, #1 of 2024-02-28
- Synthetic Data (Almost) from Scratch: Generalized Instruction Tuning for Language Models 52 upvotes, #2 of 2024-02-21
- Generative Representational Instruction Tuning 54 upvotes, #2 of 2024-02-16
- Multilingual E5 Text Embeddings: A Technical Report 23 upvotes, #5 of 2024-02-09
- K-Level Reasoning with Large Language Models 18 upvotes, #8 of 2024-02-05
- Improving Text Embeddings with Large Language Models 84 upvotes, #1 of 2024-01-02
- Democratizing Reasoning Ability: Tailored Learning from Large Language Model 16 upvotes, #5 of 2023-10-23
- Tuna: Instruction Tuning using Feedback from Large Language Models 10 upvotes, #10 of 2023-10-23
- BitNet: Scaling 1-bit Transformers for Large Language Models 108 upvotes, #1 of 2023-10-18
- Calibrating LLM-Based Evaluator 12 upvotes, #5 of 2023-09-26
- Kosmos-2.5: A Multimodal Literate Model 56 upvotes, #4 of 2023-09-21
- Adapting Large Language Models via Reading Comprehension 82 upvotes, #2 of 2023-09-19
- Large Language Model for Science: A Study on P vs. NP 22 upvotes, #4 of 2023-09-13
- Retentive Network: A Successor to Transformer for Large Language Models 173 upvotes, #1 of 2023-07-18
- Learning to Retrieve In-Context Examples for Large Language Models 23 upvotes, #3 of 2023-07-17
- In-context Autoencoder for Context Compression in a Large Language Model 29 upvotes, #2 of 2023-07-14
- Unleashing Cognitive Synergy in Large Language Models: A Task-Solving Agent through Multi-Persona Self-Collaboration 20 upvotes, #4 of 2023-07-12
- LongNet: Scaling Transformers to 1,000,000,000 Tokens 82 upvotes, #2 of 2023-07-06
- Kosmos-2: Grounding Multimodal Large Language Models to the World 36 upvotes, #1 of 2023-06-27
- Knowledge Distillation of Large Language Models 24 upvotes, #5 of 2023-06-16
- Augmenting Language Models with Long-Term Memory 19 upvotes, #4 of 2023-06-13
- Dual-Alignment Pre-training for Cross-lingual Sentence Embedding 1 upvotes, #10 of 2023-05-17
- Pre-Training to Learn in Context 2 upvotes, #8 of 2023-05-17
- Not All Languages Are Created Equal in LLMs: Improving Multilingual Capability by Cross-Lingual-Thought Prompting 1 upvotes, #9 of 2023-05-12
- Chain-of-Dictionary Prompting Elicits Translation in Large Language Models 2 upvotes, #7 of 2023-05-12
- Pre-training Language Model as a Multi-perspective Course Learner 1 upvotes, #6 of 2023-05-09
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.