Furu Wei

Furu Wei on Hugging Face Daily Papers: 49 papers, 21 in the top 3 of their day, 2,557 upvotes.

  1. VIBEVOICE-ASR Technical Report 19 upvotes, #11 of 2026-01-27
  2. Multiplex Thinking: Reasoning via Token-wise Branch-and-Merge 38 upvotes, #2 of 2026-01-20
  3. Thinking Augmented Pre-training 22 upvotes, #9 of 2025-09-26
  4. Rectified Sparse Attention 10 upvotes, #19 of 2025-06-05
  5. BitNet v2: Native 4-bit Activations with Hadamard Transformation for 1-bit LLMs 41 upvotes, #3 of 2025-04-28
  6. BitNet b1.58 2B4T Technical Report 66 upvotes, #1 of 2025-04-17
  7. mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data 13 upvotes, #15 of 2025-02-14
  8. Chain-of-Retrieval Augmented Generation 41 upvotes, #2 of 2025-01-27
  9. GeAR: Generation Augmented Retrieval 21 upvotes, #8 of 2025-01-09
  10. Multimodal Latent Language Modeling with Next-Token Diffusion 38 upvotes, #5 of 2024-12-13
  11. MH-MoE:Multi-Head Mixture-of-Experts 22 upvotes, #7 of 2024-11-26
  12. BitNet a4.8: 4-bit Activations for 1-bit LLMs 61 upvotes, #3 of 2024-11-08
  13. Data Selection via Optimal Control for Language Models 8 upvotes, #26 of 2024-10-10
  14. Self-Boosting Large Language Models with Synthetic Preference Data 14 upvotes, #15 of 2024-10-10
  15. Differential Transformer 148 upvotes, #1 of 2024-10-08
  16. Autoregressive Speech Synthesis without Vector Quantization 12 upvotes, #11 of 2024-07-12
  17. Direct Preference Knowledge Distillation for Large Language Models 21 upvotes, #3 of 2024-07-01
  18. Instruction Pre-Training: Language Models are Supervised Multitask Learners 74 upvotes, #2 of 2024-06-21
  19. VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers 11 upvotes, #5 of 2024-06-11
  20. MoRA: High-Rank Updating for Parameter-Efficient Fine-Tuning 39 upvotes, #2 of 2024-05-21
  21. Multi-Head Mixture-of-Experts 45 upvotes, #2 of 2024-04-24
  22. MathScale: Scaling Instruction Tuning for Mathematical Reasoning 13 upvotes, #6 of 2024-03-06
  23. Towards Optimal Learning of Language Models 18 upvotes, #10 of 2024-02-28
  24. The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits 630 upvotes, #1 of 2024-02-28
  25. Synthetic Data (Almost) from Scratch: Generalized Instruction Tuning for Language Models 52 upvotes, #2 of 2024-02-21
  26. Generative Representational Instruction Tuning 54 upvotes, #2 of 2024-02-16
  27. Multilingual E5 Text Embeddings: A Technical Report 23 upvotes, #5 of 2024-02-09
  28. K-Level Reasoning with Large Language Models 18 upvotes, #8 of 2024-02-05
  29. Improving Text Embeddings with Large Language Models 84 upvotes, #1 of 2024-01-02
  30. Democratizing Reasoning Ability: Tailored Learning from Large Language Model 16 upvotes, #5 of 2023-10-23
  31. Tuna: Instruction Tuning using Feedback from Large Language Models 10 upvotes, #10 of 2023-10-23
  32. BitNet: Scaling 1-bit Transformers for Large Language Models 108 upvotes, #1 of 2023-10-18
  33. Calibrating LLM-Based Evaluator 12 upvotes, #5 of 2023-09-26
  34. Kosmos-2.5: A Multimodal Literate Model 56 upvotes, #4 of 2023-09-21
  35. Adapting Large Language Models via Reading Comprehension 82 upvotes, #2 of 2023-09-19
  36. Large Language Model for Science: A Study on P vs. NP 22 upvotes, #4 of 2023-09-13
  37. Retentive Network: A Successor to Transformer for Large Language Models 173 upvotes, #1 of 2023-07-18
  38. Learning to Retrieve In-Context Examples for Large Language Models 23 upvotes, #3 of 2023-07-17
  39. In-context Autoencoder for Context Compression in a Large Language Model 29 upvotes, #2 of 2023-07-14
  40. Unleashing Cognitive Synergy in Large Language Models: A Task-Solving Agent through Multi-Persona Self-Collaboration 20 upvotes, #4 of 2023-07-12
  41. LongNet: Scaling Transformers to 1,000,000,000 Tokens 82 upvotes, #2 of 2023-07-06
  42. Kosmos-2: Grounding Multimodal Large Language Models to the World 36 upvotes, #1 of 2023-06-27
  43. Knowledge Distillation of Large Language Models 24 upvotes, #5 of 2023-06-16
  44. Augmenting Language Models with Long-Term Memory 19 upvotes, #4 of 2023-06-13
  45. Dual-Alignment Pre-training for Cross-lingual Sentence Embedding 1 upvotes, #10 of 2023-05-17
  46. Pre-Training to Learn in Context 2 upvotes, #8 of 2023-05-17
  47. Not All Languages Are Created Equal in LLMs: Improving Multilingual Capability by Cross-Lingual-Thought Prompting 1 upvotes, #9 of 2023-05-12
  48. Chain-of-Dictionary Prompting Elicits Translation in Large Language Models 2 upvotes, #7 of 2023-05-12
  49. Pre-training Language Model as a Multi-perspective Course Learner 1 upvotes, #6 of 2023-05-09

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.