Multilingual E5 Text Embeddings: A Technical Report
Liang Wang, nyanyanya, Xiaolong Huang, Linjun Yang, Rangan Majumder, Furu Wei
Multilingual E5 Text Embeddings: A Technical Report: 23 upvotes on Hugging Face Daily Papers, #5 of 15 papers on 2024-02-09. Day-by-day upvote history.
This technical report presents the training methodology and evaluation results of the open-source multilingual E5 text embedding models, released in mid-2023. Three embedding models of different sizes (small / base / large) are provided, offering a balance between the inference efficiency and embedding quality. The training procedure adheres to the English E5 model recipe, involving contrastive pre-training on 1 billion multilingual text pairs, followed by fine-tuning on a combination of labeled datasets. Additionally, we introduce a new instruction-tuned embedding model, whose performance is on par with state-of-the-art, English-only models of similar sizes. Information regarding the model release can be found at https://github.com/microsoft/unilm/tree/master/e5 .
Paper page on Hugging Face · arXiv
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.