Piccolo2: General Text Embedding with Multi-task Hybrid Loss Training

HuangJunqin, Zhongjie Hu, Zihao Jing, Gaomengya, Yichao Wu

Piccolo2: General Text Embedding with Multi-task Hybrid Loss Training: 18 upvotes on Hugging Face Daily Papers, #5 of 9 papers on 2024-05-14. Day-by-day upvote history.

In this report, we introduce Piccolo2, an embedding model that surpasses other models in the comprehensive evaluation over 6 tasks on CMTEB benchmark, setting a new state-of-the-art. Piccolo2 primarily leverages an efficient multi-task hybrid loss training approach, effectively harnessing textual data and labels from diverse downstream tasks. In addition, Piccolo2 scales up the embedding dimension and uses MRL training to support more flexible vector dimensions. The latest information of piccolo models can be accessed via: https://huggingface.co/sensenova/

Paper page on Hugging Face · arXiv

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.