Baichuan 2: Open Large-scale Language Models

Aiyuan Yang, BinXiao, Bingning Wang, Borong Zhang, Chao Yin, Chenxu Lv, Da Pan, Hellen Wang, Dong Yan, FanYang, Fei Deng, Feng Wang, Feng Liu, Guangwei Ai, Guosheng Dong Haizhou Zhao, Hang Xu, Alvin Sun, zhang hongda, Hui Liu, JiJiaming, Jian Xie, Juntao Dai, Kun Fang, Lei Su Liang Song, lili, Liyun Ru, Luyao Ma, Mang Wang, Mickel Liu, MingAn Lin, Nuolan Nie, Peidong Guo, Ruiyang Sun, Tao Zhang, Tianpeng Li, Tianyu Li, Wei Cheng, Chen, Xiangrong Zeng, 5555, xiaoxi, Xin Men, Xin Yu, Xuehai Pan, shenyanjun, Yiding Wang, Yiyu Li, Youxin Jiang, yuchen gao, Yupeng Zhang, zhouzenan, Zhiying Wu

Baichuan 2: Open Large-scale Language Models: 21 upvotes on Hugging Face Daily Papers, #4 of 9 papers on 2023-09-20. Day-by-day upvote history.

Large language models (LLMs) have demonstrated remarkable performance on a variety of natural language tasks based on just a few examples of natural language instructions, reducing the need for extensive feature engineering. However, most powerful LLMs are closed-source or limited in their capability for languages other than English. In this technical report, we present Baichuan 2, a series of large-scale multilingual language models containing 7 billion and 13 billion parameters, trained from scratch, on 2.6 trillion tokens. Baichuan 2 matches or outperforms other open-source models of similar size on public benchmarks like MMLU, CMMLU, GSM8K, and HumanEval. Furthermore, Baichuan 2 excels in vertical domains such as medicine and law. We will release all pre-training model checkpoints to benefit the research community in better understanding the training dynamics of Baichuan 2.

Paper page on Hugging Face · arXiv

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.