MH-MoE:Multi-Head Mixture-of-Experts

HUANG SHAOHAN, Xun Wu, Shuming Ma, Furu Wei

MH-MoE:Multi-Head Mixture-of-Experts: 26 upvotes on Hugging Face Daily Papers, #7 of 24 papers on 2024-11-26. Day-by-day upvote history.

Multi-Head Mixture-of-Experts (MH-MoE) demonstrates superior performance by using the multi-head mechanism to collectively attend to information from various representation spaces within different experts. In this paper, we present a novel implementation of MH-MoE that maintains both FLOPs and parameter parity with sparse Mixture of Experts models. Experimental results on language models show that the new implementation yields quality improvements over both vanilla MoE and fine-grained MoE models. Additionally, our experiments demonstrate that MH-MoE is compatible with 1-bit Large Language Models (LLMs) such as BitNet.

Paper page on Hugging Face · arXiv

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.