Knesset-DictaBERT: A Hebrew Language Model for Parliamentary Proceedings

Gili Goldin, Shuly Wintner

Knesset-DictaBERT: A Hebrew Language Model for Parliamentary Proceedings: 24 upvotes on Hugging Face Daily Papers, #5 of 10 papers on 2024-07-31. Day-by-day upvote history.

We present Knesset-DictaBERT, a large Hebrew language model fine-tuned on the Knesset Corpus, which comprises Israeli parliamentary proceedings. The model is based on the DictaBERT architecture and demonstrates significant improvements in understanding parliamentary language according to the MLM task. We provide a detailed evaluation of the model's performance, showing improvements in perplexity and accuracy over the baseline DictaBERT model.

Paper page on Hugging Face · arXiv

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.