---
title: "EuroBERT: Scaling Multilingual Encoders for European Languages"
canonical_url: "https://www.modelscope.cn/papers/124319"
md_url: "https://www.modelscope.cn/papers/124319.md"
arxiv_id: 2503.05500
published: 2026-06-01
last_updated: 2026-06-01
authors:
  - "Nicolas Boizard"
  - "Hippolyte Gisserot-Boukhlef"
  - "Duarte M. Alves"
  - "André Martins"
  - "Ayoub Hammal"
  - "Caio Corro"
  - "Céline Hudelot"
  - "Emmanuel Malherbe"
  - "Etienne Malaboeuf"
  - "Fanny Jourdan"
  - "Gabriel Hautreux"
  - "João Alves"
  - "Kevin El Haddad"
  - "Manuel Faysse"
  - "Maxime Peyrard"
  - "Nuno M. Guerreiro"
  - "Patrick Fernandes"
  - "Ricardo Rei"
  - "Pierre Colombo"
model_name: EuroBERT
model_developer: "CentraleSupélec, Instituto Superior Técnico & Universidade de Lisboa, CNRS, and others (European academic and research institutions)"
domain:
  - "自然语言处理"
  - "深度学习"
  - "多语言模型"
  - "信息检索"
  - "代码理解"
  - "数学推理"
type:
  - "自然语言处理"
  - "深度学习"
  - "多语言模型"
  - "信息检索"
  - "代码理解"
  - "数学推理"
  - "Computation and Language (cs.CL)"
  - "Artificial Intelligence (cs.AI)"
arxiv_url: "https://arxiv.org/abs/2503.05500v3"
pdf_url: "https://arxiv.org/pdf/2503.05500v3.pdf"
code_link: "https://huggingface.co/EuroBERT"
---

# EuroBERT: Scaling Multilingual Encoders for European Languages

> General-purpose multilingual vector representations, used in retrieval, regression and classification, are traditionally obtained from bidirectional encoder models. Despite their wide applicability, encoders have been recently overshadowed by advances in…

「EuroBERT: Scaling Multilingual Encoders for European Languages」是 ModelScope 魔搭社区收录的论文，arXiv 2503.05500，作者为 Nicolas Boizard, Hippolyte Gisserot-Boukhlef, Duarte M. Alves et al.，发表于 2026-06-01，属于 自然语言处理、深度学习、多语言模型 领域。

- **ArXiv**: 2503.05500
- **Published**: 2026-06-01
- **Authors**: Nicolas Boizard, Hippolyte Gisserot-Boukhlef, Duarte M. Alves, André Martins, Ayoub Hammal, Caio Corro, Céline Hudelot, Emmanuel Malherbe, Etienne Malaboeuf, Fanny Jourdan, Gabriel Hautreux, João Alves, Kevin El Haddad, Manuel Faysse, Maxime Peyrard, Nuno M. Guerreiro, Patrick Fernandes, Ricardo Rei, Pierre Colombo
- **Model**: EuroBERT
- **Developer**: CentraleSupélec, Instituto Superior Técnico & Universidade de Lisboa, CNRS, and others (European academic and research institutions)
- **Domain**: 自然语言处理, 深度学习, 多语言模型, 信息检索, 代码理解, 数学推理
- **ArXiv URL**: https://arxiv.org/abs/2503.05500v3
- **PDF**: https://arxiv.org/pdf/2503.05500v3.pdf
- **Code**: https://huggingface.co/EuroBERT

Source: https://www.modelscope.cn/papers/124319

---

> EuroBERT：融合Llama架构与多模态语料的欧洲语言大编码器，重振双向表征学习

## 摘要

本文针对当前多语言编码器发展滞后于生成式解码器的现状，系统性地将解码器领域的先进经验（如架构改进、数据增强、训练策略优化）迁移回双向编码器范式，提出EuroBERT——一个面向欧洲及全球主流语言的新型多语言编码器家族。研究背景在于通用向量表示对检索、分类、回归等NLP任务至关重要，但现有编码器（如XLM-R）已显陈旧。作者创新性地采用Llama-3风格架构（去偏置、分组查询注意力、SwiGLU、RMSNorm、RoPE），构建5T词元多语言语料（含15种语言、38种编程语言、数学公式与证明文本），并设计两阶段训练流程（预训练+退火式微调），动态调整掩码率（50%→10%）、数据分布与上下文长度（最高8,192 tokens）。在MIRACL、XNLI、CodeSearchNet、MathShepherd等20+跨领域基准上，EuroBERT全面超越XLM-R、mGTE、mDeBERTa等强基线，尤其在代码与数学任务上优势显著。研究还通过系统消融揭示了高质量教育数据未必最优、多源异构数据更利于泛化等反直觉发现，并开源全部模型、中间检查点与训练框架。

## Abstract

General-purpose multilingual vector representations, used in retrieval, regression and classification, are traditionally obtained from bidirectional encoder models. Despite their wide applicability, encoders have been recently overshadowed by advances in generative decoder-only models. However, many innovations driving this progress are not inherently tied to decoders. In this paper, we revisit the development of multilingual encoders through the lens of these advances, and introduce EuroBERT, a family of multilingual encoders covering European and widely spoken global languages. Our models outperform existing alternatives across a diverse range of tasks, spanning multilingual capabilities, mathematics, and coding, and natively supporting sequences of up to 8,192 tokens. We also examine the design decisions behind EuroBERT, offering insights into our dataset composition and training pipeline. We publicly release the EuroBERT models, including intermediate training checkpoints, together with our training framework.
