---
title: TransEvalnia
canonical_url: "https://www.modelscope.cn/datasets/SakanaAI/TransEvalnia"
md_url: "https://www.modelscope.cn/datasets/SakanaAI/TransEvalnia.md"
repository: SakanaAI/TransEvalnia
last_updated: 2025-07-19
license: "Apache License 2.0"
storage_size: "22 MB"
downloads: 276
stars: 2
---

# TransEvalnia

> TransEvalnia - SakanaAI 在 ModelScope 开源的数据集。TransEvalnia dataset

SakanaAI/TransEvalnia 是 ModelScope 魔搭社区上的数据集，存储大小 22 MB，采用 Apache License 2.0 许可。

- **Repository**: SakanaAI/TransEvalnia
- **License**: Apache License 2.0
- **Storage size**: 22 MB
- **Downloads**: 276
- **Stars**: 2
- **Last updated**: 2025-07-19

Source: https://www.modelscope.cn/datasets/SakanaAI/TransEvalnia

---

# TransEvalnia dataset

Paper: [arxiv](https://arxiv.org/abs/2507.12724) | Github: [SakanaAI/TransEvalnia](https://github.com/SakanaAI/TransEvalnia)

## Introduction

**[TransEvalnia](#)** is a prompting-based translation evaluation and ranking system that uses reasoning in performing its evaluations and ranking. This repo presents the dataset used in the work.

<img src="https://cdn-uploads.huggingface.co/production/uploads/605aebdece105fbcadcb8f3d/owp3vZDRul1406bXuEt8N.png" width="800">

The dataset consists of two parts. 
The **with_human_ranking** part includes 3,000 translation triplets with human scores. The data was mainly used for evaluating ranking accuracy. 
The **human_verification** part includes 800 model-generated evaluations and their verification from human annotators. The data was mainly used for meta-evaluation.

## The *with_human_ranking* data
The **with_human_ranking** data contains over 3,000 translation triplets (src, tgt1, tgt2) from the following 7 data sources and their reasoning-based evaluations from Qwen2.5-72B-Instruct and Claude Sonnet 3.5.
* `hard en-ja`: 47 English-Japanese translation triplets curated by expert translators
* `wmt 2021 en-ja`: 500 English-Japanese translation triplets from WMT 2021 DA
* `wmt 2021 ja-en`: 500 Japanese-English translation triplets from WMT 2021 DA
* `wmt 2022 en-ru`: 497 English-Russian translation triplets from WMT 2022 MQM
* `wmt 2023 en-de`: 500 English-German translation triplets from WMT 2023 MQM
* `wmt 2023 zh-en`: 500 Chinese-English translation triplets from WMT 2023 MQM
* `wmt 2024 en-es`: 499 English-Spanish translation triplets from WMT 2024 MQM

Every data item has the following fields:

- `dataset: str` - Dataset name.
- `src_text: str` - Source text.
- `tgt_texts: list[str]` - Texts of two translations.
- `src_lang: str` - Source language code (e.g. ja). 
- `tgt_lang: str` - Target language code (e.g. en).
- `src_lang_long: str` - Source language full name (e.g. Japanese).
- `src_lang_long: str` - Source language full name (e.g. English).
- `human_scores: list[float]` - Ground-truth human scores of the two translations.
  - For `hard en-ja`, a human score is between [0, 10]. A higher score indicates higher translation quality.
  - For datasets sourced from WMT DA, a human score is the z-score of the original direct assessment rating. A higher score indicates higher translation quality.
  - For datasets sourced from WMT MQM, a human score is the negated MQM score. A higher score indicates higher translation quality.
- `one_step_ranking/{model}: str` - Model generated ranking decision using the one-step method.
- `dim_evals/{model}: list[str]` - Model generated dimensional evaluations for each of the two translations.
- `two_step_ranking/{model}: str` - Model generated ranking decision based on `dim_evals/{model}`, using the two-step method.
- `two_step_scoring/{model}: str` - Model generated scoring decision based on `dim_evals/{model}`, using the two-step method.
- `interleaved_dim_evals: str` - Model generated interleaved dimensional evaluations, based on `dim_evals/{model}`, using the three-step method.
- `three_step_ranking/{model}: str` - Model generated ranking decision based on `interleaved_dim_evals/{model}`, using the three-step method.

The dependency between the fields can be visualized as following.
```
src_text
tgt_texts  
src_lang
tgt_lang
   ├── one_step_ranking/{model}
   └── dim_evals/{model}
         ├── two_step_ranking/{model}  
         ├── two_step_scoring/{model}
         └── interleaved_dim_evals/{model}
                 └── three_step_ranking/{model}
```

Note: `{model}` can be `qwen` or `claude`.

## The *human_verification* data
The **human_verification** data contains 800 model-generated evaluations and their human verifications. The verifications were collected from two translation service vendors.
* `human_verification-generic_claude-vendor1`: Claude Sonnet 3.5's evaluations of 200 translations from multiple domains. Annotated by vendor 1.
* `human_verification-generic_claude-vendor2`: Claude Sonnet 3.5's evaluations of 200 translations from multiple domains. Annotated by vendor 2.
* `human_verification-generic_qwen-vendor2`: Qwen2.5-72B-Instruct's evaluations of 200 translations from multiple domains. Annotated by vendor 2.
* `human_verification-haiku_qwen-vendor2`: Qwen2.5-72B-Instruct's evaluations of 200 Haiku translations. Annotated by vendor 2.

Every data item has the following fields:

- `idx: int` - Data index.
- `source: str` - Source text.
- `translation: str` - Translation text.
- `evaluation: str` - Model-generated evaluation.
- `annotations: dict` - Human verification of the model-generated evaluation. Data provided by different vendors have different structures.
- `system: str` - Model that was used to generate the translation.

## Citation

```
@misc{sproat2025transevalniareasoningbasedevaluationranking,
      title={TransEvalnia: Reasoning-based Evaluation and Ranking of Translations}, 
      author={Richard Sproat and Tianyu Zhao and Llion Jones},
      year={2025},
      eprint={2507.12724},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2507.12724}, 
}
```
