---
title: C3
canonical_url: "https://www.modelscope.cn/datasets/OmniData/C3"
md_url: "https://www.modelscope.cn/datasets/OmniData/C3.md"
repository: OmniData/C3
last_updated: 2024-07-01
license: "[C3 Custom]"
storage_size: "9.4 MB"
domain:
  - publishDate
  - displayName
  - labelTypes
  - taskTypes
  - paperUrl
  - publishUrl
  - publisher
  - mediaTypes
tasks:
  - 2019-01-01
  - C3
  - "Chinese Corpus"
  - "Common Sense Reasoning One-Shot"
  - "https://arxiv.org/pdf/1904.09679v3.pdf"
  - "https://github.com/nlpdata/c3"
  - "Tencent AI Lab"
  - Text
downloads: 176
stars: 0
---

# C3

> C3 - OmniData 在 ModelScope 开源的数据集。displayName: C3 labelTypes: Chinese Corpus license: C3 Custom mediaTypes: Text paperUrl: https://arxiv.org/pdf/1904.09679v3.pdf publishDate: "2019-01-01" publishUrl: https://github.com/nlpdata/c3 publisher: Cornell…

OmniData/C3 是 ModelScope 魔搭社区上的2019-01-01、C3、Chinese Corpus数据集，涉及 publishDate、displayName、labelTypes 领域，存储大小 9.4 MB，采用 [C3 Custom] 许可。

- **Repository**: OmniData/C3
- **License**: [C3 Custom]
- **Tasks**: 2019-01-01, C3, Chinese Corpus, Common Sense Reasoning One-Shot, https://arxiv.org/pdf/1904.09679v3.pdf, https://github.com/nlpdata/c3, Tencent AI Lab, Text
- **Domain**: publishDate, displayName, labelTypes, taskTypes, paperUrl, publishUrl, publisher, mediaTypes
- **Storage size**: 9.4 MB
- **Downloads**: 176
- **Stars**: 0
- **Last updated**: 2024-07-01

Source: https://www.modelscope.cn/datasets/OmniData/C3

---

displayName: C3
labelTypes:
- Chinese Corpus
license:
- C3 Custom
mediaTypes:
- Text
paperUrl: https://arxiv.org/pdf/1904.09679v3.pdf
publishDate: "2019-01-01"
publishUrl: https://github.com/nlpdata/c3
publisher:
- Cornell University
- Tencent AI Lab
tags: []
taskTypes:
- Machine Reading Comprehension
- Reading Comprehension
- Language Modelling
- Common Sense Reasoning Few-Shot
- Common Sense Reasoning Zero-Shot
- Common Sense Reasoning One-Shot

---
  ## 简介
  C3 是一个自由形式的多选中文机器阅读理解数据集。我们展示了第一个自由形式的多选中文机器阅读理解数据集（C^3），包含 13,369 个文档（对话或更正式的混合体裁文本）及其相关的 19,577 个从中文收集的自由形式选择题-作为第二语言的考试。我们对这些现实世界问题所需的先验知识（即语言、特定领域和一般世界知识）进行了全面分析。我们实施了基于规则和流行的神经方法，发现性能最佳的模型 (68.5%) 和人类读者 (96.0%) 之间仍然存在显着的性能差距，尤其是在需要先验知识的问题上。我们进一步研究了基于英语翻译相关数据集的干扰物合理性和数据增强对模型性能的影响。我们预计 C^3 将对现有系统提出巨大挑战，因为回答 86.8% 的问题需要随附文档内外的知识，我们希望 C^3 可以作为研究如何利用各种先验知识的平台更好地理解给定的书面或口头文本。 C^3 可在 https://dataset.org/c3/ 获得。
  ## 引文
  ```
@article{sun2020investigating,
  title={Investigating prior knowledge for challenging chinese machine reading comprehension},
  author={Sun, Kai and Yu, Dian and Yu, Dong and Cardie, Claire},
  journal={Transactions of the Association for Computational Linguistics},
  volume={8},
  pages={141--155},
  year={2020},
  publisher={MIT Press}
}
```
  
## Download dataset
:modelscope-code[]{type="git"}
