---
title: wnut17_ner
canonical_url: "https://www.modelscope.cn/datasets/iic/wnut17_ner"
md_url: "https://www.modelscope.cn/datasets/iic/wnut17_ner.md"
repository: iic/wnut17_ner
chinese_name: "wnut17命名实体识别数据集"
last_updated: 2022-10-18
license: CC-BY-4.0
storage_size: "6.5 MB"
domain:
  - text
tasks:
  - token-classification
language:
  - en
type:
  - ner
downloads: 8898
stars: 4
---

# wnut17_ner

> wnut17_ner - iic 在 ModelScope 开源的数据集。数据集概述 wnut17数据集是面向社交媒体的英文命名实体识别数据集。

iic/wnut17_ner 是 ModelScope 魔搭社区上的token-classification数据集，涉及 text 领域，存储大小 6.5 MB，采用 CC-BY-4.0 许可。

- **Repository**: iic/wnut17_ner
- **License**: CC-BY-4.0
- **Tasks**: token-classification
- **Domain**: text
- **Storage size**: 6.5 MB
- **Downloads**: 8898
- **Stars**: 4
- **Last updated**: 2022-10-18

Source: https://www.modelscope.cn/datasets/iic/wnut17_ner

---

# wnut17命名实体识别数据集

## 数据集概述
wnut17数据集是面向社交媒体的英文命名实体识别数据集。

### 数据集简介
本数据集包括训练集（3394）、验证集（1009）、测试集（1287），实体类型包括corporation, creative-work、group、location、person、product。

### 数据集的格式和结构
数据格式采用conll标准，数据分为两列，第一列是输入句中的词划分，第二列是每个词对应的命名实体类型标签。一个具体case的例子如下：

```
Visuals	O
of	O
the	O
avalanche	O
site	O
in	O
Gurez	B-location
sector	I-location
.	O
```

## 数据集版权信息

Creative Commons Attribution 4.0 International。

## 引用方式
  ```bib
    @inproceedings{derczynski-etal-2017-results,
        title = "Results of the {WNUT}2017 Shared Task on Novel and Emerging Entity Recognition",
      author = "Derczynski, Leon  and
        Nichols, Eric  and
        van Erp, Marieke  and
        Limsopatham, Nut",
      booktitle = "Proceedings of the 3rd Workshop on Noisy User-generated Text",
      month = sep,
      year = "2017",  
      address = "Copenhagen, Denmark",
      publisher = "Association for Computational Linguistics",
      url = "https://www.aclweb.org/anthology/W17-4418",
      doi = "10.18653/v1/W17-4418",
      pages = "140--147",
      abstract = "This shared task focuses on identifying unusual, previously-unseen entities in the context of emerging discussions.
                  Named entities form the basis of many modern approaches to other tasks (like event clustering and summarization),
                  but recall on them is a real problem in noisy text - even among annotators.
                  This drop tends to be due to novel entities and surface forms.
                  Take for example the tweet {``}so.. kktny in 30 mins?!{''} {--} even human experts find the entity {`}kktny{'}
                  hard to detect and resolve. The goal of this task is to provide a definition of emerging and of rare entities,
                  and based on that, also datasets for detecting these entities. The task as described in this paper evaluated the
                  ability of participating entries to detect and classify novel and emerging named entities in noisy text.",
  }
  ```
