---
title: erhu_playing_tech
canonical_url: "https://www.modelscope.cn/datasets/ccmusic-database/erhu_playing_tech"
md_url: "https://www.modelscope.cn/datasets/ccmusic-database/erhu_playing_tech.md"
repository: ccmusic-database/erhu_playing_tech
chinese_name: "二胡演奏技法数据集 Erhu Playing Technique Dataset"
last_updated: 2026-08-05
license: CC-BY-NC-ND
storage_size: "341 MB"
domain:
  - audio
tasks:
  - audio-classification
downloads: 6499
stars: 17
---

# erhu_playing_tech

> erhu_playing_tech - ccmusic-database 在 ModelScope 开源的数据集。本数据集包含1500个二胡音频片段（.wav格式），所有音频均由专业二胡演奏家演奏。根据二胡不同的演奏技法，将它们分为11类（分弓、垫弓、泛音、连弓&滑音&大滑音、击弓、拨弦、抛弓、顿弓、颤弓、颤音、揉弦）。每种演奏技法都有对应的若干个音频。音频来自： 中国传统乐器音响数据库（CTIS）。

ccmusic-database/erhu_playing_tech 是 ModelScope 魔搭社区上的audio-classification数据集，涉及 audio 领域，存储大小 341 MB，采用 CC-BY-NC-ND 许可。

- **Repository**: ccmusic-database/erhu_playing_tech
- **License**: CC-BY-NC-ND
- **Tasks**: audio-classification
- **Domain**: audio
- **Storage size**: 341 MB
- **Downloads**: 6499
- **Stars**: 17
- **Last updated**: 2026-08-05

Source: https://www.modelscope.cn/datasets/ccmusic-database/erhu_playing_tech

---

# Intro 简介
## Original Content 原始内容
This dataset was created and has been utilized for Erhu playing technique detection by [[1]](https://arxiv.org/pdf/1910.09021), which has not undergone peer review. The original dataset comprises 1,253 Erhu audio clips, all performed by professional Erhu players. These clips were annotated according to three hierarchical levels, resulting in annotations for four, seven, and 11 categories. Part of the audio data is sourced from the CTIS dataset described earlier.

原始数据集来源于 [二胡演奏技法数据集](https://ccmusic-database.github.io/database/ccm.html#shou8), 所有演奏均由专业二胡演奏家完成。这些片段由具有二胡演奏专长的注释者分为11类, 分别是: 分弓、垫弓、泛音、连弓与滑音与拨弦、击弓、拨弦、抛弓、断弓、颤弓和震音。对于某些演奏技巧, 提供了多个音频片段, 每个片段的力度都不同。该数据集已创建并用于二胡演奏技巧检测。标签系统是分层的, 包含三个级别。第一级别包括四个类别: 颤音、断音、滑音和其他; 第二级别包括七个类别: 颤音短、颤音长、断音、滑音上、滑音连音、滑音下和其他; 第三级别包括11个类别, 即前面描述的11种演奏技巧。

## Integration 集成
We first perform label cleaning to abandon the labels for the four and seven categories, since they do not strictly form a hierarchical relationship, and there are also missing data problems. This process leaves us with only the labels for the 11 categories. Then, we add Chinese character label and Chinese pinyin label to enhance comprehensibility. The 11 labels are: Detache (分弓), Diangong (垫弓), Harmonic (泛音), Legato\slide\glissando (连弓\滑音\连音), Percussive (击弓), Pizzicato (拨弦), Ricochet (抛弓), Staccato (断弓), Tremolo (震音), Trill (颤音), and Vibrato (揉弦). After integration, the data structure contains six columns: audio (with a sampling rate of 44,100 Hz), mel spectrograms, numeric label, Italian label, Chinese character label, and Chinese pinyin label. The total number of audio clips remains at 1,253, with a total duration of 25.81 minutes. The average duration is 1.24 seconds.

We constructed the [default subset](#usage-快速使用) of the current integrated version dataset based on its 11 classification data and optimized the names of the 11 categories. The data structure can be seen in the [viewer](https://www.modelscope.cn/datasets/ccmusic-database/erhu_playing_tech/dataPeview). Although the original dataset has been cited in some articles, the experiments in those articles lack reproducibility. In order to demonstrate the effectiveness of the default subset, we further processed the data and constructed the [eval subset](#usage-快速使用) to supplement the evaluation of this integrated version dataset. The results of the evaluation can be viewed in the [erhu_playing_tech](https://www.modelscope.cn/models/ccmusic-database/erhu_playing_tech). In addition, the labels of categories 4 and 7 in the original dataset were not discarded. Instead, they were separately constructed into [4_classes subset](#usage-快速使用) and [7_classes subset](#usage-快速使用). However, these two subsets have not been evaluated and therefore are not reflected in our paper.

经过对上述数据的整理, 我们基于其中的 11 分类数据构建了当前集成版本数据集的 [默认子集](#usage-快速使用), 并优化了 11 个类别的名称, 其数据结构见 [数据预览](https://www.modelscope.cn/datasets/ccmusic-database/erhu_playing_tech/dataPeview)。尽管有文章引用过原始数据集, 但该文章的实验缺乏可复现性, 为了能够呈现默认子集的有效性, 我们对其进行了进一步的数据处理, 构建出了 [评估子集](#usage-快速使用) 用于补充对本集成版本数据集的评估, 评估的结果可在 [二胡演奏技法识别模型](https://www.modelscope.cn/models/ccmusic-database/erhu_playing_tech) 中查看。除此之外, 原始数据集中的4类和7类标签也没有被丢弃, 而是分别构建成了 [4分类子集](#usage-快速使用) 和 [7分类子集](#usage-快速使用)。但这两个子集并未进行评估, 因此未在文章中体现。

## Statistics 统计
| ![](https://www.modelscope.cn/datasets/ccmusic-database/erhu_playing_tech/resolve/master/data/erhu_pie.jpg) | ![](https://www.modelscope.cn/datasets/ccmusic-database/erhu_playing_tech/resolve/master/data/erhu.jpg) | ![](https://www.modelscope.cn/datasets/ccmusic-database/erhu_playing_tech/resolve/master/data/erhu_bar.jpg) | ![](https://www.modelscope.cn/datasets/ccmusic-database/erhu_playing_tech/resolve/master/data/erhu.png) |
| :---------------------------------------------------------------------------------------------------------: | :-----------------------------------------------------------------------------------------------------: | :---------------------------------------------------------------------------------------------------------: | :-----------------------------------------------------------------------------------------------------: |
|                                                 **Fig. 1**                                                  |                                               **Fig. 2**                                                |                                                 **Fig. 3**                                                  |                                               **Fig. 4**                                                |

To begin with, **Fig. 1** presents the number of data entries per label. The Trill label has the highest data volume, with 249 instances, which accounts for 19.9% of the total dataset. Conversely, the Harmonic label has the least amount of data, with only 30 instances, representing a meager 2.4% of the total. Turning to the audio duration per category, as illustrated in **Fig. 2**, the audio data associated with the Trill label has the longest cumulative duration, amounting to 4.88 minutes. In contrast, the Percussive label has the shortest audio duration, clocking in at 0.75 minutes. These disparities clearly indicate a class imbalance problem within the dataset. Finally, as shown in **Fig. 3**, we count the frequency of audio occurrences at 550-ms intervals. The quantity of data decreases as the duration lengthens. The most populated duration range is 90-640 ms, with 422 audio clips. The least populated range is 3390-3940 ms, which contains only 12 clips. **Fig. 4** is the statistical charts for the 11_classes (Default), 7_classes, and 4_classes subsets. (图四是 11类(默认)、7类以及4类子集的统计图)

### Totals 总量统计
|         Subset 子集         | Total count 总数据量 | Total duration(s) 总时长(秒) |
| :-------------------------: | :------------------: | :--------------------------: |
| Default / 11_classes / Eval |        `1253`        |     `1548.3557823129247`     |
|    7_classes / 4_classes    |        `635`         |     `719.8175736961448`      |

### Range (Default subset) 默认子集极值
|                  Statistical items 统计项                   |      Values 值       |
| :---------------------------------------------------------: | :------------------: |
|              Mean duration(ms) 平均时长(毫秒)               | `1235.7189004891661` |
|               Min duration(ms) 最短时长(毫秒)               |  `91.7687074829932`  |
|               Max duration(ms) 最长时长(毫秒)               | `4468.934240362812`  |
| Classes in the longest audio duartion interval 最长时长类别 |  `Vibrato, Detache`  |

## Labels 标签
Compared to the raw data, we added the columns tech name, Chinese name, and Chinese pinyin during the data preprocessing stage. Additionally, we extracted the mel spectrograms from the original audio and integrated the original audio, mel spectrograms, label columns, and the aforementioned three columns into a single data table with viewer.

|      Playing tech      | Label No |  Chinese name  |          Chinese pinyin          |
| :--------------------: | :------: | :------------: | :------------------------------: |
|        vibrato         |    0     |      揉弦      |            rou2_xian2            |
|         trill          |    1     |      颤音      |            chan4_yin1            |
|        tremolo         |    2     |      震音      |            zhen4_yin1            |
|        staccato        |    3     |      断弓      |           duan4_gong1            |
|        ricochet        |    4     |      抛弓      |            pao1_gong1            |
|       pizzicato        |    5     |      拨弦      |            bo1_xian2             |
|       percussive       |    6     |      击弓      |            ji1_gong1             |
| legato_slide_glissando |    7     | 连弓/滑音/连音 | lian2_gong1/hua2_yin1/lian2_yin1 |
|        harmonic        |    8     |      泛音      |            fan4_yin1             |
|        diangong        |    9     |      垫弓      |           dian4_gong1            |
|        detache         |    10    |      分弓      |            fen1_gong1            |

相比于原始数据, 在数据整理阶段, 我们增加了*技法名*、*中文名*和*汉语拼音*列。此外, 我们从原始音频中提取了 mel 频谱, 并将原始音频、mel 频谱、标签列以及上述三列整合到同一可预览的数据表中。

## Usage 快速使用
:modelscope-code[]{type="sdk"}

## Clone with HTTP 下载
```bash
GIT_LFS_SKIP_SMUDGE=1 
```

:modelscope-code[]{type="git"}

## Mirror 镜像
<https://huggingface.co/datasets/ccmusic-database/erhu_playing_tech>

## Evaluation 评估
[1] [Wang, Zehao et al. “Musical Instrument Playing Technique Detection Based on FCN: Using Chinese Bowed-Stringed Instrument as an Example.” ArXiv abs/1910.09021 (2019): n. pag.](https://arxiv.org/pdf/1910.09021)<br>
[2] <https://www.modelscope.cn/models/ccmusic-database/erhu_playing_tech>

## Cite 引用
```bibtex
@article{abs-1910-09021,
  author     = {Zehao Wang and Jingru Li and Xiaoou Chen and Zijin Li and Shicheng Zhang and Baoqiang Han and Deshun Yang},
  title      = {Musical Instrument Playing Technique Detection Based on {FCN:} Using Chinese Bowed-Stringed Instrument as an Example},
  journal    = {CoRR},
  volume     = {abs/1910.09021},
  year       = {2019},
  url        = {http://arxiv.org/abs/1910.09021},
  eprinttype = {arXiv},
  eprint     = {1910.09021}
}
```

<div style="display:none">
