---
title: bel_canto
canonical_url: "https://www.modelscope.cn/datasets/ccmusic-database/bel_canto"
md_url: "https://www.modelscope.cn/datasets/ccmusic-database/bel_canto.md"
repository: ccmusic-database/bel_canto
chinese_name: "美声民族唱法数据集 Bel Conto & Chinese Folk Song Singing Technique Dataset"
last_updated: 2026-08-05
license: CC-BY-NC-ND
storage_size: "3.1 GB"
domain:
  - audio
tasks:
  - audio-classification
downloads: 8614
stars: 15
---

# bel_canto

> bel_canto - ccmusic-database 在 ModelScope 开源的数据集。美声民族唱法数据集专注于区分美声与民族唱法，所有音频均由专业歌唱家演唱。该数据集包含四个明确定义的类别，分别为男美声、男民族、女美声和女民族。经过处理的数据是通过对演唱录音进行去静音、合并、频谱转换以及切片而获得，总计包含9603个数据样本。该数据集的建立旨在为研究者提供一个具有代表性的音频数据集，以便支持对美声和民族唱法声学特征的深入研究，并为相关领域的深度学习模型开发提供基础。

ccmusic-database/bel_canto 是 ModelScope 魔搭社区上的audio-classification数据集，涉及 audio 领域，存储大小 3.1 GB，采用 CC-BY-NC-ND 许可。

- **Repository**: ccmusic-database/bel_canto
- **License**: CC-BY-NC-ND
- **Tasks**: audio-classification
- **Domain**: audio
- **Storage size**: 3.1 GB
- **Downloads**: 8614
- **Stars**: 15
- **Last updated**: 2026-08-05

Source: https://www.modelscope.cn/datasets/ccmusic-database/bel_canto

---

# Dataset Card for Bel Conto and Chinese Folk Song Singing Tech 简介
## Original Content 原始内容
This dataset is created by the authors and encompasses two distinct singing styles: bel canto and Chinese folk singing. Bel canto is a vocal technique frequently employed in Western classical music and opera, symbolizing the zenith of vocal artistry within the broader Western musical heritage. Chinese folk singing, for which there is no official English translation, is referred to here as a classical singing style that originated in China during the 20th century. It is a fusion of traditional Chinese vocal techniques with European bel canto singing and is currently widely utilized in the performance of various forms of Chinese folk music. The purpose of creating this dataset is to fill a gap in the current singing datasets, as none of them includes Chinese folk singing, and by incorporating both bel canto and Chinese folk singing, it provides a valuable resource for researchers to conduct cross-cultural comparative studies in vocal performance. The original dataset contains 203 acapella singing recordings sung in two styles, bel canto and Chinese folk singing style. All of them were sung by professional vocalists and were recorded in the recording studio of the China Conservatory of Music using a Schoeps MK4 microphone. Additionally, apart from singing style labels, gender labels are also included.

原始数据集来源于 [美声民族唱法数据集](https://ccmusic-database.github.io/database/ccm.html#shou9), 包含 203 段不同风格的无伴奏演唱片段, 这些片段由专业声乐家以美声和中国民族唱法风格演唱。所有演唱均由专业声乐家在专业的商业录音室录制。

## Integration 集成
Since this is a self-created dataset, we directly carry out the unified integration of the data structure. After integration, the data structure of the dataset is as follows: audio (with a sampling rate of 22,050 Hz), mel spectrograms, 4-class numerical label, gender label and singing style label. The data number remains 203, with a total duration of 5.08 hours. The average duration is 90 seconds.

We have constructed the [default subset](#usage-快速使用) of the current integrated version of the dataset, and its data structure can be viewed in the [viewer](https://www.modelscope.cn/datasets/ccmusic-database/bel_canto/dataPeview). Since the default subset has not been evaluated, to verify its effectiveness, we have built the [eval subset](#usage-快速使用) based on the default subset for the evaluation of the integrated version of the dataset. The evaluation results can be seen in [bel_canto](https://www.modelscope.cn/models/ccmusic-database/bel_canto).

基于上述原始数据集, 我们构建了当前集成版数据集的 [默认子集](#usage-快速使用), 其数据结构可以在 [数据预览](https://www.modelscope.cn/datasets/ccmusic-database/bel_canto/dataPeview) 中查看。由于默认子集并未进行评估, 为了验证其有效性, 我们基于默认子集构建了 [评估子集](#usage-快速使用), 用于对集成版数据集的评估, 评估结果可以在 [美声民族唱法分类模型](https://www.modelscope.cn/models/ccmusic-database/bel_canto) 中查看。

## Statistics 统计
| ![](https://www.modelscope.cn/datasets/ccmusic-database/bel_canto/resolve/master/data/bel_pie.jpg) | ![](https://www.modelscope.cn/datasets/ccmusic-database/bel_canto/resolve/master/data/bel_canto.jpg) | ![](https://www.modelscope.cn/datasets/ccmusic-database/bel_canto/resolve/master/data/bel_bar.jpg) |
| :------------------------------------------------------------------------------------------------: | :--------------------------------------------------------------------------------------------------: | :------------------------------------------------------------------------------------------------: |
|                                             **Fig. 1**                                             |                                              **Fig. 2**                                              |                                             **Fig. 3**                                             |

Firstly, **Fig. 1** presents the clip number of each category. The label with the highest data volume is Folk singing female, comprising 93 audio clips, which is 45.8% of the dataset. The label with the lowest data volume is Bel Canto Male, with 32 audio clips, constituting 15.8% of the dataset. Next, we assess the length of the audio for each label in **Fig. 2**. The Folk Singing Female label has the longest audio data, totaling 119.85 minutes. Bel Canto Male has the shortest audio duration, amounting to 46.61 minutes. The trend is the same as the data number difference shown in the pie chart. Lastly, in **Fig. 3**, the number of audio clips within various duration intervals is displayed. The most common duration range is observed to be 46-79 seconds, with 78 audio segments, while the least common range is 211-244 seconds, featuring only 2 audio segments.

| Split(8:1:1) / Subset | default |  eval |
| :-------------------: | ------: | ----: |
|         train         |     159 | 7,926 |
|      validation       |      21 |   990 |
|         test          |      23 |   994 |
|         total         |     203 | 9,910 |

Since the raw audio data contained missing labels for singing method and gender, we manually completed these missing labels during the data preprocessing stage with the assistance of individuals with musical expertise. In addition, we extracted the mel spectrograms from the original audio and consolidated the original audio, mel spectrograms, 2-class singing method labels, 2-class gender labels, and the 4-class labels formed by the orthogonal combination of the two labeling systems into a single data table with viewer.

|                 Statistical items 统计项                  |      Values 值       |
| :-------------------------------------------------------: | :------------------: |
|                   Total count 总数据量                    |        `203`         |
|               Total duration(s) 总时长(秒)                | `18270.477865079374` |
|               Mean duration(s) 平均时长(秒)               | `90.00235401516929`  |
|               Min duration(s) 最短时长(秒)                |        `13.7`        |
|               Max duration(s) 最长时长(秒)                |       `310.0`        |
| Class in the longest audio duartion interval 最长时长类别 |  `Bel Canto Female`  |

由于原始数据中的音频存在 singing method 和 gender 标签缺失的情况, 在数据整理阶段, 我们邀请了具有音乐背景的人员对这些缺失的标签进行了人工补全。此外, 我们从原始音频中提取了其 mel 频谱, 并将原始音频、mel 频谱、singing method 的二分类标签、gender 的二分类标签, 以及由上述两个标签体系正交生成的四分类标签整合至同一可预览的数据表中。

## Default Subset Structure 默认子数据集结构
### Data Instances 文件格式
.zip(.wav, .jpg)

### Data Fields 标签
m_bel=0, f_bel=1, m_folk=2, f_folk=3
 
## Usage 快速使用
:modelscope-code[]{type="sdk"}

## Maintenance 维护
```bash
GIT_LFS_SKIP_SMUDGE=1 
```

:modelscope-code[]{type="git"}

### Requirements(For data processing codes) 环境(仅用于数据处理代码)
```bash
cd data
conda create -n data python=3.x -y
conda activate data
pip install -r https://www.modelscope.cn/datasets/ccmusic-database/bel_canto/resolve/master/data/requirements.txt
```

### Data processor 数据处理脚本
1. Open project with `VSCode`
2. Select file `https://www.modelscope.cn/datasets/ccmusic-database/bel_canto/resolve/master/data/data.py`
3. Press `F5`

### Dataset Summary 数据集总结
For the second version, the audio first undergoes a **_silence-removal_** stage, in which a high-pass filter with a threshold of **_40dB_** is applied to filter out all silent regions. Then, silence-removed audio is concatenated into one single audio stream. Then, the audio stream is segmented to a 1.6-second clip and is then transformed to Mel, CQT and Chroma. This yields a total size of 9,910 instances.

### Supported Tasks and Leaderboards 支持任务
Audio classification, Image classification, singing method classification, voice classification

### Languages 语言
Chinese, English

## Dataset Creation 数据创建
### Curation Rationale 动机
Lack of a dataset for Bel Conto and Chinese folk song-singing tech

### Source Data 数据源
#### Initial Data Collection and Normalization 源数据搜集与正规化
Zhaorui Liu, Monan Zhou

#### Who are the source language producers? 语言支持
Students from CCMUSIC

### Annotations 标注
#### Annotation process 标注步骤
All of them are sung by professional vocalists and were recorded in professional commercial recording studios.

#### Who are the annotators? 标注者
professional vocalists

## Considerations for Using the Data 数据使用考量
### Social Impact of Dataset 社会影响
Promoting the development of AI in the music industry

### Discussion of Biases 偏好
Only for Chinese songs

### Other Known Limitations 其它限制
Some singers may not have enough professional training in classical or ethnic vocal techniques.

## Additional Information 附加信息
### Dataset Curators 策划人
Zijin Li

### Mirror 镜像
<https://huggingface.co/datasets/ccmusic-database/bel_canto>

### Evaluation 评估
<https://www.modelscope.cn/models/ccmusic-database/bel_canto>

### Citation Information 引用
```bibtex
@article{Zhou-2025,
  author  = {Monan Zhou and Shenyang Xu and Zhaorui Liu and Zhaowen Wang and Feng Yu and Wei Li and Baoqiang Han},
  title   = {CCMusic: An Open and Diverse Database for Chinese Music Information Retrieval Research},
  journal = {Transactions of the International Society for Music Information Retrieval},
  volume  = {8},
  number  = {1},
  pages   = {22--38},
  month   = {Mar},
  year    = {2025},
  url     = {https://doi.org/10.5334/tismir.194},
  doi     = {10.5334/tismir.194}
}
```

### Contributions 贡献
Provide a dataset for distinguishing Bel Conto and Chinese folk song-singing tech

<div style="display:none">
