---
title: bel_canto
canonical_url: "https://www.modelscope.cn/models/ccmusic-database/bel_canto"
md_url: "https://www.modelscope.cn/models/ccmusic-database/bel_canto.md"
repository: ccmusic-database/bel_canto
chinese_name: "美声民族唱法分类模型"
last_updated: 2026-08-05
license: "MIT License"
downloads: 1437
stars: 15
---

# bel_canto

> bel_canto - ccmusic-database 在 ModelScope 开源的模型。美声民族唱法分类模型旨在区分古典和民族声乐风格，所有音频样本均由专业歌手演唱。该模型使用一个包含四个类别的音频数据集进行微调，该数据集已经被转换为频谱图。骨干网络最初在计算机视觉（CV）领域进行预训练，随后经过专门为声乐风格分类任务设计的微调过程。在这个模型中，CV任务上的预训练为网络学习通用音频特征提供了基础，然后在微调过程中这些特征被调整以适应古典和民族声乐风格的微妙差异。这个包含四个类别的音频数据集包括来自古典和…

ccmusic-database/bel_canto 是 ModelScope 魔搭社区上的机器学习模型，采用 MIT License 许可。

- **Repository**: ccmusic-database/bel_canto
- **License**: MIT License
- **Downloads**: 1437
- **Stars**: 15
- **Last updated**: 2026-08-05

Source: https://www.modelscope.cn/models/ccmusic-database/bel_canto

---

# Intro 简介
The classification model for Bel Canto and Folk singing styles aims to distinguish between classical and folk vocal styles, with all audio samples performed by professional singers. The model is fine-tuned using an audio dataset comprising four categories, which have been converted into spectrograms. The backbone network was initially pre-trained in the field of computer vision (CV) and subsequently fine-tuned specifically for the task of vocal style classification. In this model, pre-training on CV tasks provides the network with a foundation for learning general audio features, which are then adjusted during the fine-tuning process to capture the subtle differences between classical and folk vocal styles. This audio dataset, which includes samples from classical and various folk singing traditions, enables the model to identify unique patterns associated with each vocal style. By using spectrograms as input representations, the model effectively analyzes the temporal and spectral components of audio signals. Through the fine-tuning process, the model continually enhances its ability to discern the nuanced differences in vocal delivery and style between classical and folk genres. This specialized model holds significant potential in the music industry and cultural preservation, as it accurately classifies vocal performances into these two broad categories. Its foundation in pre-trained computer vision principles demonstrates the versatility and adaptability of neural networks across different domains, enhancing the model's ability to capture the complex characteristics of vocal performances.

美声民族唱法分类模型旨在区分古典和民族声乐风格, 所有音频样本均由专业歌手演唱。该模型使用一个包含四个类别的音频数据集进行微调, 该数据集已经被转换为频谱图。骨干网络最初在计算机视觉 (CV) 领域进行预训练, 随后经过专门为声乐风格分类任务设计的微调过程。在这个模型中, CV任务上的预训练为网络学习通用音频特征提供了基础, 然后在微调过程中这些特征被调整以适应古典和民族声乐风格的微妙差异。这个包含四个类别的音频数据集包括来自古典和各种民族歌唱传统的样本, 使模型能够捕捉与每种声乐风格相关的独特模式。将频谱图作为输入表示使模型能够有效地分析音频信号的时域和频域成分。通过微调过程, 模型不断提升其辨别古典和民族风格之间声音传递和风格微妙差异的能力。这一专业模型在音乐产业和文化保护方面具有巨大潜力, 因为它能够准确地将声乐表演分类为这两个广泛的类别。其基于预训练计算机视觉原理的基础展示了神经网络在不同领域的多功能性和适应性, 增强了模型捕捉声乐表演复杂特征的能力。

## Demo 在线演示 (推理代码)
<https://www.modelscope.cn/studios/ccmusic-database/bel_canto>

## Usage 使用
:modelscope-code[]{type="sdk"}

## Maintenance 维护
```bash
GIT_LFS_SKIP_SMUDGE=1 
```

:modelscope-code[]{type="git"}

## Results 训练结果
|   Backbone    |                 Mel                  |     CQT     |   Chroma    |
| :-----------: | :----------------------------------: | :---------: | :---------: |
|    Swin-S     |             **_0.928_**              | **_0.936_** | **_0.787_** |
|    Swin-T     |                0.906                 |    0.863    |    0.731    |
|               |                                      |             |             |
|    AlexNet    |                0.919                 |    0.920    |    0.746    |
|  ConvNeXt-T   |                0.895                 |    0.925    |    0.714    |
|   GoogleNet   | [**_0.948_**](#best-result-最佳结果) |    0.921    |    0.739    |
|  MNASNet1.3   |                0.931                 | **_0.931_** | **_0.765_** |
| SqueezeNet1.1 |                0.923                 |    0.914    |    0.685    |
|    Average    |                0.921                 |    0.916    |    0.738    |

### Best Result 最佳结果
<table>
    <tr>
        <th>Loss curve</th>
        <td><img src="./googlenet_mel_2024-07-30_00-51-26/loss.jpg"></td>
    </tr>
    <tr>
        <th>Training and validation accuracy</th>
        <td><img src="./googlenet_mel_2024-07-30_00-51-26/acc.jpg"></td>
    </tr>
    <tr>
        <th>Confusion matrix</th>
        <td><img src="./googlenet_mel_2024-07-30_00-51-26/mat.jpg"></td>
    </tr>
</table>

## Dataset 数据集
<https://www.modelscope.cn/datasets/ccmusic-database/bel_canto>

## Mirror 镜像
<https://huggingface.co/ccmusic-database/bel_canto>

## Evaluation 校验
<https://github.com/monetjoe/ccmusic_eval>

## Cite 引用
```bibtex
@article{Zhou-2025,
  author  = {Monan Zhou and Shenyang Xu and Zhaorui Liu and Zhaowen Wang and Feng Yu and Wei Li and Baoqiang Han},
  title   = {CCMusic: An Open and Diverse Database for Chinese Music Information Retrieval Research},
  journal = {Transactions of the International Society for Music Information Retrieval},
  volume  = {8},
  number  = {1},
  pages   = {22--38},
  month   = {Mar},
  year    = {2025},
  url     = {https://doi.org/10.5334/tismir.194},
  doi     = {10.5334/tismir.194}
}
```
