---
title: "S2S-Arena, Evaluating Speech2Speech Protocols on Instruction Following with Paralinguistic Information"
canonical_url: "https://www.modelscope.cn/papers/124126"
md_url: "https://www.modelscope.cn/papers/124126.md"
arxiv_id: 2503.05085
published: 2025-03-07
last_updated: 2025-03-07
authors:
  - "Feng Jiang"
  - "Zhiyu Lin"
  - "Fan Bu"
  - "Yuhao Du"
  - "Benyou Wang"
  - "Haizhou Li"
model_name: S2S-Arena
model_developer: "香港中文大学（深圳）"
domain:
  - "自然语言处理"
  - "计算机语音技术"
  - "深度学习"
type:
  - "自然语言处理"
  - "计算机语音技术"
  - "深度学习"
  - "Computation and Language (cs.CL)"
  - "Sound (cs.SD)"
  - "Audio and Speech Processing (eess.AS)"
arxiv_url: "https://arxiv.org/abs/2503.05085"
pdf_url: "https://arxiv.org/pdf/2503.05085.pdf"
---

# S2S-Arena, Evaluating Speech2Speech Protocols on Instruction Following with Paralinguistic Information

> The rapid development of large language models (LLMs) has brought significant attention to speech models, particularly recent progress in speech2speech protocols supporting speech input and output. However, the existing benchmarks adopt automatic text-based…

「S2S-Arena, Evaluating Speech2Speech Protocols on Instruction Following with Paralinguistic Information」是 ModelScope 魔搭社区收录的论文，arXiv 2503.05085，作者为 Feng Jiang, Zhiyu Lin, Fan Bu et al.，发表于 2025-03-07，属于 自然语言处理、计算机语音技术、深度学习 领域。

- **ArXiv**: 2503.05085
- **Published**: 2025-03-07
- **Authors**: Feng Jiang, Zhiyu Lin, Fan Bu, Yuhao Du, Benyou Wang, Haizhou Li
- **Model**: S2S-Arena
- **Developer**: 香港中文大学（深圳）
- **Domain**: 自然语言处理, 计算机语音技术, 深度学习
- **ArXiv URL**: https://arxiv.org/abs/2503.05085
- **PDF**: https://arxiv.org/pdf/2503.05085.pdf

Source: https://www.modelscope.cn/papers/124126

---

> S2S-Arena：引入副语言信息的语音到语音指令跟随能力评估新标杆

## 摘要

本文提出了S2S-Arena，一个用于评估语音模型在语音到语音协议中指令跟随能力的新型基准。该基准特别关注包含副语言信息（如情感、语速、语调等）的语音理解和生成。现有基准主要依赖自动文本评估器，忽视了副语言信息的重要性。为了解决这一问题，S2S-Arena设计了154个融合TTS和真实录音的样本，涵盖四个实际领域的21项任务，并通过人工对四种流行语音模型进行了全面比较。

研究背景方面，随着大规模语言模型（LLMs）的发展，语音模型的研究也取得了显著进展。然而，现有的评估基准未能充分考虑副语言信息的影响，导致评估结果不够全面。为此，本文提出了一种新的评估框架，旨在更准确地反映语音模型的真实性能。

在提出的方法中，S2S-Arena采用三阶段构建过程：任务定义、指令设计和样本录制。它不仅要求模型理解语音输入中的副语言线索（如节奏），还要根据语义指令生成带有副语言特征的语音输出。具体来说，该基准包括四个难度级别，从仅考虑指令跟随到同时考虑输入和输出中的副语言信息。

实验结果显示，在语音到语音协议中，级联式ASR-LLM-TTS模型在对齐文本和语音后优于联合训练模型。此外，副语言信息的理解主要依赖于LLM骨干，而多语言支持则受限于语音模块。虽然现有优秀模型已经能够理解语音输入中的副语言信息，但在生成带副语言信息的适当音频方面仍面临挑战。这项工作为未来语音模型的设计提供了重要参考，特别是在多模态和多语言支持方面。

## Abstract

The rapid development of large language models (LLMs) has brought significant attention to speech models, particularly recent progress in speech2speech protocols supporting speech input and output. However, the existing benchmarks adopt automatic text-based evaluators for evaluating the instruction following ability of these models lack consideration for paralinguistic information in both speech understanding and generation. To address these issues, we introduce S2S-Arena, a novel arena-style S2S benchmark that evaluates instruction-following capabilities with paralinguistic information in both speech-in and speech-out across real-world tasks. We design 154 samples that fused TTS and live recordings in four domains with 21 tasks and manually evaluate existing popular speech models in an arena-style manner. The experimental results show that: (1) in addition to the superior performance of GPT-4o, the speech model of cascaded ASR, LLM, and TTS outperforms the jointly trained model after text-speech alignment in speech2speech protocols; (2) considering paralinguistic information, the knowledgeability of the speech model mainly depends on the LLM backbone, and the multilingual support of that is limited by the speech module; (3) excellent speech models can already understand the paralinguistic information in speech input, but generating appropriate audio with paralinguistic information is still a challenge.
