---
title: "WildIFEval: Instruction Following in the Wild"
canonical_url: "https://www.modelscope.cn/papers/125130"
md_url: "https://www.modelscope.cn/papers/125130.md"
arxiv_id: 2503.06573
published: 2025-03-09
last_updated: 2025-03-09
authors:
  - "Gili Lior"
  - "Asaf Yehudai"
  - "Ariel Gera"
  - "Liat Ein-Dor"
model_name: WILDIFEVAL
model_developer: "希伯来大学耶路撒冷分校，IBM研究院"
domain:
  - "自然语言处理"
  - "深度学习"
  - "机器学习"
type:
  - "自然语言处理"
  - "深度学习"
  - "机器学习"
  - "Computation and Language (cs.CL)"
  - "Artificial Intelligence (cs.AI)"
arxiv_url: "https://arxiv.org/abs/2503.06573"
pdf_url: "https://arxiv.org/pdf/2503.06573.pdf"
---

# WildIFEval: Instruction Following in the Wild

> Recent LLMs have shown remarkable success in following user instructions, yet handling instructions with multiple constraints remains a significant challenge. In this work, we introduce WildIFEval - a large-scale dataset of 12K real user instructions with…

「WildIFEval: Instruction Following in the Wild」是 ModelScope 魔搭社区收录的论文，arXiv 2503.06573，作者为 Gili Lior, Asaf Yehudai, Ariel Gera et al.，发表于 2025-03-09，属于 自然语言处理、深度学习、机器学习 领域。

- **ArXiv**: 2503.06573
- **Published**: 2025-03-09
- **Authors**: Gili Lior, Asaf Yehudai, Ariel Gera, Liat Ein-Dor
- **Model**: WILDIFEVAL
- **Developer**: 希伯来大学耶路撒冷分校，IBM研究院
- **Domain**: 自然语言处理, 深度学习, 机器学习
- **ArXiv URL**: https://arxiv.org/abs/2503.06573
- **PDF**: https://arxiv.org/pdf/2503.06573.pdf

Source: https://www.modelscope.cn/papers/125130

---

> WILDIFEVAL：探索真实世界多约束指令下的语言模型表现

## 摘要

本文提出了一种名为WILDIFEVAL的大规模数据集，旨在评估大型语言模型（LLMs）在处理多约束指令时的表现。现有数据集通常采用自底向上的方法，通过人工或合成生成的指令来评估模型的指令遵循能力，但这些方法可能无法捕捉到真实用户指令的复杂性和多样性。WILDIFEVAL包含12,000条来自真实用户的多约束指令，涵盖了广泛的词汇和主题范围，并将这些约束分为八个高层次类别。通过该数据集，作者对14个不同的LLMs进行了基准测试，发现所有模型在处理更多约束时性能都会下降，特别是在涉及长度相关约束的任务中。此外，不同类型约束对模型性能的影响也有所不同。这项工作不仅为复杂指令遵循任务提供了一个具有挑战性的基准，还揭示了现实世界中多约束指令的特点和分布。

## Abstract

Recent LLMs have shown remarkable success in following user instructions, yet handling instructions with multiple constraints remains a significant challenge. In this work, we introduce WildIFEval - a large-scale dataset of 12K real user instructions with diverse, multi-constraint conditions. Unlike prior datasets, our collection spans a broad lexical and topical spectrum of constraints, in natural user prompts. We categorize these constraints into eight high-level classes to capture their distribution and dynamics in real-world scenarios. Leveraging WildIFEval, we conduct extensive experiments to benchmark the instruction-following capabilities of leading LLMs. Our findings reveal that all evaluated models experience performance degradation with an increasing number of constraints. Thus, we show that all models have a large room for improvement on such tasks. Moreover, we observe that the specific type of constraint plays a critical role in model performance. We release our dataset to promote further research on instruction-following under complex, realistic conditions.
