---
title: "Adversarial Data Collection: Human-Collaborative Perturbations for Efficient and Robust Robotic Imitation Learning"
canonical_url: "https://www.modelscope.cn/papers/126948"
md_url: "https://www.modelscope.cn/papers/126948.md"
arxiv_id: 2503.11646
published: 2025-03-14
last_updated: 2025-03-14
authors:
  - "Siyuan Huang"
  - "Yue Liao"
  - "Siyuan Feng"
  - "Shu Jiang"
  - "Si Liu"
  - "Hongsheng Li"
  - "Maoqing Yao"
  - "Guanghui Ren"
model_name: "Adversarial Data Collection (ADC)"
model_developer: "上海交通大学, 香港中文大学多媒体实验室(MMLab), Agibot, 北京航空航天大学"
domain:
  - "机器人学"
  - "自然语言处理"
  - "计算机视觉"
  - "机器学习"
type:
  - "机器人学"
  - "自然语言处理"
  - "计算机视觉"
  - "机器学习"
  - "Robotics (cs.RO)"
arxiv_url: "https://arxiv.org/abs/2503.11646"
pdf_url: "https://arxiv.org/pdf/2503.11646.pdf"
---

# Adversarial Data Collection: Human-Collaborative Perturbations for Efficient and Robust Robotic Imitation Learning

> The pursuit of data efficiency, where quality outweighs quantity, has emerged as a cornerstone in robotic manipulation, especially given the high costs associated with real-world data collection. We propose that maximizing the informational density of…

「Adversarial Data Collection: Human-Collaborative Perturbations for Efficient and Robust Robotic Imitation Learning」是 ModelScope 魔搭社区收录的论文，arXiv 2503.11646，作者为 Siyuan Huang, Yue Liao, Siyuan Feng et al.，发表于 2025-03-14，属于 机器人学、自然语言处理、计算机视觉 领域。

- **ArXiv**: 2503.11646
- **Published**: 2025-03-14
- **Authors**: Siyuan Huang, Yue Liao, Siyuan Feng, Shu Jiang, Si Liu, Hongsheng Li, Maoqing Yao, Guanghui Ren
- **Model**: Adversarial Data Collection (ADC)
- **Developer**: 上海交通大学, 香港中文大学多媒体实验室(MMLab), Agibot, 北京航空航天大学
- **Domain**: 机器人学, 自然语言处理, 计算机视觉, 机器学习
- **ArXiv URL**: https://arxiv.org/abs/2503.11646
- **PDF**: https://arxiv.org/pdf/2503.11646.pdf

Source: https://www.modelscope.cn/papers/126948

---

> 对抗性数据收集：用实时人机协作扰动提升机器人模仿学习的效率与鲁棒性

## 摘要

本文针对机器人操作中数据收集效率低、成本高的问题，提出了一种名为Adversarial Data Collection (ADC)的方法。传统的数据收集方式通常依赖于大量静态演示数据，而这些数据往往包含冗余信息且缺乏多样性。ADC通过引入人机协作的对抗性扰动范式，重新定义了机器人数据采集过程。具体而言，在单个演示过程中，对抗性操作员动态改变物体状态、环境条件和语言指令，而远程操作员则需要适应这些变化并调整动作。这种方法将多样化的失败恢复行为、任务变体和环境扰动压缩到最小的演示中。

实验结果表明，基于ADC训练的模型在面对未见过的任务指令、感知扰动以及错误恢复能力方面表现出更强的泛化性和鲁棒性。特别地，使用仅20% ADC收集的数据量即可显著优于传统方法使用完整数据集的效果。这表明，战略性数据采集比单纯的后处理更能提升数据质量，从而实现更高效的机器人学习。

此外，作者正在构建一个大规模的ADC-Robotics数据集，该数据集包含真实世界中的操纵任务及其对抗性扰动，旨在推动机器人模仿学习的研究进展。

## Abstract

The pursuit of data efficiency, where quality outweighs quantity, has emerged as a cornerstone in robotic manipulation, especially given the high costs associated with real-world data collection. We propose that maximizing the informational density of individual demonstrations can dramatically reduce reliance on large-scale datasets while improving task performance. To this end, we introduce Adversarial Data Collection, a Human-in-the-Loop (HiL) framework that redefines robotic data acquisition through real-time, bidirectional human-environment interactions. Unlike conventional pipelines that passively record static demonstrations, ADC adopts a collaborative perturbation paradigm: during a single episode, an adversarial operator dynamically alters object states, environmental conditions, and linguistic commands, while the tele-operator adaptively adjusts actions to overcome these evolving challenges. This process compresses diverse failure-recovery behaviors, compositional task variations, and environmental perturbations into minimal demonstrations. Our experiments demonstrate that ADC-trained models achieve superior compositional generalization to unseen task instructions, enhanced robustness to perceptual perturbations, and emergent error recovery capabilities. Strikingly, models trained with merely 20% of the demonstration volume collected through ADC significantly outperform traditional approaches using full datasets. These advances bridge the gap between data-centric learning paradigms and practical robotic deployment, demonstrating that strategic data acquisition, not merely post-hoc processing, is critical for scalable, real-world robot learning. Additionally, we are curating a large-scale ADC-Robotics dataset comprising real-world manipulation tasks with adversarial perturbations. This benchmark will be open-sourced to facilitate advancements in robotic imitation learning.
