---
title: "Motion Anything: Any to Motion Generation"
canonical_url: "https://www.modelscope.cn/papers/125267"
md_url: "https://www.modelscope.cn/papers/125267.md"
arxiv_id: 2503.06955
published: 2025-03-10
last_updated: 2025-03-10
authors:
  - "Zeyu Zhang"
  - "Yiran Wang"
  - "Wei Mao"
  - "Danning Li"
  - "Rui Zhao"
  - "Biao Wu"
  - "Zirui Song"
  - "Bohan Zhuang"
  - "Ian Reid"
  - "Richard Hartley"
model_name: "Motion Anything"
model_developer: "澳大利亚国立大学, 悉尼大学, 腾讯, 麦吉尔大学, 京东, 悉尼科技大学, MBZUAI, 浙江大学, 谷歌"
domain:
  - "计算机视觉"
  - "自然语言处理"
  - "深度学习"
type:
  - "计算机视觉"
  - "自然语言处理"
  - "深度学习"
  - "Computer Vision and Pattern Recognition (cs.CV)"
arxiv_url: "https://arxiv.org/abs/2503.06955"
pdf_url: "https://arxiv.org/pdf/2503.06955.pdf"
code_link: "https://steve-zeyu-zhang.github.io/MotionAnything"
---

# Motion Anything: Any to Motion Generation

> Conditional motion generation has been extensively studied in computer vision, yet two critical challenges remain. First, while masked autoregressive methods have recently outperformed diffusion-based approaches, existing masking models lack a mechanism to…

「Motion Anything: Any to Motion Generation」是 ModelScope 魔搭社区收录的论文，arXiv 2503.06955，作者为 Zeyu Zhang, Yiran Wang, Wei Mao et al.，发表于 2025-03-10，属于 计算机视觉、自然语言处理、深度学习 领域。

- **ArXiv**: 2503.06955
- **Published**: 2025-03-10
- **Authors**: Zeyu Zhang, Yiran Wang, Wei Mao, Danning Li, Rui Zhao, Biao Wu, Zirui Song, Bohan Zhuang, Ian Reid, Richard Hartley
- **Model**: Motion Anything
- **Developer**: 澳大利亚国立大学, 悉尼大学, 腾讯, 麦吉尔大学, 京东, 悉尼科技大学, MBZUAI, 浙江大学, 谷歌
- **Domain**: 计算机视觉, 自然语言处理, 深度学习
- **ArXiv URL**: https://arxiv.org/abs/2503.06955
- **PDF**: https://arxiv.org/pdf/2503.06955.pdf
- **Code**: https://steve-zeyu-zhang.github.io/MotionAnything

Source: https://www.modelscope.cn/papers/125267

---

> Motion Anything：多模态条件下的人体运动生成新突破

## 摘要

本文提出了Motion Anything，一个能够生成高质量、可控的人体运动的多模态框架。该方法解决了现有条件运动生成中的两个关键挑战：1) 现有的掩码自回归模型缺乏根据给定条件优先处理动态帧和身体部位的机制；2) 不同条件模态的方法无法有效集成多种模态，限制了生成运动的控制性和连贯性。为了解决这些问题，作者引入了一种基于注意力的掩码建模方法，能够在时空维度上对关键帧和动作进行细粒度控制。此外，还提出了Text-Motion-Dance (TMD) 数据集，包含2,153对文本、音乐和舞蹈样本，填补了社区在多模态条件下运动生成数据集的空白。实验结果表明，Motion Anything在多个基准测试中超越了现有方法，在HumanML3D上的FID指标提高了15%，并在AIST++和TMD数据集上表现出一致的性能提升。

## Abstract

Conditional motion generation has been extensively studied in computer vision, yet two critical challenges remain. First, while masked autoregressive methods have recently outperformed diffusion-based approaches, existing masking models lack a mechanism to prioritize dynamic frames and body parts based on given conditions. Second, existing methods for different conditioning modalities often fail to integrate multiple modalities effectively, limiting control and coherence in generated motion. To address these challenges, we propose Motion Anything, a multimodal motion generation framework that introduces an Attention-based Mask Modeling approach, enabling fine-grained spatial and temporal control over key frames and actions. Our model adaptively encodes multimodal conditions, including text and music, improving controllability. Additionally, we introduce Text-Motion-Dance (TMD), a new motion dataset consisting of 2,153 pairs of text, music, and dance, making it twice the size of AIST++, thereby filling a critical gap in the community. Extensive experiments demonstrate that Motion Anything surpasses state-of-the-art methods across multiple benchmarks, achieving a 15% improvement in FID on HumanML3D and showing consistent performance gains on AIST++ and TMD. See our project website this https URL
