---
title: "Unified Reward Model for Multimodal Understanding and Generation"
canonical_url: "https://www.modelscope.cn/papers/124332"
md_url: "https://www.modelscope.cn/papers/124332.md"
arxiv_id: 2503.05236
published: 2025-03-07
last_updated: 2025-03-07
authors:
  - "Yibin Wang"
  - "Yuhang Zang"
  - "Hao Li"
  - "Cheng Jin"
  - "Jiaqi Wang"
model_name: UNIFIEDREWARD
model_developer: "复旦大学, 上海人工智能创新研究院, 上海人工智能实验室, 上海科学院人工智能研究所"
domain:
  - "计算机视觉"
  - "深度学习"
  - "自然语言处理"
type:
  - "计算机视觉"
  - "深度学习"
  - "自然语言处理"
  - "Computer Vision and Pattern Recognition (cs.CV)"
arxiv_url: "https://arxiv.org/abs/2503.05236"
pdf_url: "https://arxiv.org/pdf/2503.05236.pdf"
---

# Unified Reward Model for Multimodal Understanding and Generation

> Recent advances in human preference alignment have significantly enhanced multimodal generation and understanding. A key approach is training reward models to guide preference optimization. However, existing models are often task-specific, limiting their…

「Unified Reward Model for Multimodal Understanding and Generation」是 ModelScope 魔搭社区收录的论文，arXiv 2503.05236，作者为 Yibin Wang, Yuhang Zang, Hao Li et al.，发表于 2025-03-07，属于 计算机视觉、深度学习、自然语言处理 领域。

- **ArXiv**: 2503.05236
- **Published**: 2025-03-07
- **Authors**: Yibin Wang, Yuhang Zang, Hao Li, Cheng Jin, Jiaqi Wang
- **Model**: UNIFIEDREWARD
- **Developer**: 复旦大学, 上海人工智能创新研究院, 上海人工智能实验室, 上海科学院人工智能研究所
- **Domain**: 计算机视觉, 深度学习, 自然语言处理
- **ArXiv URL**: https://arxiv.org/abs/2503.05236
- **PDF**: https://arxiv.org/pdf/2503.05236.pdf

Source: https://www.modelscope.cn/papers/124332

---

> UNIFIEDREWARD：多模态理解和生成的首个统一奖励模型

## 摘要

本文提出了一种名为UNIFIEDREWARD的统一奖励模型，旨在解决现有奖励模型在多模态理解和生成任务中适应性差的问题。现有的奖励模型通常是针对特定任务设计的，限制了它们在不同视觉应用中的通用性和适应性。为此，作者构建了一个大规模的人类偏好数据集，涵盖了图像和视频的理解与生成任务，并基于此数据集训练了UNIFIEDREWARD模型。该模型能够进行成对排名和点评分，适用于多种视觉任务的偏好对齐。

具体来说，UNIFIEDREWARD通过三个阶段的流程实现：1) 统一奖励模型训练；2) 偏好数据构建；3) 生成/理解模型对齐。首先，使用大规模人类偏好数据集训练模型；其次，利用训练好的模型自动构建高质量的偏好对数据，通过成对排名和点过滤逐步筛选输出；最后，使用这些偏好对数据进行直接偏好优化（DPO），以使视觉模型的输出更符合人类偏好。

实验结果表明，联合学习多个视觉任务可以显著提高各任务的表现，特别是在图像和视频的理解和生成方面。该方法不仅提高了模型性能，还展示了跨任务协同效应，使得不同视觉领域的评估能力得到增强。

## Abstract

Recent advances in human preference alignment have significantly enhanced multimodal generation and understanding. A key approach is training reward models to guide preference optimization. However, existing models are often task-specific, limiting their adaptability across diverse visual applications. We also argue that jointly learning to assess multiple tasks may foster a synergistic effect, where improved image understanding enhances image generation assessment, and refined image evaluation benefits video assessment through better frame analysis. To this end, this paper proposes UnifiedReward, the first unified reward model for multimodal understanding and generation assessment, enabling both pairwise ranking and pointwise scoring, which can be employed for vision model preference alignment. Specifically, (1) we first develop UnifiedReward on our constructed large-scale human preference dataset, including both image and video generation/understanding tasks. (2) Then, it is utilized to automatically construct high-quality preference pair data based on the vision models, fine-gradually filtering their outputs through pair ranking and point sifting. (3) Finally, these data are used for their preference alignment through Direct Preference Optimization (DPO). Experimental results demonstrate that joint learning to assess diverse visual tasks can lead to substantial mutual benefits and we apply our pipeline to both image and video understanding/generation tasks, significantly improving the performance in each domain.
