---
title: "GTR: Guided Thought Reinforcement Prevents Thought Collapse in RL-based VLM Agent Training"
canonical_url: "https://www.modelscope.cn/papers/125459"
md_url: "https://www.modelscope.cn/papers/125459.md"
arxiv_id: 2503.08525
published: 2025-03-11
last_updated: 2025-03-11
authors:
  - "Tong Wei"
  - "Yijun Yang"
  - "Junliang Xing"
  - "Yuanchun Shi"
  - "Zongqing Lu"
  - "Deheng Ye"
model_name: "GTR (Guided Thought Reinforcement)"
model_developer: "清华大学, 腾讯, 北京大学"
domain:
  - "自然语言处理"
  - "计算机视觉"
  - "深度学习"
  - "机器学习"
type:
  - "自然语言处理"
  - "计算机视觉"
  - "深度学习"
  - "机器学习"
  - "Computer Vision and Pattern Recognition (cs.CV)"
  - "Artificial Intelligence (cs.AI)"
arxiv_url: "https://arxiv.org/abs/2503.08525"
pdf_url: "https://arxiv.org/pdf/2503.08525.pdf"
---

# GTR: Guided Thought Reinforcement Prevents Thought Collapse in RL-based VLM Agent Training

> Reinforcement learning with verifiable outcome rewards (RLVR) has effectively scaled up chain-of-thought (CoT) reasoning in large language models (LLMs). Yet, its efficacy in training vision-language model (VLM) agents for goal-directed action reasoning in…

「GTR: Guided Thought Reinforcement Prevents Thought Collapse in RL-based VLM Agent Training」是 ModelScope 魔搭社区收录的论文，arXiv 2503.08525，作者为 Tong Wei, Yijun Yang, Junliang Xing et al.，发表于 2025-03-11，属于 自然语言处理、计算机视觉、深度学习 领域。

- **ArXiv**: 2503.08525
- **Published**: 2025-03-11
- **Authors**: Tong Wei, Yijun Yang, Junliang Xing, Yuanchun Shi, Zongqing Lu, Deheng Ye
- **Model**: GTR (Guided Thought Reinforcement)
- **Developer**: 清华大学, 腾讯, 北京大学
- **Domain**: 自然语言处理, 计算机视觉, 深度学习, 机器学习
- **ArXiv URL**: https://arxiv.org/abs/2503.08525
- **PDF**: https://arxiv.org/pdf/2503.08525.pdf

Source: https://www.modelscope.cn/papers/125459

---

> 破解思维崩溃：GTR引领VLM智能体迈向高效推理与决策新纪元

## 摘要

本文针对在强化学习（RL）训练中视觉-语言模型（VLM）代理出现的“思维崩溃”问题，提出了一种新的方法——引导思维强化（GTR）。研究背景表明，在复杂任务环境中，仅基于最终动作结果的奖励机制会导致VLM代理的推理过程失去多样性，产生不合理的推理和错误的动作决策。为解决这一问题，作者提出了GTR框架，该框架通过引入自动化推理修正机制，结合策略优化算法PPO，同时对代理的推理过程和动作进行训练。GTR无需密集的人工标注，即可提供更有效的过程监督。实验部分在24点卡牌游戏和ALFWorld环境中验证了GTR的有效性。结果显示，与现有方法相比，GTR显著提高了LLaVA-7b模型的任务成功率和泛化能力，特别是在复杂视觉环境中表现尤为突出，成功率达到SOTA方法的3-5倍，同时模型规模更小。此外，GTR在不同任务中的成功应用证明了其广泛的适用性。

## Abstract

Reinforcement learning with verifiable outcome rewards (RLVR) has effectively scaled up chain-of-thought (CoT) reasoning in large language models (LLMs). Yet, its efficacy in training vision-language model (VLM) agents for goal-directed action reasoning in visual environments is less established. This work investigates this problem through extensive experiments on complex card games, such as 24 points, and embodied tasks from ALFWorld. We find that when rewards are based solely on action outcomes, RL fails to incentivize CoT reasoning in VLMs, instead leading to a phenomenon we termed thought collapse, characterized by a rapid loss of diversity in the agent's thoughts, state-irrelevant and incomplete reasoning, and subsequent invalid actions, resulting in negative rewards. To counteract thought collapse, we highlight the necessity of process guidance and propose an automated corrector that evaluates and refines the agent's reasoning at each RL step. This simple and scalable GTR (Guided Thought Reinforcement) framework trains reasoning and action simultaneously without the need for dense, per-step human labeling. Our experiments demonstrate that GTR significantly enhances the performance and generalization of the LLaVA-7b model across various visual environments, achieving 3-5 times higher task success rates compared to SoTA models with notably smaller model sizes.
