---
title: Qwen3-8B-RL-W-GT-WebArena
canonical_url: "https://www.modelscope.cn/models/JunxuanLi/Qwen3-8B-RL-W-GT-WebArena"
md_url: "https://www.modelscope.cn/models/JunxuanLi/Qwen3-8B-RL-W-GT-WebArena.md"
repository: JunxuanLi/Qwen3-8B-RL-W-GT-WebArena
last_updated: 2026-09-04
license: Apache-2.0
model_type:
  - qwen3
architectures:
  - Qwen3ForCausalLM
base_model:
  - Qwen/Qwen3-8B
base_model_relation: finetune
parameters: 8.2B
tensor_type:
  - BF16
library_name:
  - safetensors
downloads: 27
stars: 0
---

# Qwen3-8B-RL-W-GT-WebArena

> Qwen3-8B-RL-W-GT-WebArena - JunxuanLi 在 ModelScope 开源的模型。Qwen3-8B-RL-W-GT-WebArena

JunxuanLi/Qwen3-8B-RL-W-GT-WebArena 是 ModelScope 魔搭社区上的 8.2B 参数机器学习模型，采用 Apache-2.0 许可，基于 Qwen/Qwen3-8B 构建。

- **Repository**: JunxuanLi/Qwen3-8B-RL-W-GT-WebArena
- **License**: Apache-2.0
- **Parameters**: 8.2B
- **Base model**: Qwen/Qwen3-8B
- **Downloads**: 27
- **Stars**: 0
- **Last updated**: 2026-09-04

Source: https://www.modelscope.cn/models/JunxuanLi/Qwen3-8B-RL-W-GT-WebArena

---

# Qwen3-8B-RL-W-GT-WebArena

<div align="center">
  <a href="https://github.com/THUNLP-MT/TTED">💻 GitHub</a>
  &nbsp;|&nbsp;
  <a href="https://arxiv.org/abs/xxxx.xxxxx">📄 arXiv</a>
</div>

## TTED

![TTED](./TTED.png)

**TTED** (**Test-Time Environment Decomposition**) is a label-free test-time adaptation framework for LLM web agents that addresses compositional generalization failures by decomposing a complex task state into aligned sub-goals and task-relevant sub-observations. It gathers self-assessed interaction experience within these localized decision contexts and then adapts the agent through query-level in-context learning for single-turn static webpages or task-level reinforcement learning for multi-turn realistic environments, allowing locally acquired improvements in decomposition and action generation to be composed when solving the original task.  

Empirically, TTED outperforms alternative test-time adaptation strategies, indicating that learning within decomposed local decision contexts is more effective than directly optimizing behavior in the full, compositionally complex environment:
* **Training with Ground-Truth Labels:** Despite requiring no ground-truth supervision, TTED can outperform supervised SFT or RL conducted in the original environment, where sparse and delayed task-level rewards provide weak learning signals.
* **Standard Test-Time Training (TTT):** Unlike monolithic TTT, TTED explicitly decomposes both the task goal and observation into aligned local contexts, reducing irrelevant information, and noise in self-assessed rewards.
* **Test-Time Reinforcement Learning (TTRL):** Whereas TTRL relies on matching-based self-consistency rewards that primarily supervise action outputs, TTED uses principled generative self-assessment to evaluate sub-goal alignment, sub-observation relevance, and action correctness, yielding more comprehensive and reliable signals for scalable test-time learning. 

> This model is trained from Qwen3-8B via reinforcement learning (RL) with ground-truth rewards on WebArena, which is specialized for complex web automation tasks. 

* The `record_RL_W_GT.zip` contains the experimental trajectories of this model on WebArena reported in our paper *Learning Simple Test-Time Environments for LLM Web Agents*.

## Citation

If you find this model useful, please cite the following paper:

```bibtex
@misc{li2026learning,
    title={Learning Simple Test-Time Environments for LLM Web Agents}, 
    author={Junxuan Li and Zijun Liu and Ziyi Huang and Peng Li and Yuzhou Liu and Ming Yan and Yang Liu},
    year={2026},
    eprint={2608.29305},
    archivePrefix={arXiv},
    primaryClass={cs.CL},
    url={https://arxiv.org/abs/2608.29305}, 
}
```
