---
title: AgenticArXiv-RL-Qwen2.5-1.5B-GRPO-DAPO
canonical_url: "https://www.modelscope.cn/models/Algorineko/AgenticArXiv-RL-Qwen2.5-1.5B-GRPO-DAPO"
md_url: "https://www.modelscope.cn/models/Algorineko/AgenticArXiv-RL-Qwen2.5-1.5B-GRPO-DAPO.md"
repository: Algorineko/AgenticArXiv-RL-Qwen2.5-1.5B-GRPO-DAPO
last_updated: 2026-10-01
license: mit
model_type:
  - qwen2
architectures:
  - Qwen2ForCausalLM
base_model:
  - Qwen/Qwen2.5-1.5B-Instruct
base_model_relation: finetune
parameters: 1.5B
tensor_type:
  - F32
library_name:
  - safetensors
downloads: 4
stars: 0
tags:
  - agent
  - reinforcement-learning
  - tool-use
  - react
  - arxiv
  - grpo
  - dapo
---

# AgenticArXiv-RL-Qwen2.5-1.5B-GRPO-DAPO

> AgenticArXiv-RL-Qwen2.5-1.5B-GRPO-DAPO - Algorineko 在 ModelScope 开源的模型。AgenticArXiv-RL-Qwen2.5-1.5B-GRPO-DAPO

Algorineko/AgenticArXiv-RL-Qwen2.5-1.5B-GRPO-DAPO 是 ModelScope 魔搭社区上的 1.5B 参数机器学习模型，采用 mit 许可，基于 Qwen/Qwen2.5-1.5B-Instruct 构建。

- **Repository**: Algorineko/AgenticArXiv-RL-Qwen2.5-1.5B-GRPO-DAPO
- **License**: mit
- **Parameters**: 1.5B
- **Base model**: Qwen/Qwen2.5-1.5B-Instruct
- **Tags**: agent, reinforcement-learning, tool-use, react, arxiv, grpo, dapo
- **Downloads**: 4
- **Stars**: 0
- **Last updated**: 2026-10-01

Source: https://www.modelscope.cn/models/Algorineko/AgenticArXiv-RL-Qwen2.5-1.5B-GRPO-DAPO

---

# AgenticArXiv-RL-Qwen2.5-1.5B-GRPO-DAPO

**中文** · [English](#english)

## 中文

[AgenticArXiv-RL](https://github.com/Algorineko/AgenticArXiv-RL) 的 DAPO 变体发布权重，也是三变体对照（GSPO / Dr.GRPO / DAPO，见仓库 `docs/rl_paradigm_comparison.md`）中**综合表现最好**的一个：在与 [GRPO 基线](https://www.modelscope.cn/models/Algorineko/AgenticArXiv-RL-Qwen2.5-1.5B-GRPO)相同的任务、奖励与算力协议下，改用 **DAPO 目标**（clip-higher `epsilon_high=0.28`、截断完成掩码、无 KL 约束 β=0）。

- **基座**：[AgenticArXiv-RL-Qwen2.5-1.5B-SFT](https://www.modelscope.cn/models/Algorineko/AgenticArXiv-RL-Qwen2.5-1.5B-SFT)（与基线同一起点）
- **训练任务**：`data/splits/v6_grpo_train.json` 的 33 条信号任务，离线快照回放
- **配方**：公共配方与基线一致（`num_generations=4`、`lr=5e-6`、`batch_size=16`、`max_turns=6`、无奖励课程、seed 42、60 步）；**范式差异：`--dapo --beta 0`**（β=0 是 DAPO 预设的组成部分，基线为 0.1）
- **⚠ 与完整 DAPO 的差异**：**不含动态采样**。完整 DAPO（含动态采样）在本 33 任务池上第 5 步即耗尽提示池（SFT 起点零方差组占比过高，29/32 次重采样全被拒）；故本权重保留 DAPO 其余组件、关闭动态采样（manifest `dynamic_sampling.enabled=false` 如实记录）。

### 离线评测（seed 45 / repeat 3 / 离线快照回放 / regex agent）

严格成功率（pass³）：

| 切分 | SFT | GRPO 基线 | **DAPO（无动态采样）** |
|---|---|---|---|
| rl_train（训练任务诊断） | 0.081 | 0.636 | **0.667** |
| dev | 0.000 | 0.375 | **0.500** |
| iid_test | 0.056 | 0.444 | **0.481** |
| ood_test | 0.000 | 0.500 | **0.500** |

辅助指标：rl_train 工具准确率 0.85 / 虚假完成率 0.12（基线 0.85/0.15）；iid_test 0.65/0.35（基线 0.50/0.50）——泛化侧优势明显；ood_test 0.50/0.50 与基线持平。阶段验证通过（mean_reward 0.295，阈值 −0.2）。结论：**去掉 KL 约束 + clip-higher 在 dev/iid 上带来真实提升**，且训练全程无失稳。

### 限制

- 同基线：离线回放口径、33 条信号任务子集、dev/ood 样本量小（24/12）、单种子单次运行
- 非完整 DAPO：缺动态采样组件（原因见上），不应直接与原论文 DAPO 数字对标
- dev 的 +0.125 在 24 样本下置信区间宽

## English

DAPO variant of [AgenticArXiv-RL](https://github.com/Algorineko/AgenticArXiv-RL), and the strongest of the three-variant comparison (GSPO / Dr.GRPO / DAPO, see `docs/rl_paradigm_comparison.md`): the **DAPO objective** (clip-higher `epsilon_high=0.28`, truncated-completion masking, no KL with β=0) on the same tasks, reward and protocol as the [GRPO baseline](https://www.modelscope.cn/models/Algorineko/AgenticArXiv-RL-Qwen2.5-1.5B-GRPO).

- Base: [Algorineko/AgenticArXiv-RL-Qwen2.5-1.5B-SFT](https://www.modelscope.cn/models/Algorineko/AgenticArXiv-RL-Qwen2.5-1.5B-SFT)
- Recipe: shared recipe identical to baseline; paradigm delta: `--dapo --beta 0`
- **Caveat**: **without dynamic sampling** — full DAPO exhausted the 33-task prompt pool at step 5 (SFT-start zero-variance groups too frequent; 29/32 resamples rejected). The other DAPO components are kept; the manifest records `dynamic_sampling.enabled=false`.

Strict success (pass³, offline, seed 45, repeat 3): rl_train **0.667** / dev **0.500** / iid **0.481** / ood 0.500 — best of all variants on three of four splits (baseline: 0.636 / 0.375 / 0.444 / 0.500). iid tool accuracy 0.65 / false-finish 0.35 (baseline 0.50/0.50). Limitations: offline replay, 33-task subset, small held-out sets, single seed, not full DAPO.
