---
title: AgenticArXiv-RL-Qwen2.5-1.5B-SFT-T5
canonical_url: "https://www.modelscope.cn/models/Algorineko/AgenticArXiv-RL-Qwen2.5-1.5B-SFT-T5"
md_url: "https://www.modelscope.cn/models/Algorineko/AgenticArXiv-RL-Qwen2.5-1.5B-SFT-T5.md"
repository: Algorineko/AgenticArXiv-RL-Qwen2.5-1.5B-SFT-T5
last_updated: 2026-09-26
license: mit
model_type:
  - qwen2
architectures:
  - Qwen2ForCausalLM
base_model:
  - Qwen/Qwen2.5-1.5B-Instruct
base_model_relation: finetune
parameters: 1.5B
tensor_type:
  - F32
library_name:
  - safetensors
downloads: 5
stars: 0
tags:
  - agent
  - tool-use
  - react
  - arxiv
  - sft
  - figure-analysis
---

# AgenticArXiv-RL-Qwen2.5-1.5B-SFT-T5

> AgenticArXiv-RL-Qwen2.5-1.5B-SFT-T5 - Algorineko 在 ModelScope 开源的模型。AgenticArXiv-RL-Qwen2.5-1.5B-SFT-T5

Algorineko/AgenticArXiv-RL-Qwen2.5-1.5B-SFT-T5 是 ModelScope 魔搭社区上的 1.5B 参数机器学习模型，采用 mit 许可，基于 Qwen/Qwen2.5-1.5B-Instruct 构建。

- **Repository**: Algorineko/AgenticArXiv-RL-Qwen2.5-1.5B-SFT-T5
- **License**: mit
- **Parameters**: 1.5B
- **Base model**: Qwen/Qwen2.5-1.5B-Instruct
- **Tags**: agent, tool-use, react, arxiv, sft, figure-analysis
- **Downloads**: 5
- **Stars**: 0
- **Last updated**: 2026-09-26

Source: https://www.modelscope.cn/models/Algorineko/AgenticArXiv-RL-Qwen2.5-1.5B-SFT-T5

---

# AgenticArXiv-RL-Qwen2.5-1.5B-SFT-T5

**中文** · [English](#english)

## 中文

[AgenticArXiv-RL](https://github.com/Algorineko/AgenticArXiv-RL) 阶段 1 SFT 的 T5 版：在 Qwen2.5-1.5B-Instruct 上全参监督微调，训练数据为 v3_81 训练集派生的参数化专家轨迹（86 条派生任务，含 T5 图表分析 `analyze_figure` 的检索→下载→抽图→分析四步链）。相比已发布的 SFT 权重，本模型学会了图表分析工具链；T5 的观察结果由本项目的 FigureQA VLM 录制（`AgenticArXiv-RL-Qwen3-VL-4B-FigureQA`）。

- **基座**：[Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
- **数据**：249 条确定性专家轨迹 × 12 个语言视角 = 2988 行；派生自 `data/splits/v3_81.json:train`（heldout 重叠 0）
- **配方**：全参（无 LoRA）、2 epoch、lr 1e-4、batch 4 × grad_accum 2、max_length 2656、seed 42；单卡约 47 分钟
- **阶段验证**：通过（parse_rate 0.625 ≥ 0.3；同口径对照：已发布 SFT 0.750）
- **与已发布 SFT 的差异**：数据切分 v3_77 → v3_81、训练数据含 T5、max_length 2560 → 2656

### 离线评测（seed 45 / repeat 3 / 离线快照回放 / regex agent，与已发布模型同协议）

严格成功率（pass^3）：

| 切分 | n | SFT（已发布） | GRPO（已发布） | **本模型** |
|---|---|---|---|---|
| v6 rl_train（诊断，含 3 条 T5） | 99 | 0.081 | 0.636 | **0.485** |
| dev | 24 | 0.000 | 0.375 | **0.250** |
| iid_test（含 1 条 T5） | 54 | 0.056 | 0.444 | **0.278** |
| ood_test | 12 | 0.000 | 0.500 | **0.250** |

T5 图表分析任务细节（发布前两个模型在全部 4 条 T5 任务上 tool 正确率均为 0——轨迹显示执行到抽图就提前 FINISH）：

| 任务 | 结果 |
|---|---|
| analyze_cv5_ref3_trend | strict 3/3，完全学会 |
| analyze_cv5_ref1_desc | 工具调用步骤正确（tool 1.0），但 `question` 传成 `"tldr"`（与 summarize 的 style 枚举混淆）→ strict 0 |
| analyze_ai5_ref4_desc（iid） | 同上，`"tldr"` → strict 0 |
| analyze_ro5_ref2_axes | `question` 传成 `"trend"`/`"tldr"` → strict 0 |

### 用法

与已发布 SFT 完全一致（ReAct 文本 Agent，环境侧工具）：

```bash
python -m AgenticArxiv.benchmark.run_benchmark \
  --backend transformers --model <本模型路径> \
  --agents regex --repeat 3 --seed 45 \
  --offline --task-set expanded --split data/splits/v3_81.json:iid_test
```

### 诚实说明与限制

- T5 训练数据只覆盖 3 篇论文的派生问法（不复制基任务原文，防泄漏）；`question` 枚举在未见过的「父任务问法 × 图」组合上泛化不全——上述 `tldr`/`trend` 混淆如实报告，不做挑选。
- 整体严格成功率介于已发布 SFT 与 GRPO 之间：本模型是 SFT 阶段产物；其后可继续做 GRPO。
- 训练与评测均为**离线快照回放**，不代表连接真实 arXiv 的在线表现。
- 基座 Qwen2.5-1.5B-Instruct 为 Apache-2.0；本项目以 MIT 许可发布。项目地址：[AgenticArXiv-RL](https://github.com/Algorineko/AgenticArXiv-RL)。

## English

Stage-1 SFT (T5 edition) of [AgenticArXiv-RL](https://github.com/Algorineko/AgenticArXiv-RL): full-parameter SFT of Qwen2.5-1.5B-Instruct on parametric expert trajectories derived from the v3_81 train split (86 derived tasks, including the 4-step figure-analysis chain `analyze_figure`; observations recorded with this project's FigureQA VLM). Unlike the previously released SFT, this checkpoint has learned the figure-analysis tool.

Held-out strict success (pass^3, offline, seed 45, repeat 3): v6 rl_train 0.485 (released SFT 0.081 / GRPO 0.636), dev 0.250, iid_test 0.278, ood_test 0.250. On the four T5 tasks the previously released models scored tool-accuracy 0.0 (early FINISH); this model passes `analyze_cv5_ref3_trend` 3/3 and calls the tool at the right step on two more, but binds the wrong `question` enum (`"tldr"`/`"trend"`) on unseen parent-phrasing × figure combinations — reported as-is. Training/eval are offline snapshot replays. Base Apache-2.0, release MIT.
