---
title: AgenticArXiv-RL-Qwen3-VL-4B-FigureQA
canonical_url: "https://www.modelscope.cn/models/Algorineko/AgenticArXiv-RL-Qwen3-VL-4B-FigureQA"
md_url: "https://www.modelscope.cn/models/Algorineko/AgenticArXiv-RL-Qwen3-VL-4B-FigureQA.md"
repository: Algorineko/AgenticArXiv-RL-Qwen3-VL-4B-FigureQA
last_updated: 2026-09-26
license: mit
model_type:
  - qwen3_vl
architectures:
  - Qwen3VLForConditionalGeneration
base_model:
  - Qwen/Qwen3-VL-4B-Instruct
base_model_relation: finetune
parameters: 4.4B
tensor_type:
  - BF16
library_name:
  - transformer
  - safetensors
downloads: 5
stars: 0
tags:
  - vision-language
  - figure-analysis
  - arxiv
  - agent
  - sft
---

# AgenticArXiv-RL-Qwen3-VL-4B-FigureQA

> AgenticArXiv-RL-Qwen3-VL-4B-FigureQA - Algorineko 在 ModelScope 开源的模型。AgenticArXiv-RL-Qwen3-VL-4B-FigureQA

Algorineko/AgenticArXiv-RL-Qwen3-VL-4B-FigureQA 是 ModelScope 魔搭社区上的 4.4B 参数机器学习模型，采用 mit 许可，基于 Qwen/Qwen3-VL-4B-Instruct 构建。

- **Repository**: Algorineko/AgenticArXiv-RL-Qwen3-VL-4B-FigureQA
- **License**: mit
- **Parameters**: 4.4B
- **Base model**: Qwen/Qwen3-VL-4B-Instruct
- **Tags**: vision-language, figure-analysis, arxiv, agent, sft
- **Downloads**: 5
- **Stars**: 0
- **Last updated**: 2026-09-26

Source: https://www.modelscope.cn/models/Algorineko/AgenticArXiv-RL-Qwen3-VL-4B-FigureQA

---

# AgenticArXiv-RL-Qwen3-VL-4B-FigureQA

**中文** · [English](#english)

## 中文

[AgenticArXiv-RL](https://github.com/Algorineko/AgenticArXiv-RL) 工具集演进 T5（`analyze_figure` 图表分析）的 **env 侧 VLM**：在 Qwen3-VL-4B-Instruct 上用 arXiv 论文图表 + caption 做 LoRA 后训练（r=32），作为离线快照的图表分析录制引擎。策略模型仍是纯文本小模型——本模型只活在环境侧，不进入策略梯度。

- **基座**：[Qwen/Qwen3-VL-4B-Instruct](https://huggingface.co/Qwen/Qwen3-VL-4B-Instruct)（Apache-2.0）
- **训练数据**：项目离线快照 T4 记录中 **451 张 arXiv 图表（73 篇论文）**，监督目标全部来自论文作者写的 caption：
  - `describe` → caption 开头 1–2 句
  - `axes` / `trend` → caption 中命中坐标轴/趋势线索词的句子（与项目抽取式后端同一套规则）
- **训练**：LoRA r=32（仅语言层，1.47% 可训练参数）、3 epoch、lr 1e-4、seed 42；单卡 64GB 训练 34.5 分钟、峰值显存 17.9 GiB；训练后已合并为全量权重

### 留出评测（11 篇论文 94 图，按论文切分，贪心解码）

对 caption 证据的 ROUGE F1（**代理指标**）：

| 问法 | n | base R1 | 训练后 R1 | base RL | 训练后 RL |
|---|---|---|---|---|---|
| describe | 65 | 0.108 | **0.136** | 0.081 | **0.111** |
| axes | 24 | 0.081 | **0.114** | 0.075 | **0.087** |
| trend | 5 | 0.178 | 0.049 | 0.118 | 0.049 |
| 合计 | 94 | 0.105 | **0.126** | 0.081 | **0.102** |

训练后回答收敛为 caption 风格的一句话（describe 平均 47 → 18 token）；空回答/异常 0。完整逐条结果与样例见仓库 `artifacts/vlm_figureqa_eval/`。

### 用法

项目内（录制裁剪后的离线快照，用于 T5 任务与 SFT 数据）：

```bash
FIGURE_ANALYSIS_BACKEND=vlm VLM_MODEL_PATH=<本模型目录> \
python -m AgenticArxiv.rl.backfill_figure_analysis \
  --snapshot data/mock_arxiv_snapshot.json --force --paper-id <arXiv id>
```

直接推理（三种问题提示词与项目一致：describe / axes / trend，回答一句到两句）：

```python
from PIL import Image
from transformers import AutoProcessor, AutoModelForImageTextToText

model_dir = "Algorineko/AgenticArXiv-RL-Qwen3-VL-4B-FigureQA"
processor = AutoProcessor.from_pretrained(model_dir)
model = AutoModelForImageTextToText.from_pretrained(model_dir, dtype="bfloat16").eval()

image = Image.open("figure.png").convert("RGB")
messages = [{"role": "user", "content": [
    {"type": "image", "image": image},
    {"type": "text", "text": "Describe what this figure shows in one or two sentences."},
]}]
inputs = processor(text=[processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)],
                   images=[image], return_tensors="pt")
out = model.generate(**inputs, max_new_tokens=96, do_sample=False)
print(processor.batch_decode(out[:, inputs["input_ids"].shape[1]:], skip_special_tokens=True)[0])
```

### 诚实说明与限制

- 评测是 ROUGE 对 caption 的**代理指标**，不是人工质量评判；回答可以正确但不与 caption 重叠，也可能照抄 caption 而答非所问。
- `trend` 留出仅 5 条且本次测量回退——按字面报告，不做挑选；`axes`/`trend` 的监督目标来自 caption 关键词匹配，本身带噪声。
- T5 任务引用的 4 篇论文均在训练集内（快照录制属 in-sample；泛化能力由留出集承担）。
- 只理解 arXiv 论文内嵌图表（项目抽取器产出的 PNG/JPG）；训练与评测全部离线，不含任何在线服务。
- 基座 Qwen3-VL-4B-Instruct 为 Apache-2.0；本项目以 MIT 许可发布。项目地址：[AgenticArXiv-RL](https://github.com/Algorineko/AgenticArXiv-RL)。

## English

The env-side VLM for tool T5 (`analyze_figure`) of [AgenticArXiv-RL](https://github.com/Algorineko/AgenticArXiv-RL): a LoRA post-train (r=32, merged) of Qwen3-VL-4B-Instruct on arXiv figures + author captions, used as the offline snapshot recording engine. The policy stays text-only; this model lives only inside the environment.

- Data: 451 figures from 73 papers (project offline snapshot, T4 records); targets are the authors' captions (first 1–2 sentences for `describe`; caption sentences matching axes/trend hints for the other two)
- Training: LoRA r=32 on the language layers only (1.47% of parameters), 3 epochs, lr 1e-4, seed 42; 34.5 min on a single 64GB accelerator, peak 17.9 GiB
- Held-out (11 papers, 94 figure-question pairs, greedy): overall ROUGE-1 0.105 → **0.126**, ROUGE-L 0.081 → **0.102** against caption evidence; `describe` 0.108 → **0.136**, `axes` 0.081 → **0.114**, `trend` regressed on n=5 and is reported as-is
- Use it via `FIGURE_ANALYSIS_BACKEND=vlm VLM_MODEL_PATH=<model dir>` with the project's `backfill_figure_analysis`, or directly with `AutoModelForImageTextToText` (same three prompt templates as above)
- Caveats: ROUGE-against-caption is a proxy, not human judgement; caption-hint targets are noisy; the four T5 task papers are in-distribution; figures are arXiv embedded images only. Base model Apache-2.0, this release MIT.
