---
title: SUM
canonical_url: "https://www.modelscope.cn/datasets/KAIWANG/SUM"
md_url: "https://www.modelscope.cn/datasets/KAIWANG/SUM.md"
repository: KAIWANG/SUM
chinese_name: "SUM｜高分辨率细粒度视觉感知评测集"
last_updated: 2026-10-08
license: other
downloads: 183258
stars: 0
---

# SUM

> SUM - KAIWANG 在 ModelScope 开源的数据集。700 道选择题与 29 种对齐视图，用于比较上下文、分辨率和背景对细粒度视觉感知的影响。魔搭镜像暂缺两张经平台审核撤回的图片，详见说明。 700 questions with 29 aligned image views for controlled fine-grained visual perception evaluation. Paper forthcoming.

KAIWANG/SUM 是 ModelScope 魔搭社区上的数据集，采用 other 许可。

- **Repository**: KAIWANG/SUM
- **License**: other
- **Downloads**: 183258
- **Stars**: 0
- **Last updated**: 2026-10-08

Source: https://www.modelscope.cn/datasets/KAIWANG/SUM

---

<table align="center" role="presentation">
  <tr>
    <td align="center"><img src="./assets/rl-mind-logo.webp" alt="RL-MIND research group logo" width="160" /></td>
    <td align="center"><img src="./assets/sum-logo.png" alt="SUM dataset logo" width="160" /></td>
  </tr>
</table>
<h1 align="center">SUM</h1>
<p align="center"><b>High-Resolution Fine-Grained Perception Across Controlled Image Views</b></p>
<p align="center">
  <a href="https://huggingface.co/datasets/RL-MIND/SUM">🤗 Hugging Face</a> ·
  <a href="https://modelscope.cn/datasets/KAIWANG/SUM">🤖 ModelScope</a> ·
  <a href="./README.md">English</a> · <a href="./README_ZH.md">中文</a>
</p>
<p align="center">700 questions · 29 aligned views · 5 task categories</p>

<!-- kaiwang-zh-overview:start -->
## 中文介绍：高分辨率细粒度视觉感知评测集

**SUM** 面向高分辨率图像的细粒度视觉感知评测，包含 **700 道选择题、29 种对齐视图和 5 类任务**。同一问题分别对应原图、局部裁剪、白底贴回、全图降采样以及仅保留目标框像素的视图，可用于配对比较不同视觉条件下的模型表现。

完整发布含 **20,300 张视图图像**，问题来自 VStarBench、ZoomBench 和 HR-Bench。29 种视图共享基础问题 ID，不应视为独立题目。**魔搭默认分支目前缺少两张被平台自动审核撤回的图像，具体路径见下方镜像说明。** 配套论文尚未发布至 arXiv，暂以数据集链接和使用版本注明来源。

[阅读完整中文说明](https://www.modelscope.cn/datasets/KAIWANG/SUM/file/view/master/README_ZH.md?status=1)

<!-- kaiwang-zh-overview:end -->

## 🔎 Introduction

**SUM** is an evaluation collection for studying fine-grained visual perception under controlled changes to image context, resolution, and background. Each of its **700 multiple-choice questions** is paired with **29 aligned image views**, so the same question can be evaluated on the full image, a region crop, a white canvas, a downsampled image, or a view retaining only the target bounding-box pixels.

The collection draws on selected samples from **VStarBench (115)**, **ZoomBench (485)**, and **HR-Bench (100)**. The release contains **20,300 view images**, representing 700 questions and **664 distinct original-image file hashes**. These are different units: the transformed views are not additional independent questions.

**Paper status:** the accompanying paper has not yet been posted on arXiv. A paper link and formal paper citation will be added when available. This card describes the released files and does not report unannounced experimental results.

![Dataset statistics](./assets/dataset-overview.png)

## 🎯 What can be compared?

| View family | Variants | Images | What changes |
| --- | --- | ---: | --- |
| Original | Full image | 700 | Reference scene and resolution |
| Crop | 1, 1.5, 3, 6, 12, 18 × bbox area | 4,200 | Amount of context around the target |
| Whiteground | The same 6 crop settings | 4,200 | Crop pasted at its original position on a full-size white canvas |
| Downsample | 11 linear scales from 0.9375 to 0.0625 | 7,700 | Full-scene spatial resolution |
| Target-only | 1, 3, 6, 18 × crop canvas, plus original canvas | 3,500 | Only target-bbox pixels retained; all other pixels replaced with white |

Crop multipliers describe **area**, while downsample factors describe **width and height**. These quantities should not be compared as if they were the same scale. The target-only view preserves pixels inside a bounding box; it is not an instance-segmentation mask. Whiteground preserves all content inside the selected crop, including its surrounding context.

![Matched examples across the five view families](./assets/view-examples.png)

Thumbnails are resized for display. Pixel dimensions below each thumbnail describe the stored image. Released benchmark images retain their original encoding and resolution.

## 📊 Dataset composition

| Task | Questions |
| --- | ---: |
| Color Attributes | 347 |
| OCR | 118 |
| Structural Attributes | 94 |
| Object Identification | 84 |
| Material Attributes | 57 |
| **Total** | **700** |

There are **679 four-option** and **21 two-option** questions. Questions are primarily in English; some OCR targets and a small number of source options contain Chinese or Japanese text. VStarBench samples come from `direct_attributes`; HR-Bench samples come from the local `FSP_8k_100` selection. The category mapping used by this release is preserved in the legacy metadata, including the historical merge of Visual Prompting and Fine-Grained Counting into Object Identification.

## 📦 Files and configurations

```text
README.md / README_ZH.md          Dataset cards
data/<view>/test/metadata.jsonl   700 image-linked rows per view
data/<view>/test/images/          Original encoded image files
dataset.json                     700 canonical questions; original-view image paths
metadata/image_manifest.json     Per-image SHA-256, byte size and dimensions
metadata/validation_summary.json Full release validation summary
metadata/errata.json              Known issues and historical corrections
metadata/legacy/                  Legacy metadata with portable paths
metadata/path_mapping.json        Mapping from the previous directory layout
scripts/prompts.py                Normalized and legacy prompt constructors
scripts/evaluate.py               Exact option-letter accuracy
assets/                          Logos and data-derived figures
THIRD_PARTY_NOTICES.md            Source attribution and terms
```

The default configuration is **`original`**. Every configuration has a single **`test`** split with 700 rows. Examples include `crop_3x`, `whiteground_3x`, `downsample_0_25`, and `target_only_original`. The complete configuration list appears in the dataset selector and `metadata/validation_summary.json`.

| Field | Meaning |
| --- | --- |
| `id`, `base_sample_id` | Stable question ID shared across views |
| `image` | Decoded image when loaded through ImageFolder |
| `file_name` | Image path relative to the corresponding metadata file, in raw JSONL |
| `image_group_id` | SHA-256 of the original image file; identical-image questions share this ID |
| `source_dataset`, `source_question_id`, `source_image_id` | Source provenance |
| `question`, `question_raw` | Normalized question stem and original source question text |
| `choices`, `answer`, `answer_text` | Option list, correct letter and corresponding text |
| `task_type` | One of the five released task categories |
| `view_id`, `view_family` | Exact transformation setting and family |
| `width`, `height` | Dimensions of the current view |
| `bbox_original`, `bbox_clipped`, `bbox_view` | Source, in-bounds original-image, and current-view coordinates |
| `crop_box_original` | Crop or paste region in original-image coordinates |
| `crop_area_multiplier`, `downsample_linear_scale` | Separate parameters for the two scale definitions |

Coordinates are pixel-space `xyxy`. `bbox_view` for downsampling uses the actual output width and height, including rounding. Legacy metadata retains the crop boxes, effective area multipliers and pixel rules used to generate the released images. Paths beginning with `upstream://` inside legacy metadata are provenance identifiers, not additional bundled files.

## 🚀 Quick start

### Hugging Face: load one view

```python
from datasets import load_dataset

dataset = load_dataset("RL-MIND/SUM", "original", split="test", streaming=True)
sample = next(iter(dataset))
print(sample["question"], sample["choices"], sample["answer"])
image = sample["image"]

# Use the same question IDs for a paired comparison.
crops = load_dataset("RL-MIND/SUM", "crop_3x", split="test", streaming=True)
```

### Download an individual sample

```python
import json
from huggingface_hub import hf_hub_download

annotations = hf_hub_download("RL-MIND/SUM", "dataset.json", repo_type="dataset")
with open(annotations, encoding="utf-8") as f:
    sample = json.load(f)[0]
image_path = hf_hub_download("RL-MIND/SUM", sample["image"], repo_type="dataset")
```

### ModelScope mirror

**Mirror availability (2026-10-08):** ModelScope's automated review reverted two image files after upload. The default ModelScope branch currently contains **20,298 of the 20,300 images**. The complete release is available on Hugging Face. All annotations and question IDs are retained; use the complete release for paired evaluations involving these files:

- `data/crop_1_5x/test/images/000602.png`
- `data/downsample_0_75/test/images/000075.jpg`

```python
import json
from modelscope_hub import HubApi

api = HubApi()
annotation_path = api.download_file("KAIWANG/SUM", "dataset", "dataset.json")
with open(annotation_path, encoding="utf-8") as f:
    sample = json.load(f)[0]
image_path = api.download_file("KAIWANG/SUM", "dataset", sample["image"])
```

Install `datasets`, `huggingface_hub`, `Pillow`, or `modelscope-hub` as needed. Available image files and annotations are identical across the two hosts; the two ModelScope omissions are listed above. Download a selected configuration or individual images when you do not need the full collection (approximately 44.4 GB before platform storage deduplication).

## ✅ Evaluation and release notes

- Evaluate the same IDs under each view and report accuracy by source, task and view. `scripts/evaluate.py` scores explicit option letters; absent or invalid answers count as incorrect.
- The 485 ZoomBench source question strings already include choices. `question` separates their stem from the choices; `question_raw` preserves the old text. `scripts/prompts.py` offers both the normalized prompt and the legacy prompt that repeats those choices. Prompt changes can affect results, so record which version was used.
- Thirty source bounding boxes extend slightly beyond image boundaries. The source coordinates are retained alongside clipped coordinates; source image bytes and existing transformed views are preserved.
- Sample `000669` contains the duplicate distractors `red` and `Red`; its recorded answer is `C` / `Brown`. The source options are retained and the issue is flagged in `quality_notes` and the errata. No replacement distractor was invented.
- Repeated original images can support different questions. Keep the same `image_group_id` and all of its views together if creating any downstream split. Do not treat the 29 paired views as independent data points.
- Every released image was fully decoded, checked against its expected dimensions and SHA-256 hashed during packaging. This structural validation is not a new human review of every semantic answer or a rerun of all transformation pixel-equivalence checks.

## 📜 Sources, terms and citation

Please credit [V*Bench](https://github.com/penghao-wu/vstar), [ZoomBench](https://huggingface.co/datasets/inclusionAI/ZoomBench), and [HR-Bench](https://huggingface.co/datasets/DreamMr/HR-Bench), as applicable. SUM provides a selected, normalized and transformed evaluation collection; it does not claim that the source questions or imagery were newly authored for this release.

Source-specific terms and unresolved upstream license details are documented in [THIRD_PARTY_NOTICES.md](./THIRD_PARTY_NOTICES.md). The metadata deliberately uses `license: other`; this is not a blanket permissive license for all underlying images. Until the accompanying paper is public, cite the dataset URL and the exact repository revision used. A formal paper citation will be added later.
