---
title: XS-VID-v2
canonical_url: "https://www.modelscope.cn/datasets/lanlanlanrr/XS-VID-v2"
md_url: "https://www.modelscope.cn/datasets/lanlanlanrr/XS-VID-v2.md"
repository: lanlanlanrr/XS-VID-v2
last_updated: 2026-10-05
license: cc-by-4.0
storage_size: "31 GB"
downloads: 1132
stars: 0
---

# XS-VID-v2

> XS-VID-v2 - lanlanlanrr 在 ModelScope 开源的数据集。XS-VID v2 is a benchmark for extremely small object detection and tracking in videos. This release uses one canonical COCO-VID-style annotation contract across the Detection, MOT, and SOT tracks.

lanlanlanrr/XS-VID-v2 是 ModelScope 魔搭社区上的数据集，存储大小 31 GB，采用 cc-by-4.0 许可。

- **Repository**: lanlanlanrr/XS-VID-v2
- **License**: cc-by-4.0
- **Storage size**: 31 GB
- **Downloads**: 1132
- **Stars**: 0
- **Last updated**: 2026-10-05

Source: https://www.modelscope.cn/datasets/lanlanlanrr/XS-VID-v2

---

# XS-VID v2

XS-VID v2 is a benchmark for extremely small object detection and tracking in videos. This release uses one canonical COCO-VID-style annotation contract across the Detection, MOT, and SOT tracks.

The `annotations/`, `archives/`, `sot/`, `tools/`, and `manifests/` directories
form this v2 release. All 222,924 benchmark images are packaged in 16 archives.

## Download and Reproduce

Project: https://gjhhust.github.io/XS-VID/  
Code and task commands: https://github.com/gjhhust/YOLOFT  
Dataset mirrors: [Hugging Face](https://huggingface.co/datasets/lanlanlan23/XS-VID-v2),
[ModelScope](https://www.modelscope.cn/datasets/lanlanlanrr/XS-VID-v2).  
Checkpoints: https://huggingface.co/lanlanlan23/YOLOFT-XSVID-v2

From the YOLOFT repository, either command downloads all dataset files, checks
the SHA-256 of every media archive, extracts the images and prepares portable
labels/splits. ModelScope is the domestic mirror; neither command needs a token
once the repositories are public. Allow about 34 GB for download plus extracted
images and labels.

```bash
git clone https://github.com/gjhhust/YOLOFT.git
cd YOLOFT
python tools/xsvid/download_release.py --hub ms --destination ./downloads
# Alternative dataset mirror:
python tools/xsvid/download_release.py --hub hf --destination ./downloads
# Task checkpoints, portable TransT initialization and OSNet:
python tools/xsvid/download_release.py --hub hf --component model --destination ./downloads
```

See the code README for full VID/MOT/SOT inference commands and the bounded
training-startup command. The corrected paper SOT checkpoint has SHA-256
`c59639f868a5fdfa5474f072d657bafb072430c7386b12901e3bc1ac7a42188a`.
Do not substitute the early 57.09-AUC SOT checkpoint.

## Layout

```text
XS-VID-v2/
  annotations/
    train.json
    test.json
    train_dev.json
    val_dev.json
    protocols/
      paper_detection_test.json
  images/
  sot/
  tools/
  manifests/
```

The downloadable v2 media are divided into 16 video-aligned tar archives.
Extract every file in `archives/` into the same dataset root:

```bash
for archive in archives/XSVID_v2_images_*_of_16.tar; do
  tar -xf "$archive"
done
```

This creates the `images/<video>/` tree expected by the released loaders.
Verify every archive against `archives/manifest.json` before use.

`train.json` and `test.json` are the corrected canonical benchmark partitions. `train_dev.json` and `val_dev.json` are a development-only partition of the training data. Final benchmark models train on the full `train.json` and are evaluated on `test.json`.

`annotations/protocols/paper_detection_test.json` is a compact detection-only compatibility view of the evaluation annotations used for the Detection results reported in the paper. It is provided to reproduce the published table values exactly; new benchmark results should use the corrected canonical `annotations/test.json`.

## Annotation contract

The files follow COCO-VID conventions (`videos`, `images`, `annotations`, and `categories`) and retain XS-VID fields needed by the three tracks.

- `track_id` is the corrected dataset-global trajectory identity used by MOT and TAO-compatible evaluation; the accompanying `tracks` array provides its video and category metadata.
- In `test.json`, `paper_track_id` is present only when the trajectory identity used for the submitted paper differs from `track_id`. The paper-compatible evaluator falls back to `track_id` for all other annotations.
- `segment` is retained when available for rich per-instance annotations. The loader falls back to `segmentation` if `segment` is absent.
- `frame_index` is retained for temporal ordering. The loader falls back to `frame_id` if `frame_index` is absent.
- `ignore` is preserved both as category ID 4 and as an annotation flag.

The sparse compatibility field avoids duplicating identity metadata on unaffected annotations while keeping canonical and paper-result evaluation in the same annotation view.

## Native video loading

The legacy Detection loader requires the portable labels and split lists prepared
by the download command. MOT uses canonical JSON-derived video indices; SOT uses
the sequence view of the same image tree. For training, the code release provides
`tools/xsvid/prepare_unified_train_data.py` and a bounded startup smoke command.
Canonical JSON remains the source of truth; derived labels do not replace it.

## YOLO-compatible labels

The canonical JSON is the source of truth. Generate a portable YOLO layout after downloading:

```bash
python tools/prepare_yolo_labels.py --root /path/to/XS-VID-v2
```

The command creates `labels/`, `splits/`, and `xsvid.yaml`. It has no symbolic-link or machine-path dependency. The default export matches the original YOLOFT recipe and retains the seven categories, including `ignore`. Use `--exclude-ignore` only for the explicit six-class no-ignore ablation. The canonical JSON remains unchanged in either case.

Use `--splits-only` to regenerate only the small split-list files after relocating an existing release.

## Track protocols

- **Detection:** use `annotations/test.json` and the detection evaluator in the code release. To reproduce the Detection values reported in the paper exactly, evaluate the released paper predictions against `annotations/protocols/paper_detection_test.json`.
- **MOT:** use canonical `track_id` in `annotations/test.json` for new results. Use sparse `paper_track_id` with fallback to `track_id` only to reproduce the submitted paper result.
- **SOT:** `sot/test.txt` lists the official sequences, `sot/image_links.json` maps their frames to the shared image tree, and `sot/sequences/<name>/groundtruth.txt` contains the targets. `sot/sequences.json` records their canonical trajectory mapping. Run `python tools/materialize_sot_images.py --sot-root sot --images images` only when a conventional per-sequence image layout is needed.

The `manifests/` directory contains SHA-256 checksums and record counts for the canonical annotations. Validate the benchmark annotations and media after download:

```bash
python tools/validate_release_annotations.py annotations/train.json --output /tmp/xsvid_train_validation.json
python tools/validate_release_annotations.py annotations/test.json --output /tmp/xsvid_test_validation.json
python tools/validate_release_media.py --root /path/to/XS-VID-v2 --output /tmp/xsvid_media_validation.json
```

## Citation

```bibtex
@article{guo2026xsvid,
  title={XS-VID: A Large-Scale Benchmark for Small Object Detection and Tracking in Videos},
  author={Guo, Jiahao and Xu, Ziyang and Wu, Lianjun and Gao, Fei and Liu, Wenyu and Wang, Xinggang},
  journal={IEEE Transactions on Pattern Analysis and Machine Intelligence},
  year={2026},
  doi={10.1109/TPAMI.2026.3741044}
}
```

## License

XS-VID v2 is released under [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/).
