---
title: Echo-4o-Image
canonical_url: "https://www.modelscope.cn/datasets/AI-ModelScope/Echo-4o-Image"
md_url: "https://www.modelscope.cn/datasets/AI-ModelScope/Echo-4o-Image.md"
repository: AI-ModelScope/Echo-4o-Image
last_updated: 2025-10-10
license: mit
storage_size: "313 GB"
downloads: 5556
stars: 2
---

# Echo-4o-Image

> Echo-4o-Image - AI-ModelScope 在 ModelScope 开源的数据集。Echo-4o-Image Dataset

AI-ModelScope/Echo-4o-Image 是 ModelScope 魔搭社区上的数据集，存储大小 313 GB，采用 mit 许可。

- **Repository**: AI-ModelScope/Echo-4o-Image
- **License**: mit
- **Storage size**: 313 GB
- **Downloads**: 5556
- **Stars**: 2
- **Last updated**: 2025-10-10

Source: https://www.modelscope.cn/datasets/AI-ModelScope/Echo-4o-Image

---

# Echo-4o-Image Dataset

[Paper](https://huggingface.co/papers/2508.09987) | [Project Page](https://yejy53.github.io/Echo-4o) | [Code](https://github.com/yejy53/Echo-4o)

## Introduction

Echo-4o-Image is a 180K-scale synthetic dataset generated by GPT-4o, designed to advance open-source models in image generation. While real-world image datasets are valuable, synthetic images offer crucial advantages, especially in addressing blind spots in real-world coverage:

*   **Complementing Rare Scenarios:** Synthetic data can generate examples for scenarios less represented in real-world datasets, such as surreal fantasy or multi-reference image generation, which are common in user queries.
*   **Clean and Controllable Supervision:** Unlike real-world data, which often contains complex background noise and misalignment between text and image, synthetic images provide pure backgrounds and long-tailed supervision signals, facilitating more accurate text-to-image alignment.

This dataset was instrumental in fine-tuning the unified multimodal generation baseline Bagel to obtain Echo-4o, demonstrating strong performance across standard benchmarks. Furthermore, Echo-4o-Image consistently enhances other foundation models (e.g., OmniGen2, BLIP3-o), highlighting its strong transferability.

## Echo-4o-Image Dataset Details

Echo-4o-Image is a large-scale synthetic dataset distilled from GPT-4o, containing approximately 179,000 samples. It spans three distinct task types:

*   **38K surreal fantasy generation tasks:** Designed to address imaginative content.
*   **73K multi-reference image generation tasks:** For scenarios requiring multiple visual cues.
*   **68K complex instruction execution tasks:** To improve adherence to detailed textual prompts.

For better visualization, an online gallery showcasing representative samples from our dataset is available: [Online Gallery](https://yejy53.github.io/Echo-4o/)

## Data Structure

The dataset typically organizes data within compressed packages (e.g., `.tar.gz` files referenced in `configs`). Inside these packages, data is arranged as follows:

```
- package_idx/
--- package_idx.json # metadata for samples in this package
--- images/
----- 00001.png
----- 00002.png
...
```


## Usage

This dataset can be used to train and fine-tune text-to-image models, extending capabilities to support multi-reference datasets.

### Training

The training process extends existing frameworks (e.g., Bagel's capabilities).
1.  **Data Preparation:** Follow data preparation guidelines, ensuring multi-reference data adheres to the expected format.
2.  **Training Process:** Training scripts use interfaces and parameters similar to established models (e.g., Bagel), allowing for seamless integration with existing training commands and configurations.

### Inference

*   **Text-to-Image Tasks:** For standard text-to-image generation, follow the inference process of base models (e.g., Bagel).
*   **Multi-Reference Tasks:** Specific examples and guides for tasks involving multiple references are provided in the [official GitHub repository](https://github.com/yejy53/Echo-4o).

### Code and Supporting Files

The associated GitHub repository provides crucial supporting files for working with the dataset:

*   **Attributes and Subjects:** `./code/attributes_and_subjects.json` contains dictionaries defining various attributes and subjects used in the dataset.
*   **Range-sensitive filtering:** `./code/range_sensitive_filter.json` contains metadata for data filtering, and `./code/data_filter.py` converts it for use in dataloaders.
*   **Data Loader:** `./code/dataloader.py` provides an example of how to load the data into image pairs, incorporating filtering and balanced resampling.

## Evaluation Benchmarks

The paper introduces two novel benchmarks for rigorously evaluating image generation capabilities:

*   **GenEval++:** Increases instruction complexity and uses an automated evaluator (powered by GPT-4.1) to mitigate score saturation and provide a more accurate assessment of text-to-image instruction following.
*   **Imagine-Bench:** Focuses on imaginative content, offering a comprehensive evaluation of conceptual creativity and visual consistency across dimensions like fantasy fulfillment, identity preservation, and aesthetic quality.

Detailed guides for these benchmarks can be found in the [EVAL section of the GitHub repository](https://github.com/yejy53/Echo-4o/blob/main/EVAL.md).

## Acknowledgements

We would like to thank the following open-source projects and research works:

*   [Bagel](https://github.com/ByteDance-Seed/Bagel)
*   [BLIP3o](https://github.com/JiuhaiChen/BLIP3o)
*   [OmniGen2](https://github.com/VectorSpaceLab/OmniGen2?tab=readme-ov-file)

## Citation

If you find this dataset or the associated work useful for your research, please cite the paper:

```bib
@article{ye2025echo4o,
      title={Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation}, 
      author={Junyan Ye, Dongzhi Jiang, Zihao Wang, Leqi Zhu, Zhenghao Hu, Zilong Huang, Jun He, Zhiyuan Yan, Jinghua Yu, Hongsheng Li, Conghui He, Weijia Li},
      journal={https://arxiv.org/abs/2508.09987},
      year={2025},
}
```
