---
title: OpenAI-4o_t2i_human_preference
canonical_url: "https://www.modelscope.cn/datasets/AI-ModelScope/OpenAI-4o_t2i_human_preference"
md_url: "https://www.modelscope.cn/datasets/AI-ModelScope/OpenAI-4o_t2i_human_preference.md"
repository: AI-ModelScope/OpenAI-4o_t2i_human_preference
last_updated: 2025-04-01
license: "Apache License 2.0"
storage_size: "4.8 GB"
downloads: 891
stars: 0
---

# OpenAI-4o_t2i_human_preference

> OpenAI-4o_t2i_human_preference - AI-ModelScope 在 ModelScope 开源的数据集。.vertical-container { display: flex; flex-direction: column; gap: 60px; } .horizontal-container { display: flex; flex-direction: row; justify-content: center; gap: 60px; }

AI-ModelScope/OpenAI-4o_t2i_human_preference 是 ModelScope 魔搭社区上的数据集，存储大小 4.8 GB，采用 Apache License 2.0 许可。

- **Repository**: AI-ModelScope/OpenAI-4o_t2i_human_preference
- **License**: Apache License 2.0
- **Storage size**: 4.8 GB
- **Downloads**: 891
- **Stars**: 0
- **Last updated**: 2025-04-01

Source: https://www.modelscope.cn/datasets/AI-ModelScope/OpenAI-4o_t2i_human_preference

---

<style>

.vertical-container {
    display: flex;  
    flex-direction: column;
    gap: 60px;  
}
  
.horizontal-container {
    display: flex;  
    flex-direction: row;
    justify-content: center;
    gap: 60px;  
}

.image-container img {
  max-height: 250px; /* Set the desired height */
  margin:0;
  object-fit: contain; /* Ensures the aspect ratio is maintained */
  width: auto; /* Adjust width automatically based on height */
  box-sizing: content-box;
}

.image-container img.big {
  max-height: 350px; /* Set the desired height */
}


.image-container {
  display: flex; /* Aligns images side by side */
  justify-content: space-around; /* Space them evenly */
  align-items: center; /* Align them vertically */
  gap: .5rem
}

  .container {
    width: 90%;
    margin: 0 auto;
  }

  .text-center {
    text-align: center;
  }

  .score-amount {
margin: 0;
margin-top: 10px;
  }

  .score-percentage {Score: 
    font-size: 12px;
    font-weight: semi-bold;
  }
  
</style>

# Rapidata OpenAI 4o Preference

<a href="https://www.rapidata.ai">
<img src="https://cdn-uploads.huggingface.co/production/uploads/66f5624c42b853e73e0738eb/jfxR79bOztqaC6_yNNnGU.jpeg" width="400" alt="Dataset visualization">
</a>

This T2I dataset contains over 200'000 human responses from over ~45,000 individual annotators, collected in less than half a day using the [Rapidata Python API](https://docs.rapidata.ai), accessible to anyone and ideal for large scale evaluation.
Evaluating OpenAI 4o (version from 26.3.2025) across three categories: preference, coherence, and alignment.

Explore our latest model rankings on our [website](https://www.rapidata.ai/benchmark).

If you get value from this dataset and would like to see more in the future, please consider liking it ❤️

## Overview

The evaluation consists of 1v1 comparisons between OpenAI 4o (version from 26.3.2025) and 12 other models: Ideogram V2, Recraft V2, Lumina-15-2-25, Frames-23-1-25, Imagen-3, Flux-1.1-pro, Flux-1-pro, DALL-E 3, Midjourney-5.2, Stable Diffusion 3, Aurora, and Janus-7b.

Below, you'll find key visualizations that highlight how these models compare in terms of prompt alignment and coherence, where OpenAI 4o (version from 26.3.2025) significantly outperforms the other models.

<div style="width: 100%; display: flex; justify-content: center; align-items: center; gap: 20px;">
    <div style="width: 90%; max-width: 1000px;">
        <img src="https://cdn-uploads.huggingface.co/production/uploads/664dcc6296d813a7e15e170e/fMf6_uredbYDY7Hzuyk9J.png" style="width: 95%; height: auto; display: block;">
    </div>
    <div style="width: 90%; max-width: 1000px;">
        <img src="https://cdn-uploads.huggingface.co/production/uploads/664dcc6296d813a7e15e170e/rMjvWjG8HFql65D47TGsZ.png" style="width: 100%; height: auto; display: block;">
    </div>
</div>

## Master of Absurd Prompts
The benchmark intentially includes a range of absurd or conflicting prompts that aim to target situations or scenes that are very unlikely to occur in the training data 
such as *'A Chair on a cat'* or *'Car is bigger than the airplane.'*. Most other models struggle to adhere to these prompts consistently, but the 4o image generation model 
appears to be significantly ahead of the competition in this regard.

<div class="horizontal-container">
  <div clas="container">
   <div class="text-center">
      <q>A chair on a cat.</q>
    </div>
    <div class="image-container">
      <div>
        <h3 class="score-amount">OpenAI 4o</h3>
        <img class="big" src="https://cdn-uploads.huggingface.co/production/uploads/6710d82fd3a72fc574ea620f/kS2uE91Q3QAKxR205DxS_.webp" width=300>
      </div>
      <div>
        <h3 class="score-amount">Imagen 3 </h3>
        <img class="big" src="https://cdn-uploads.huggingface.co/production/uploads/6710d82fd3a72fc574ea620f/KKQRsy9xzJVs7QsYyhuzp.jpeg" width=300>
      </div>
    </div>
  </div>
  <div clas="container">
   <div class="text-center">
      <q>Car is bigger than the airplane.</q>
    </div>
    <div class="image-container">
      <div>
        <h3 class="score-amount">OpenAI 4o</h3>
        <img class="big" src="https://cdn-uploads.huggingface.co/production/uploads/6710d82fd3a72fc574ea620f/TWSsbPFxVJgaHW0gVCR2a.webp" width=300>
      </div>
      <div>
        <h3 class="score-amount">Flux1.1-pro</h3>
        <img class="big" src="https://cdn-uploads.huggingface.co/production/uploads/6710d82fd3a72fc574ea620f/7w3Ls8a6PmuR1ZR1J72Zk.jpeg" width=300>
      </div>
    </div>
  </div>
</div>

That being said, some of the 'absurd' prompts are still not fully solved.

<div class="horizontal-container">
  <div clas="container">
   <div class="text-center">
      <q>A fish eating a pelican.</q>
    </div>
    <div class="image-container">
      <div>
        <h3 class="score-amount">OpenAI 4o</h3>
        <img class="big" src="https://cdn-uploads.huggingface.co/production/uploads/6710d82fd3a72fc574ea620f/xsJ2E_0Kx5gJjIO6C29-Q.webp" width=300>
      </div>
      <div>
        <h3 class="score-amount">Recraft V2</h3>
        <img class="big" src="https://cdn-uploads.huggingface.co/production/uploads/6710d82fd3a72fc574ea620f/R7Zf5dmhvjUTgkBMEx9Ns.webp" width=300>
      </div>
    </div>
  </div>
  <div clas="container">
   <div class="text-center">
      <q>A horse riding an astronaut.</q>
    </div>
    <div class="image-container">
      <div>
        <h3 class="score-amount">OpenAI 4o</h3>
        <img class="big" src="https://cdn-uploads.huggingface.co/production/uploads/6710d82fd3a72fc574ea620f/RretHzxWGlXsjD9gXmg2k.webp" width=300>
      </div>
      <div>
        <h3 class="score-amount">Ideogram</h3>
        <img class="big" src="https://cdn-uploads.huggingface.co/production/uploads/6710d82fd3a72fc574ea620f/qtVKN58c0JgCYKK2Xc5bT.png" width=300>
      </div>
    </div>
  </div>
</div>


## Alignment

The alignment score quantifies how well an video matches its prompt. Users were asked: "Which image matches the description better?".

<div class="vertical-container">
  <div class="container">
   <div class="text-center">
      <q>A baseball player in a blue and white uniform is next to a player in black and white .</q>
    </div>
    <div class="image-container">
      <div>
        <h3 class="score-amount">OpenAI 4o</h3>
        <div class="score-percentage">Score: 100%</div>
        <img style="border: 5px solid #18c54f;" src="https://cdn-uploads.huggingface.co/production/uploads/664dcc6296d813a7e15e170e/pzKcqdCXwVDZi5lwgoeGv.jpeg" width=500>
      </div>
      <div>
        <h3 class="score-amount">Stable Diffusion 3 </h3>
        <div class="score-percentage">Score: 0%</div>
        <img src="https://cdn-uploads.huggingface.co/production/uploads/664dcc6296d813a7e15e170e/rbxFhkeir8TUTK-vYDn6Q.jpeg" width=500>
      </div>
    </div>
  </div>

  <div class="container">
   <div class="text-center">
      <q>A couple of glasses are sitting on a table.</q>
    </div>
    <div class="image-container">
      <div>
        <h3 class="score-amount">OpenAI 4o</h3>
        <div class="score-percentage">Score: 2.8%</div>
        <img  src="https://cdn-uploads.huggingface.co/production/uploads/664dcc6296d813a7e15e170e/AY-I6WqgUF4Eh3thLkAqJ.jpeg" width=500>
      </div>
      <div>
        <h3 class="score-amount">Dalle-3</h3>
        <div class="score-percentage">Score: 97.2%</div>
        <img style="border: 5px solid #18c54f;" src="https://cdn-uploads.huggingface.co/production/uploads/664dcc6296d813a7e15e170e/3ygGq2P4dS6rfh5q-x3jb.jpeg" width=500>
      </div>
    </div>
  </div>
</div>

## Coherence

The coherence score measures whether the generated video is logically consistent and free from artifacts or visual glitches. Without seeing the original prompt, users were asked: "Which image has **more** glitches and is **more** likely to be AI generated?"

<div class="vertical-container">
  <div class="container">
    <div class="image-container">
      <div>
        <h3 class="score-amount">OpenAI 4o </h3>
        <div class="score-percentage">Glitch Rating: 0%</div>
        <img style="border: 5px solid #18c54f;" src="https://cdn-uploads.huggingface.co/production/uploads/664dcc6296d813a7e15e170e/DzuAiklD3R_pwe-yFtRM7.jpeg" width=500>
      </div>
      <div>
        <h3 class="score-amount">Lumina-15-2-25 </h3>
        <div class="score-percentage">Glitch Rating: 100%</div>
        <img src="https://cdn-uploads.huggingface.co/production/uploads/664dcc6296d813a7e15e170e/iAn4zphOEL_cpOorp0JNZ.jpeg" width=500>
      </div>
    </div>
  </div>

   <div class="container">
    <div class="image-container">
      <div>
        <h3 class="score-amount">OpenAI 4o </h3>
        <div class="score-percentage">Glitch Rating: 98.6%</div>
        <img src="https://cdn-uploads.huggingface.co/production/uploads/664dcc6296d813a7e15e170e/IeJHwzNc77tjVAKf8nGEk.jpeg" width=500>
      </div>
      <div>
        <h3 class="score-amount">Recraft V2</h3>
        <div class="score-percentage">Glitch Rating: 1.4%</div>
        <img style="border: 5px solid #18c54f;" src="https://cdn-uploads.huggingface.co/production/uploads/664dcc6296d813a7e15e170e/iCuVaPrVGbDeLHuqbMgkc.jpeg" width=500>
      </div>
    </div>
  </div>   
</div>

## Preference

The preference score reflects how visually appealing participants found each image, independent of the prompt. Users were asked: "Which image do you prefer?"

<div class="vertical-container">
  <div class="container">
    <div class="image-container">
      <div>
        <h3 class="score-amount">OpenAI 4o</h3>
        <div class="score-percentage">Score: 100%</div>
        <img style="border: 5px solid #18c54f;" src="https://cdn-uploads.huggingface.co/production/uploads/664dcc6296d813a7e15e170e/ve4DVzU0kZznjA9N0AdkO.jpeg" width=500>
      </div>
      <div>
        <h3 class="score-amount">Lumina-15-2-25</h3>
        <div class="score-percentage">Score: 0%</div>
        <img src="https://cdn-uploads.huggingface.co/production/uploads/664dcc6296d813a7e15e170e/zTZRillcEV85C9gfLa25L.jpeg" width=500>
      </div>
    </div>
  </div>

   <div class="container">
    <div class="image-container">
      <div>
        <h3 class="score-amount">OpenAI 4o </h3>
        <div class="score-percentage">Score: 0%</div>
        <img  src="https://cdn-uploads.huggingface.co/production/uploads/664dcc6296d813a7e15e170e/0EmcYSDQeseS1XSWyG-lb.jpeg" width=500>
      </div>
      <div>
        <h3 class="score-amount">Flux-1.1 Pro </h3>
        <div class="score-percentage">Score: 100%</div>
        <img style="border: 5px solid #18c54f;" src="https://cdn-uploads.huggingface.co/production/uploads/664dcc6296d813a7e15e170e/MO7RnVUWC0gR84PIKDuyI.jpeg" width=500>
      </div>
    </div>
  </div>
</div>

## About Rapidata

Rapidata's technology makes collecting human feedback at scale faster and more accessible than ever before. Visit [rapidata.ai](https://www.rapidata.ai/) to learn more about how we're revolutionizing human feedback collection for AI development.
