---
title: autobencher-qa-33k
canonical_url: "https://www.modelscope.cn/datasets/allenai/autobencher-qa-33k"
md_url: "https://www.modelscope.cn/datasets/allenai/autobencher-qa-33k.md"
repository: allenai/autobencher-qa-33k
last_updated: 2025-08-25
license: cc-by-4.0
storage_size: "5.1 MB"
downloads: 128
stars: 0
---

# autobencher-qa-33k

> autobencher-qa-33k - allenai 在 ModelScope 开源的数据集。These are 33K questions generated using Autobencher. The questions come from randomly sampled Wikipedia articles, which are further filtered and transformed into questions by GPT-4o.

allenai/autobencher-qa-33k 是 ModelScope 魔搭社区上的数据集，存储大小 5.1 MB，采用 cc-by-4.0 许可。

- **Repository**: allenai/autobencher-qa-33k
- **License**: cc-by-4.0
- **Storage size**: 5.1 MB
- **Downloads**: 128
- **Stars**: 0
- **Last updated**: 2025-08-25

Source: https://www.modelscope.cn/datasets/allenai/autobencher-qa-33k

---

These are 33K questions generated using [Autobencher](https://arxiv.org/abs/2407.08351). The questions come from randomly sampled Wikipedia articles, which are further filtered and transformed into questions by GPT-4o.

This benchmark is used in the [signal and noise](https://huggingface.co/datasets/allenai/signal-and-noise) project to demonstrate the impact of a large sample size on the modeling noise of a benchmark.

### Citation

Please cite the original authors of Autobencher, and our work which generated this particular evaluation set:

```
@article{li2024autobencher,
  title={Autobencher: Towards declarative benchmark construction},
  author={Li, Xiang Lisa and Kaiyom, Farzaan and Liu, Evan Zheran and Mai, Yifan and Liang, Percy and Hashimoto, Tatsunori},
  journal={arXiv preprint arXiv:2407.08351},
  year={2024}
}
```

```
@article{heineman2025signal,
  title={Signal and Noise: A Framework for Reducing Uncertainty in Language Model Evaluation},
  author={Heineman, David and Hofmann, Valentin and Magnusson, Ian and Gu, Yuling and Smith, Noah A and Hajishirzi, Hannaneh and Lo, Kyle and Dodge, Jesse},
  journal={arXiv preprint arXiv:2508.13144},
  year={2025}
}
```

### Dataset Description

- **Developed by:** Allen Institute for AI (Ai2)
- **Language(s) (NLP):** English
- **License:** This dataset contains model outputs generated from GPT-4o, which is subject to OpenAI's [Terms of Use](https://openai.com/policies/row-terms-of-use/). This dataset is licensed under CC BY 4.0. It is intended for research and educational use in accordance with Ai2's [Responsible Use Guidelines](https://allenai.org/responsible-use)
- **Contact:** Technical inquiries: `davidh@allenai.org`. Press: `press@allenai.org`
