---
title: VideoMathQA
canonical_url: "https://www.modelscope.cn/datasets/MBZUAI/VideoMathQA"
md_url: "https://www.modelscope.cn/datasets/MBZUAI/VideoMathQA.md"
repository: MBZUAI/VideoMathQA
last_updated: 2025-06-06
license: "Apache License 2.0"
storage_size: "14 GB"
downloads: 1232
stars: 0
---

# VideoMathQA

> VideoMathQA - MBZUAI 在 ModelScope 开源的数据集。VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos

MBZUAI/VideoMathQA 是 ModelScope 魔搭社区上的数据集，存储大小 14 GB，采用 Apache License 2.0 许可。

- **Repository**: MBZUAI/VideoMathQA
- **License**: Apache License 2.0
- **Storage size**: 14 GB
- **Downloads**: 1232
- **Stars**: 0
- **Last updated**: 2025-06-06

Source: https://www.modelscope.cn/datasets/MBZUAI/VideoMathQA

---

# VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos

[![Paper](https://img.shields.io/badge/📄_arXiv-Paper-blue)](https://arxiv.org/abs/2506.05349)
[![Website](https://img.shields.io/badge/🌐_Project-Website-87CEEB)](https://mbzuai-oryx.github.io/VideoMathQA)
[![🏅 Leaderboard (Reasoning)](https://img.shields.io/badge/🏅_Leaderboard-Reasoning-red)](https://hanoonar.github.io/VideoMathQA/#leaderboard-2)
[![🏅 Leaderboard (Direct)](https://img.shields.io/badge/🏅_Leaderboard-Direct-yellow)](https://hanoonar.github.io/VideoMathQA/#leaderboard)
[![📊 Eval (LMMs-Eval)](https://img.shields.io/badge/📊_Eval-LMMs--Eval-orange)](https://github.com/EvolvingLMMs-Lab/lmms-eval/tree/main/lmms_eval/tasks/videomathqa)


## 📣 Announcement

Note that the Official evaluation for **VideoMathQA** is supported in the [`lmms-eval`](https://github.com/EvolvingLMMs-Lab/lmms-eval/tree/main/lmms_eval/tasks/videomathqa) framework. Please use the GitHub repository [`mbzuai-oryx/VideoMathQA`](https://github.com/mbzuai-oryx/VideoMathQA) to create or track any issues related to VideoMathQA that you may encounter.

---

## 💡 VideoMathQA

**VideoMathQA** is a benchmark designed to evaluate mathematical reasoning in real-world educational videos. It requires models to interpret and integrate information from **three modalities**, visuals, audio, and text, across time. The benchmark tackles the **needle-in-a-multimodal-haystack** problem, where key information is sparse and spread across different modalities and moments in the video.

<p align="center">
  <img src="images/intro_fig.png" alt="Highlight Figure"><br>
  <em>The foundation of our benchmark is the needle-in-a-multimodal-haystack challenge, capturing the core difficulty of cross-modal reasoning across time from visual, textual, and audio streams. Built on this, VideoMathQA categorizes each question along four key dimensions: reasoning type, mathematical concept, video duration, and difficulty.</em>
</p>

---
## 🔥 Highlights

- **Multimodal Reasoning Benchmark:** VideoMathQA introduces a challenging **needle-in-a-multimodal-haystack** setup where models must reason across **visuals, text and audio**. Key information is **sparsely distributed across modalities and time**, requiring strong performance in fine-grained visual understanding, multimodal integration, and reasoning.

- **Three Types of Reasoning:** Questions are categorized into: **Problem Focused**, where the question is explicitly stated and solvable via direct observation and reasoning from the video; **Concept Transfer**, where a demonstrated method or principle is adapted to a newly posed problem; **Deep Instructional Comprehension**, which requires understanding long-form instructional content, interpreting partially worked-out steps, and completing the solution.

- **Diverse Evaluation Dimensions:** Each question is evaluated across four axes, which captures diversity in content, length, complexity, and reasoning depth.
   **mathematic concepts**, 10 domains such as geometry, statistics, arithmetics and charts; **video duration** ranging from 10s to 1 hour long categorized as short, medium, long; **difficulty level**; and **reasoning type**.  

- **High-Quality Human Annotations:** The benchmark includes **420 expert-curated questions**, each with five answer choices, a correct answer, and detailed **chain-of-thought (CoT) steps**. Over **2,945 reasoning steps** have been manually written, reflecting **920+ hours** of expert annotation effort with rigorous quality control.


## 🔍 Examples from the Benchmark
We present example questions from <strong>VideoMathQA</strong> illustrating the three reasoning types: Problem Focused, Concept Transfer, and Deep Comprehension. The benchmark includes evolving dynamics in a video, complex text prompts, five multiple-choice options, the expert-annotated step-by-step reasoning to solve the given problem, and the final correct answer as shown above.
<p align="center">
  <img src="images/data_fig.png" alt="Figure 1" width="90%">
</p>

---

## 📈 Overview of VideoMathQA
We illustrate an overview of the <strong>VideoMathQA</strong> benchmark through: <strong>a)</strong>&nbsp;The distribution of questions and model performance across ten mathematical concepts, which highlights a significant gap in the current multimodal models and their ability to perform mathematical reasoning over videos. <strong>b)</strong>&nbsp;The distribution of video durations, spanning from short clips of 10s to long videos up to 1hr. <strong>c)</strong>&nbsp;Our three-stage annotation pipeline performed by expert science graduates, who annotate detailed step-by-step reasoning trails, with strict quality assessment at each stage.
<p align="center">
  <img src="images/stat_fig.png" alt="Figure 2" width="90%">
</p>
