---
title: MMRA
canonical_url: "https://www.modelscope.cn/datasets/m-a-p/MMRA"
md_url: "https://www.modelscope.cn/datasets/m-a-p/MMRA.md"
repository: m-a-p/MMRA
chinese_name: MMRA
last_updated: 2024-11-21
license: "Apache License 2.0"
storage_size: "545 MB"
downloads: 2761
stars: 0
---

# MMRA

> MMRA - m-a-p 在 ModelScope 开源的数据集。We define a multi-image relation association task, and meticulously curate MMRA benchmark, a Multi-granularity Multi-image Relational Association benchmark, consisted of 1,024 samples. In order to systematically and…

m-a-p/MMRA 是 ModelScope 魔搭社区上的数据集，存储大小 545 MB，采用 Apache License 2.0 许可。

- **Repository**: m-a-p/MMRA
- **License**: Apache License 2.0
- **Storage size**: 545 MB
- **Downloads**: 2761
- **Stars**: 0
- **Last updated**: 2024-11-21

Source: https://www.modelscope.cn/datasets/m-a-p/MMRA

---

# Introduction

We define a multi-image relation association task, and meticulously curate **MMRA** benchmark, a **M**ulti-granularity **M**ulti-image **R**elational **A**ssociation benchmark, consisted of **1,024** samples.
In order to systematically and comprehensively evaluate mainstream LVLMs, we establish an associational relation system among images that contain **11 subtasks** (e.g, UsageSimilarity, SubEvent, etc.) at two granularity levels (i.e., "**image**" and "**entity**") according to the relations in ConceptNet.
Our experiments reveal that on the MMRA benchmark, current multi-image LVLMs exhibit distinct advantages and disadvantages across various subtasks. Notably, fine-grained, entity-level multi-image perception tasks pose a greater challenge for LVLMs compared to image-level tasks. Tasks that involve spatial perception are especially difficult for LVLMs to handle.
Additionally, our findings indicate that while LVLMs demonstrate a strong capability to perceive image details, enhancing their ability to associate information across multiple images hinges on improving the reasoning capabilities of their language model component.
Moreover, we explored the ability of LVLMs to perceive image sequences within the context of our multi-image association task. Our experiments indicate that the majority of current LVLMs do not adequately model image sequences during the pre-training process.

![framework](./imgs/framework.png)

![main_result](./imgs/main_result.png)


---
# Evaluateion Codes

The codes of this paper can be found in our [GitHub](https://github.com/Wusiwei0410/MMRA/tree/main)



---

# Using Datasets

You can load our datasets by following codes:

```python
MMRA_data = datasets.load_dataset('m-a-p/MMRA')['train']
print(MMRA_data[0])
```

---
# Citation

BibTeX:
```
@article{wu2024mmra,
  title={MMRA: A Benchmark for Multi-granularity Multi-image Relational Association},
  author={Wu, Siwei and Zhu, Kang and Bai, Yu and Liang, Yiming and Li, Yizhi and Wu, Haoning and Liu, Jiaheng and Liu, Ruibo and Qu, Xingwei and Cheng, Xuxin and others},
  journal={arXiv preprint arXiv:2407.17379},
  year={2024}
}
```
