---
title: solyanka
canonical_url: "https://www.modelscope.cn/datasets/ai-forever/solyanka"
md_url: "https://www.modelscope.cn/datasets/ai-forever/solyanka.md"
repository: ai-forever/solyanka
last_updated: 2025-05-26
license: "Apache License 2.0"
storage_size: "6.3 GB"
downloads: 729
stars: 0
---

# solyanka

> solyanka - ai-forever 在 ModelScope 开源的数据集。Dataset card for Solyanka

ai-forever/solyanka 是 ModelScope 魔搭社区上的数据集，存储大小 6.3 GB，采用 Apache License 2.0 许可。

- **Repository**: ai-forever/solyanka
- **License**: Apache License 2.0
- **Storage size**: 6.3 GB
- **Downloads**: 729
- **Stars**: 0
- **Last updated**: 2025-05-26

Source: https://www.modelscope.cn/datasets/ai-forever/solyanka

---

# Dataset card for Solyanka

This is a dataset collection of ~10 million weakly-supervised pairs for training text embedding models. Any dataset in collection can be used in SentenceTransformers with an InfoNCE loss.

## Data processing

The initial pool of pairs were deduplified, filtered by length and quality. Most of documents are less than 512 tokens ([FRIDA](https://huggingface.co/ai-forever/FRIDA) tokenizer). Some pairs were filtered by manual rules (e.g. by post votes, rating, views). We applied consistency filtering with specific N for every datasets (refer to E5 [paper](https://arxiv.org/abs/2212.03533)) to discard low quality pairs.

## Datasets

- 9111_questions_qa ([9111-questions](https://huggingface.co/datasets/nyuuzyou/9111-questions))
- fishkinet_posts ([fishkinet-posts](https://huggingface.co/datasets/nyuuzyou/fishkinet-posts))
- habr_qna_qa ([habr_qna](https://huggingface.co/datasets/its5Q/habr_qna))
- habr_qna_title_body ([habr_qna](https://huggingface.co/datasets/its5Q/habr_qna))
- habr_title_text ([habr](https://huggingface.co/datasets/IlyaGusev/habr))
- mail_ru_qa ([otvetmailru-full](https://www.kaggle.com/datasets/atleast6characterss/otvetmailru-full))
- msmarco_en_ru ([mmarco](https://huggingface.co/datasets/unicamp-dl/mmarco))
- msmarco_ru_en ([mmarco](https://huggingface.co/datasets/unicamp-dl/mmarco))
- msmarco_ru_ru ([mmarco](https://huggingface.co/datasets/unicamp-dl/mmarco))
- pikabu_title_text ([pikabu](https://huggingface.co/datasets/IlyaGusev/pikabu))
- ru_sci_bench ([ru_sci_bench](https://huggingface.co/datasets/mlsa-iai-msu-lab/ru_sci_bench))
- stackoverflow_qa ([ru_stackoverflow](https://huggingface.co/datasets/IlyaGusev/ru_stackoverflow))
- stackoverflow_title_body ([ru_stackoverflow](https://huggingface.co/datasets/IlyaGusev/ru_stackoverflow))
- swim_ir_ru_en ([swim-ir-cross-lingual](https://huggingface.co/datasets/nthakur/swim-ir-cross-lingual))
- buriy ([ru_news](https://huggingface.co/datasets/IlyaGusev/ru_news))
- lenta ([ru_news](https://huggingface.co/datasets/IlyaGusev/ru_news))
- ods_tass ([ru_news](https://huggingface.co/datasets/IlyaGusev/ru_news))
- taiga_fontanka ([ru_news](https://huggingface.co/datasets/IlyaGusev/ru_news))
- telegram_contest ([ru_news](https://huggingface.co/datasets/IlyaGusev/ru_news))
- wikiomnia ([wikiomnia](https://huggingface.co/datasets/RussianNLP/wikiomnia))
- xlsum_summary_text ([xlsum](https://huggingface.co/datasets/csebuetnlp/xlsum))
- xlsum_title_text ([xlsum](https://huggingface.co/datasets/csebuetnlp/xlsum))
- yandex_q_qa ([yandex_q_full](https://huggingface.co/datasets/IlyaGusev/yandex_q_full))
- yandex_q_title_body ([yandex_q_full](https://huggingface.co/datasets/IlyaGusev/yandex_q_full))

## License for the Dataset Collection

This dataset collection is provided under the MIT license, except in cases where a specific dataset has a more restrictive license that may limit the use of the data (e.g., licenses that prohibit commercial use or have other restrictions).

## Terms of Use

1. The user assumes responsibility for checking and complying with the terms of the licenses for each of the datasets, links to which are provided above.
2. Use of this collection is permitted only if the source licenses for the datasets allow such use.
3. In cases where a specific dataset has more restrictive terms, those terms take precedence over the MIT license for this collection.

## Language

Russian is primary language, but some datasets contain English for cross-lingual retrieval experiments.

## Authors

- [SaluteDevices](https://sberdevices.ru/) AI for B2C RnD Team.
- Artem Snegirev: [HF profile](https://huggingface.co/artemsnegirev), [Github](https://github.com/artemsnegirev);
- Anna Maksimova [HF profile](https://huggingface.co/anpalmak);
- Aleksandr Abramov: [HF profile](https://huggingface.co/Andrilko), [Github](https://github.com/Ab1992ao), [Kaggle Competitions Master](https://www.kaggle.com/andrilko)

## Citation

...
