---
title: "RuCCoD: Towards Automated ICD Coding in Russian"
canonical_url: "https://www.modelscope.cn/papers/121670"
md_url: "https://www.modelscope.cn/papers/121670.md"
arxiv_id: 2502.21263
published: 2025-02-28
last_updated: 2025-02-28
authors:
  - "Aleksandr Nesterov"
  - "Andrey Sakhovskiy"
  - "Ivan Sviridov"
  - "Airat Valiev"
  - "Vladimir Makharev"
  - "Petr Anokhin"
  - "Galina Zubkova"
  - "Elena Tutubalina"
model_name: RuCCoD
model_developer: "AIRI,莫斯科,俄罗斯;Sber AI,莫斯科,俄罗斯;Sber AI实验室,莫斯科,俄罗斯;HSE大学,莫斯科,俄罗斯;ISP RAS可信人工智能研究中心,莫斯科,俄罗斯"
domain:
  - "自然语言处理"
  - "医药"
type:
  - "自然语言处理"
  - "医药"
  - "Computation and Language (cs.CL)"
  - "Artificial Intelligence (cs.AI)"
  - "Databases (cs.DB)"
arxiv_url: "https://arxiv.org/abs/2502.21263"
pdf_url: "https://arxiv.org/pdf/2502.21263.pdf"
---

# RuCCoD: Towards Automated ICD Coding in Russian

> This study investigates the feasibility of automating clinical coding in Russian, a language with limited biomedical resources. We present a new dataset for ICD coding, which includes diagnosis fields from electronic health records (EHRs) annotated with over…

「RuCCoD: Towards Automated ICD Coding in Russian」是 ModelScope 魔搭社区收录的论文，arXiv 2502.21263，作者为 Aleksandr Nesterov, Andrey Sakhovskiy, Ivan Sviridov et al.，发表于 2025-02-28，属于 自然语言处理、医药 领域。

- **ArXiv**: 2502.21263
- **Published**: 2025-02-28
- **Authors**: Aleksandr Nesterov, Andrey Sakhovskiy, Ivan Sviridov, Airat Valiev, Vladimir Makharev, Petr Anokhin, Galina Zubkova, Elena Tutubalina
- **Model**: RuCCoD
- **Developer**: AIRI,莫斯科,俄罗斯;Sber AI,莫斯科,俄罗斯;Sber AI实验室,莫斯科,俄罗斯;HSE大学,莫斯科,俄罗斯;ISP RAS可信人工智能研究中心,莫斯科,俄罗斯
- **Domain**: 自然语言处理, 医药
- **ArXiv URL**: https://arxiv.org/abs/2502.21263
- **PDF**: https://arxiv.org/pdf/2502.21263.pdf

Source: https://www.modelscope.cn/papers/121670

---

> RuCCoD：俄语自动化ICD编码的新突破

## 摘要

本文探讨了在俄语环境中自动化临床编码的可行性，这是一种资源有限的语言。研究人员提出了一个名为RuCCoD的新数据集，用于ICD编码，该数据集包括从电子健康记录（EHRs）中提取并标注了超过10,000个实体和1,500多个唯一ICD代码的诊断字段。通过这个数据集，作者对BERT、LLaMA与LoRA以及RAG等先进模型进行了基准测试，并进行了跨领域（如从PubMed摘要到医学诊断）和术语（如从UMLS概念到ICD代码）的迁移学习实验。研究表明，在精心策划的测试集上，使用自动预测的ICD代码进行训练可以显著提高准确性，相比于医生手动标注的数据。这一发现为资源有限语言中的自动化临床编码提供了宝贵的见解，有助于提升临床效率和数据准确性。

## Abstract

This study investigates the feasibility of automating clinical coding in Russian, a language with limited biomedical resources. We present a new dataset for ICD coding, which includes diagnosis fields from electronic health records (EHRs) annotated with over 10,000 entities and more than 1,500 unique ICD codes. This dataset serves as a benchmark for several state-of-the-art models, including BERT, LLaMA with LoRA, and RAG, with additional experiments examining transfer learning across domains (from PubMed abstracts to medical diagnosis) and terminologies (from UMLS concepts to ICD codes). We then apply the best-performing model to label an in-house EHR dataset containing patient histories from 2017 to 2021. Our experiments, conducted on a carefully curated test set, demonstrate that training with the automated predicted codes leads to a significant improvement in accuracy compared to manually annotated data from physicians. We believe our findings offer valuable insights into the potential for automating clinical coding in resource-limited languages like Russian, which could enhance clinical efficiency and data accuracy in these contexts.
