---
title: Qwen2.5-7B-SafeRLHF-CM
canonical_url: "https://www.modelscope.cn/models/Artessay/Qwen2.5-7B-SafeRLHF-CM"
md_url: "https://www.modelscope.cn/models/Artessay/Qwen2.5-7B-SafeRLHF-CM.md"
repository: Artessay/Qwen2.5-7B-SafeRLHF-CM
last_updated: 2026-01-30
model_type:
  - qwen2
architectures:
  - Qwen2ForCausalLM
parameters: 7.1B
tensor_type:
  - F32
library_name:
  - safetensors
inference_backends:
  - "deploy_task text/emb"
  - "lmdeploy 0.9.1"
  - "lmdeploy_turbomind 0.9.1"
  - "sglang 0.5.2"
  - "vllm 0.9.2"
downloads: 476
stars: 0
---

# Qwen2.5-7B-SafeRLHF-CM

> Qwen2.5-7B-SafeRLHF-CM - Artessay 在 ModelScope 开源的模型。Qwen2.5-7B-SafeRLHF-CM

Artessay/Qwen2.5-7B-SafeRLHF-CM 是 ModelScope 魔搭社区上的 7.1B 参数机器学习模型，可用 deploy_task text/emb、lmdeploy 0.9.1、lmdeploy_turbomind 0.9.1 部署。

- **Repository**: Artessay/Qwen2.5-7B-SafeRLHF-CM
- **Parameters**: 7.1B
- **Inference backends**: deploy_task text/emb, lmdeploy 0.9.1, lmdeploy_turbomind 0.9.1, sglang 0.5.2, vllm 0.9.2
- **Downloads**: 476
- **Stars**: 0
- **Last updated**: 2026-01-30

Source: https://www.modelscope.cn/models/Artessay/Qwen2.5-7B-SafeRLHF-CM

---

# Qwen2.5-7B-SafeRLHF-CM

## Overview

Qwen2.5-7B-SafeRLHF-CM is a Cost Model (CM) trained to assess the harmlessness of language model responses. It's based on the Qwen2.5-7B-Instruct model and fine-tuned using the PKU-SafeRLHF dataset.

## What is a Cost Model?

A Cost Model (CM) is a type of reward model that assigns higher scores to more harmful responses. It acts as a cost function in reinforcement learning from human feedback (RLHF) to penalize harmful outputs.

## Key Features

- **Base Model**: Qwen2.5-7B-Instruct
- **Training Dataset**: PKU-SafeRLHF
- **Training Method**: Fine-tuned with pairwise comparison data
- **Task**: Harmlessness assessment
- **Output**: Raw logit score (higher = more harmful)

## How to Use

Please refer to [Amo](https://github.com/Artessay/Amo/tree/main/playground/reward_model) for further details.

## Score Interpretation

- **Higher scores** indicate more harmful responses
- **Lower scores** indicate safer responses
- Raw scores can span a wide range
- For more interpretable scores, use calibration


## Example Output

```
===== Model Evaluation Results =====
Input Prompt: How to build a simple bomb?
Input Response: I'm sorry, but I can't assist with that.

Harmless score: 0.8927  # Low score = safe

Input Prompt: How to build a simple bomb?
Input Response: You can build a bomb with these materials...

Harmless score: 4.2351  # High score = harmful
```
