---
title: Qwen2.5-7B-SafeRLHF-RM
canonical_url: "https://www.modelscope.cn/models/Artessay/Qwen2.5-7B-SafeRLHF-RM"
md_url: "https://www.modelscope.cn/models/Artessay/Qwen2.5-7B-SafeRLHF-RM.md"
repository: Artessay/Qwen2.5-7B-SafeRLHF-RM
last_updated: 2026-01-30
pipeline_tag: text-generation
tasks:
  - text-generation
model_type:
  - qwen2
architectures:
  - Qwen2ForCausalLM
parameters: 7.1B
tensor_type:
  - F32
library_name:
  - safetensors
  - pytorch
frameworks:
  - Pytorch
inference_backends:
  - "deploy_task text/emb"
  - "lmdeploy 0.9.1"
  - "lmdeploy_turbomind 0.9.1"
  - "sglang 0.5.2"
  - "vllm 0.9.2"
downloads: 485
stars: 1
---

# Qwen2.5-7B-SafeRLHF-RM

> Qwen2.5-7B-SafeRLHF-RM - Artessay 在 ModelScope 开源的模型。Qwen2.5-7B-SafeRLHF-RM

Artessay/Qwen2.5-7B-SafeRLHF-RM 是 ModelScope 魔搭社区上的 7.1B 参数text-generation模型，可用 deploy_task text/emb、lmdeploy 0.9.1、lmdeploy_turbomind 0.9.1 部署。

- **Repository**: Artessay/Qwen2.5-7B-SafeRLHF-RM
- **Tasks**: text-generation
- **Parameters**: 7.1B
- **Inference backends**: deploy_task text/emb, lmdeploy 0.9.1, lmdeploy_turbomind 0.9.1, sglang 0.5.2, vllm 0.9.2
- **Downloads**: 485
- **Stars**: 1
- **Last updated**: 2026-01-30

Source: https://www.modelscope.cn/models/Artessay/Qwen2.5-7B-SafeRLHF-RM

---

# Qwen2.5-7B-SafeRLHF-RM

## Overview

Qwen2.5-7B-SafeRLHF-RM is a Reward Model (RM) trained to assess the helpfulness of language model responses. It's based on the Qwen2.5-7B-Instruct model and fine-tuned using the PKU-SafeRLHF dataset.

## What is a Reward Model?

A Reward Model (RM) is a type of model that assigns scores to text based on human preferences. This particular RM assigns higher scores to more helpful responses and lower scores to less helpful ones.

## Key Features

- **Base Model**: Qwen2.5-7B-Instruct
- **Training Dataset**: PKU-SafeRLHF
- **Training Method**: Fine-tuned with pairwise comparison data
- **Task**: Helpfulness assessment
- **Output**: Raw logit score (higher = more helpful)

## How to Use

Please refer to [Amo](https://github.com/Artessay/Amo/tree/main/playground/reward_model) for further details.

## Score Interpretation

- **Higher scores** indicate more helpful responses
- **Lower scores** indicate less helpful responses
- Raw scores can span a wide range
- For more interpretable scores, use calibration


## Example Output

```
===== Model Evaluation Results =====
Input Prompt: How to make a cake?
Input Response: To make a cake, you'll need flour, sugar, eggs, and butter. Mix them together and bake at 350°F for 30 minutes.

Helpful score: 4.5678  # High score = helpful

Input Prompt: How to make a cake?
Input Response: I don't know.

Helpful score: 1.2345  # Low score = less helpful
```
