---
title: qwen3_8b_sft_lora
canonical_url: "https://www.modelscope.cn/models/mc36473/qwen3_8b_sft_lora"
md_url: "https://www.modelscope.cn/models/mc36473/qwen3_8b_sft_lora.md"
repository: mc36473/qwen3_8b_sft_lora
chinese_name: qwen3_lora
last_updated: 2025-09-13
license: other
pipeline_tag: text-generation
tasks:
  - text-generation
base_model:
  - /root/Qwen3-8B
base_model_relation: adapter
parameters: 21.8M
tensor_type:
  - BF16
library_name:
  - lora
  - safetensors
  - pytorch
frameworks:
  - Pytorch
downloads: 46
stars: 0
tags:
  - llama-factory
  - lora
  - generated_from_trainer
---

# qwen3_8b_sft_lora

> qwen3_8b_sft_lora - mc36473 在 ModelScope 开源的模型。This model is a fine-tuned version of /root/Qwen3-8B on the stancetasksharegpt dataset.

mc36473/qwen3_8b_sft_lora 是 ModelScope 魔搭社区上的 21.8M 参数text-generation模型，采用 other 许可，基于 /root/Qwen3-8B 构建。

- **Repository**: mc36473/qwen3_8b_sft_lora
- **License**: other
- **Tasks**: text-generation
- **Parameters**: 21.8M
- **Base model**: /root/Qwen3-8B
- **Tags**: llama-factory, lora, generated_from_trainer
- **Downloads**: 46
- **Stars**: 0
- **Last updated**: 2025-09-13

Source: https://www.modelscope.cn/models/mc36473/qwen3_8b_sft_lora

---

<!-- This model card has been generated automatically according to the information the Trainer had access to. You
should probably proofread and complete it, then remove this comment. -->

# sft2

This model is a fine-tuned version of [/root/Qwen3-8B](https://huggingface.co//root/Qwen3-8B) on the stance_task_sharegpt dataset.

## Model description

More information needed

## Intended uses & limitations

More information needed

## Training and evaluation data

More information needed

## Training procedure

### Training hyperparameters

The following hyperparameters were used during training:
- learning_rate: 0.0001
- train_batch_size: 1
- eval_batch_size: 8
- seed: 42
- distributed_type: multi-GPU
- num_devices: 2
- gradient_accumulation_steps: 8
- total_train_batch_size: 16
- total_eval_batch_size: 16
- optimizer: Use OptimizerNames.ADAMW_TORCH with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
- lr_scheduler_type: cosine
- lr_scheduler_warmup_ratio: 0.1
- num_epochs: 3.0

### Training results



### Framework versions

- PEFT 0.15.2
- Transformers 4.55.0
- Pytorch 2.5.1+cu124
- Datasets 3.6.0
- Tokenizers 0.21.1
