---
title: Qwen3-8B-LoRA
canonical_url: "https://www.modelscope.cn/models/mc36473/Qwen3-8B-LoRA"
md_url: "https://www.modelscope.cn/models/mc36473/Qwen3-8B-LoRA.md"
repository: mc36473/Qwen3-8B-LoRA
last_updated: 2025-09-27
license: other
pipeline_tag: text-generation
tasks:
  - text-generation
model_type:
  - qwen3
architectures:
  - Qwen3ForCausalLM
base_model:
  - /root/Qwen3-8B
base_model_relation: adapter
parameters: 8.2B
tensor_type:
  - BF16
library_name:
  - lora
  - safetensors
  - pytorch
frameworks:
  - Pytorch
inference_backends:
  - "deploy_task text/emb"
  - "lmdeploy_turbomind 0.9.1"
  - "sglang 0.5.2"
  - "vllm 0.9.2"
downloads: 18
stars: 0
tags:
  - llama-factory
  - lora
  - generated_from_trainer
---

# Qwen3-8B-LoRA

> Qwen3-8B-LoRA - mc36473 在 ModelScope 开源的模型。This model is a fine-tuned version of /root/Qwen3-8B on the stancetasksharegpt dataset.

mc36473/Qwen3-8B-LoRA 是 ModelScope 魔搭社区上的 8.2B 参数text-generation模型，采用 other 许可，基于 /root/Qwen3-8B 构建，可用 deploy_task text/emb、lmdeploy_turbomind 0.9.1、sglang 0.5.2 部署。

- **Repository**: mc36473/Qwen3-8B-LoRA
- **License**: other
- **Tasks**: text-generation
- **Parameters**: 8.2B
- **Base model**: /root/Qwen3-8B
- **Inference backends**: deploy_task text/emb, lmdeploy_turbomind 0.9.1, sglang 0.5.2, vllm 0.9.2
- **Tags**: llama-factory, lora, generated_from_trainer
- **Downloads**: 18
- **Stars**: 0
- **Last updated**: 2025-09-27

Source: https://www.modelscope.cn/models/mc36473/Qwen3-8B-LoRA

---

<!-- This model card has been generated automatically according to the information the Trainer had access to. You
should probably proofread and complete it, then remove this comment. -->

# sft2

This model is a fine-tuned version of [/root/Qwen3-8B](https://huggingface.co//root/Qwen3-8B) on the stance_task_sharegpt dataset.

## Model description

More information needed

## Intended uses & limitations

More information needed

## Training and evaluation data

More information needed

## Training procedure

### Training hyperparameters

The following hyperparameters were used during training:
- learning_rate: 0.0001
- train_batch_size: 1
- eval_batch_size: 8
- seed: 42
- distributed_type: multi-GPU
- num_devices: 2
- gradient_accumulation_steps: 8
- total_train_batch_size: 16
- total_eval_batch_size: 16
- optimizer: Use OptimizerNames.ADAMW_TORCH with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
- lr_scheduler_type: cosine
- lr_scheduler_warmup_ratio: 0.1
- num_epochs: 3.0

### Training results



### Framework versions

- PEFT 0.15.2
- Transformers 4.55.0
- Pytorch 2.5.1+cu124
- Datasets 3.6.0
- Tokenizers 0.21.1
