---
title: laya-multilingual-f16.gguf
canonical_url: "https://www.modelscope.cn/models/xzh131421/laya-multilingual-f16.gguf"
md_url: "https://www.modelscope.cn/models/xzh131421/laya-multilingual-f16.gguf.md"
repository: xzh131421/laya-multilingual-f16.gguf
chinese_name: "gguf格式laya-multilingual（多语言版-支持中文），llama.cpp测试通过。"
last_updated: 2026-09-22
pipeline_tag: sentence-embedding
tasks:
  - sentence-embedding
library_name:
  - gguf
downloads: 3419
stars: 5
---

# laya-multilingual-f16.gguf

> laya-multilingual-f16.gguf - xzh131421 在 ModelScope 开源的模型。aya-multilingual（多语言版-支持中文），采用官方 Python 库发布的模型转换格式生成。实测普通笔记本4G核显，llama.cpp-b11012 Vulkan 版启动成功，四语言完整推理跑通，也验证了显式 token 数组不会被偷偷插入 BOS（传 11 个 token 返回 11…

- **Repository**: xzh131421/laya-multilingual-f16.gguf
- **Tasks**: sentence-embedding
- **Downloads**: 3419
- **Stars**: 5
- **Last updated**: 2026-09-22

Source: https://www.modelscope.cn/models/xzh131421/laya-multilingual-f16.gguf

---

# Laya Multilingual 0.3.4 — GGUF (llama.cpp)

[中文](#中文) | [English](#english)

---

# 中文

## 这是什么

`convaiinnovations/laya-multilingual`（laya 0.3.4）的 **GGUF 格式转换**，可被
llama.cpp 加载运行。

Laya 是一个**非自回归的决策模型**（System 1）：给它一个「状态」（文本、邮件、工单、
JSON）和若干「带类型的提问」，它在**一次前向传播**中返回带概率的结构化答案。
不生成文本，所以没有解析环节，也不会幻觉。

> **重要：这不是完整推理所需的一个文件。** 请务必先读下面的「本转换的范围」。

## 本转换的范围（请务必先读）

Laya **不是**一个标准的 HuggingFace 检查点，因此 `convert_hf_to_gguf.py` 无法直接读取它。
它由两部分组成：

| 组成 | 内容 | 在哪里运行 |
|---|---|---|
| `encoder.*` | mmBERT-base 双向编码器：22 层、hidden 768、12 头、FFN 1152（GeGLU）、全注意力与滑窗注意力交替（窗口 128）、RoPE theta 160000、256k 词表 | **llama.cpp**（`modern-bert` 计算图） |
| `head.layers.*` | 额外 2 层 `nn.TransformerEncoderLayer`（d=768、nhead=12、ffn=3072、`norm_first`） | 运行时（PyTorch） |
| `scorer.*`、`act_head.*`、`type_emb.*` | 选项标记打分器、act/escalate 头、问题类型嵌入 | 运行时（PyTorch） |

判断过程是：在每个选项自己的 `[MASK]` 位置上打分，再对该题的选项做 softmax。
**这个决策头不属于 llama.cpp 的 `modern-bert` 计算图，也无法塞进去**（llama.cpp 的加载器
要求文件里的张量数与它计算图里创建的张量数严格一致，多出来的张量会直接导致加载失败）。

所以：**3 亿参数的编码器跑在 llama.cpp 里，决策头跑在配套的 PyTorch 运行时里。**

### 为什么是两个文件

llama.cpp 的加载器会校验张量数量，多一个都不行。所以决策头权重不能和编码器放在同一个
GGUF 里，必须单独一个配套文件。

## 文件清单

本仓库共 6 个文件，使用时**必须放在同一个目录下**：

| 文件 | 大小 | 说明 |
|---|---|---|
| `laya-multilingual-f16.gguf` | 600 MB | `modern-bert` 编码器 + 256k BPE 分词器。**llama.cpp 直接加载这个。** |
| `laya-multilingual-f16-head.gguf` | 57 MB | 决策头权重，由运行时读取 |
| `tokenizer.json` | 33 MB | 运行时用它做分词（用显式 token id 调 llama.cpp） |
| `laya_gguf.py` | 11 KB | 运行时主程序（约 200 行） |
| `gguf_weights.py` | 1 KB | 从决策头 GGUF 读出权重的辅助模块 |
| `README.md` | — | 本文件 |

下载后目录结构应为：

```
laya/
├── laya-multilingual-f16.gguf
├── laya-multilingual-f16-head.gguf
├── tokenizer.json
├── laya_gguf.py
├── gguf_weights.py
└── README.md          (本文件)
```

> **只下载 `laya-multilingual-f16.gguf` 是不够的。** 编码器能加载，但没有任何决策能力 ——
> 决策头、分词器和运行时代码都在其余文件里。

### 运行环境

Python 依赖：`torch`、`numpy`、`tokenizers`、`gguf`（`pip install torch numpy tokenizers gguf`）。
编码器部分由 llama.cpp 承担，不需要 `transformers`。

### 为什么必须三个权重文件

`laya-multilingual-f16.gguf` 里只有编码器。llama.cpp 的加载器会校验张量数量，多一个张量
都会导致加载失败，所以决策头权重无法和编码器放进同一个 GGUF。运行时分词则需要
`tokenizer.json`（它自己分词后把 token id 数组发给 llama.cpp）。

## 快速开始

### 1. 启动 llama.cpp 服务

```bash
llama-server -m laya-multilingual-f16.gguf \
             --embeddings --pooling none \
             --ctx-size 1024 --port 8081
```

> Windows PowerShell 下需要写 `.\llama-server.exe`，PowerShell 默认不从当前目录加载命令。

### 2. 运行

需要一个配套运行时 `laya_gguf.py`（约 200 行 Python，负责分词、调 llama.cpp 取隐状态、
再在 PyTorch 里跑决策头）。三个权重文件与本文件放在同一目录，在该目录下运行：

```python
from pathlib import Path
from laya_gguf import LayaGGUF

HERE = Path(__file__).parent          # 三个权重文件所在目录

agent = LayaGGUF(
    HERE / "laya-multilingual-f16.gguf",
    HERE / "laya-multilingual-f16-head.gguf",
    HERE,                             # tokenizer.json 所在目录
    base_url="http://127.0.0.1:8081",
)

result = agent.predict(
    {"body": "我 3 号已经取消了订阅，但你们还是扣了我 4 月份的 49 美元。请退回我的银行卡。"},
    {
        "department": {
            "type": "choice",
            "instructions": "这个 `body` 应该由哪个团队负责？",
            "criteria": {
                "billing": "发票、扣费、退款",
                "technical": "软件故障和缺陷",
                "sales": "购买和升级",
            },
        },
        "wants_refund": {
            "type": "noul",
            "instructions": "发件人是否在要求退回钱款？",
            "criteria": {"false": "他们没有要求退款", "true": "他们要求把钱退回来"},
        },
    },
)

print(result["answers"]["department"]["choice"])       # billing
print(result["answers"]["wants_refund"]["noul"])       # 0.99
```

### 三种提问类型

| 类型 | 输出 | 适用 |
|---|---|---|
| `choice` | 选项概率分布 + 置信度 | 部门路由、意图分类、主题归类 |
| `score` | 有序量表上的期望等级 | 紧急程度、情绪强度、严重性 |
| `noul` | 校准后的 P(true)，0.0–1.0 | 是否类判断：垃圾邮件、越狱、流失风险 |

## 关于分词器

编码器 GGUF 里的 `tokenizer.ggml.pre` 写的是 `gemma4`，这是为了兼容**原版 llama.cpp**。
mmBERT 的分词方案（Metaspace + U+2581 规范化 + 全文 BPE）与 Gemma 4 同源，所以原版能识别。

已知差异：原版构建下，**直接对文本使用 `llama-cli` / `llama-tokenize` 时**，开头会少一个
空格前缀，且连续空格会被切成 `▁▁` 而不是 `▁`+`▁`。

**这不影响 Laya 推理** —— 运行时自己用 `tokenizer.json` 分词，再把 token id 数组发给
`/embeddings`（llama-server 接受原始 token 数组），完全绕过了 llama.cpp 的分词器。

要让文本分词也完全对齐，需要给 llama.cpp 打一个小补丁（新增 `mmbert` 预分词器 + BPE
路径的补前导空格，共 24 行）。补丁不在本仓库提供，如需要可自行实现。

## 验证结果

以下均为转换后实测，不是推断：

| 项目 | 结果 |
|---|---|
| GGUF 结构 | 134 个编码器张量，`modern-bert` 元数据正确 |
| llama.cpp 加载并执行编码器 | 通过，输出 768 维隐状态 |
| 编码器 vs 从检查点重构的 PyTorch 参考实现 | 最大绝对误差 **1.1**（数值量级 **41**，22 层 fp16 累加） |
| 滑窗注意力是否真的生效 | 局部模式参考差 **1.1**，全注意力参考差 **30.4** |
| 决策头权重经 GGUF 往返 | 36/36 张量完全一致，0 处不匹配 |
| 完整流程 vs 官方 laya 0.3.4（相同 token id，4 道题） | 概率差 **≤ 0.0006**，argmax 全部一致 |
| 速度（CPU / Vulkan） | 约 110 ms / 约 80 ms 每题 |

## 能力边界（实测）

### 可靠的

三道**信号明写、无需推断**的工单（取消订阅后仍被扣费并要求退款 / 点导出就崩溃 /
第二次投诉配送并明确说会取消），中英同义，每题一个分类题 + 一个是非题：

| 题目 | 英文 | 中文 |
|---|---|---|
| 该哪个团队处理 | billing ✓ (1.000) | billing ✓ (0.999) |
| 是否要求退款 | P=0.990 ✓ | P=0.998 ✓ |
| 该哪个团队处理 | technical ✓ (1.000) | technical ✓ (0.998) |
| 是否报软件问题 | P=0.979 ✓ | P=0.980 ✓ |
| 该哪个团队处理 | support ✓ (0.923) | support ✓ (0.986) |
| 是否说会取消 | P=0.983 ✓ | P=0.958 ✓ |

**英文 6/6，中文 6/6，判断结果完全一致。**

### 不可靠的

同一批隐含语义题（文中未直说、需要推断）：

| 题目 | 英文 | 中文 |
|---|---|---|
| 是否可能流失（"宁愿不要把时间花在换供应商上"） | 0.158 ✗ | 0.051 ✗ |
| 是否在质疑收费 | 0.004 ✗ | 0.094 ✗ |
| 是否情绪不满（开头是"感谢"） | —— | 0.120 ✗ |

分层探测（中文，题面与工单同语言）：

| 层次 | 得分 |
|---|---|
| 字面层（文中是否出现 CFO / 40% / 28号 / CSV） | **6/6** |
| 明确语义（文中直说的事实） | 3/4 |
| 隐含语义（文中未直说，需推断） | **1/4** |

### 四条可复现的结论

1. **中文能读，但两种语言都不会推断。** 字面层满分说明全链路正常；崩塌点在推断上，
   且中英一致。**边界不在语言，在"是否需要推断"。**
2. **模型是确定性的。** 同一输入重复调用，极差为 `0.000000`。
3. **它对措辞的敏感度高于对语义的敏感度。** 同一道题只改选项措辞，P 从 0.043 变到 0.930。
   把问题里的关键词原样放进选项能救回一部分（`needs_human` 0.227 → 0.611），
   但救不回流失风险（0.218 → 0.168）。
4. **置信度不可用于兜底。** 简单题答对时 P 在 0.92–1.00；而难题上它给 `sales` 打
   **0.791（错误）**、给 `is_bug_report` 打 **0.956（错误）**。**错的时候不会告诉你它不确定。**

## 官方定位

以上结果与官方文档一致：base 检查点在具体工作流上接近随机
（typed-decisions 基准 0.342，随机 0.318，多数类基线 0.461），
**它比多数类基线还差**。价值来自按你自己的数据微调：

- 微调后同一基准从 0.362 提到 **0.766**
- 温度校准后 ECE 从 0.314 降到 0.106

## 适用与不适用

**适合：**

- 明确的分类与是否判断，中英文均可，P ≥ 0.9 时自动放行是安全的
- 量大、单条价值低、需要一致口径的粗筛分流
- 本地部署、无 API 成本、数据不出内网
- 作为微调的起点

**不适合：**

- 直接拿来做业务判断（新领域上接近随机，**不要开箱即用**）
- 任何需要"话里没明说但意思是那样"的推断
- 依赖置信度做自动化门控（出错时置信度不下降）

## 许可

原模型 Apache 2.0，作者 Convai Innovations。本转换仅改变序列化格式，不改变权重。

## 引用

- 模型主页：https://huggingface.co/convaiinnovations/laya-multilingual
- PyPI：https://pypi.org/project/laya/
- GitHub：https://github.com/NandaKishorM/laya

---

## 上传到魔搭（仓库维护者）

三个权重文件和本 README **必须一次上传**。只上传 README 单个文件不会带上权重。

`modelscope upload` 作用于**整个文件夹**，不是单个文件，所以直接指向目录即可：

```bash
pip install modelscope
modelscope login --token <你的 SDK 令牌>     # 令牌在 modelscope.cn/my/myaccesstoken 获取

modelscope upload <用户名>/<仓库名> "D:\Ai\model\laya"
```

上传后用 `modelscope download` 校验文件确实都在：

```bash
modelscope download <用户名>/<仓库名> --local_dir ./check
ls -la ./check
```

应当看到 6 个文件（3 个权重 + 2 个 .py + README.md）。若只剩 README.md，说明上传时路径指到了文件
而不是目录。

> 目录中若有多余文件（例如空的"新建 文本文档.txt"），建议先删除，或用
> `--exclude "*.txt"` 排除，避免污染仓库。

---

# English

## What this is

A **GGUF conversion** of `convaiinnovations/laya-multilingual` (laya 0.3.4), loadable by
llama.cpp.

Laya is a **non-autoregressive decision model** (System 1): given a *state* (text, email,
ticket, JSON) and typed questions, it returns typed answers with probabilities in a
**single forward pass**. No text generation, so nothing to parse and nothing to
hallucinate.

> **Important: this is not a single self-contained file.** Read "Scope of this
> conversion" first.

## Scope of this conversion (read this first)

Laya is **not** a stock HuggingFace checkpoint, so `convert_hf_to_gguf.py` cannot read it
as-is. It consists of:

| part | what it is | runs where |
|---|---|---|
| `encoder.*` | mmBERT-base bidirectional encoder: 22 layers, hidden 768, 12 heads, FFN 1152 (GeGLU), alternating full/sliding attention (window 128), RoPE theta 160000, 256k vocab | **llama.cpp** (`modern-bert` graph) |
| `head.layers.*` | 2 extra `nn.TransformerEncoderLayer`s (d=768, nhead=12, ffn=3072, `norm_first`) | runtime (PyTorch) |
| `scorer.*`, `act_head.*`, `type_emb.*` | option-marker scorer, act/escalate head, question-type embedding | runtime (PyTorch) |

Decisions are produced by scoring every answer option at its own `[MASK]` token and then
softmaxing over that question's options. **This head is not part of llama.cpp's
`modern-bert` graph and cannot be added to it** — llama.cpp's loader requires the tensor
count in the file to match exactly the tensors its graph creates, so extra tensors make
loading fail outright.

So: **the 300M-parameter encoder runs inside llama.cpp, and the decision head runs in a
companion PyTorch runtime.**

### Why two files

llama.cpp's loader validates the tensor count; not even one extra tensor is tolerated.
The head weights therefore cannot ride along in the encoder GGUF and must be a separate
companion file.

## Files

This repository has 6 files. They **must sit in the same directory**:

| file | size | purpose |
|---|---|---|
| `laya-multilingual-f16.gguf` | 600 MB | `modern-bert` encoder + 256k BPE tokenizer. **llama.cpp loads this directly.** |
| `laya-multilingual-f16-head.gguf` | 57 MB | decision-head weights, read by the runtime |
| `tokenizer.json` | 33 MB | the runtime tokenizes with this (it drives llama.cpp with explicit token ids) |
| `laya_gguf.py` | 11 KB | the runtime itself (~200 lines) |
| `gguf_weights.py` | 1 KB | helper that reads the head GGUF into tensors |
| `README.md` | - | this file |

After downloading, the layout should be:

```
laya/
├── laya-multilingual-f16.gguf
├── laya-multilingual-f16-head.gguf
├── tokenizer.json
├── laya_gguf.py
├── gguf_weights.py
└── README.md          (this file)
```

> **Downloading only `laya-multilingual-f16.gguf` is not enough.** The encoder will load,
> but there is no decision capability - the head, the tokenizer and the runtime code live
> in the other files.

### Requirements

Python packages: `torch`, `numpy`, `tokenizers`, `gguf`
(`pip install torch numpy tokenizers gguf`). llama.cpp handles the encoder, so
`transformers` is not needed.

### Why three weight files are required

`laya-multilingual-f16.gguf` contains the encoder only. llama.cpp's loader validates the
tensor count and fails on even one extra tensor, so the head weights cannot be packed into
the same GGUF. The runtime needs `tokenizer.json` because it tokenizes itself and sends
the token id array to llama.cpp.

## Quick start

### 1. Start the llama.cpp server

```bash
llama-server -m laya-multilingual-f16.gguf \
             --embeddings --pooling none \
             --ctx-size 1024 --port 8081
```

> On Windows PowerShell use `.\llama-server.exe`; PowerShell does not load commands from
> the current directory by default.

### 2. Run it

This needs a companion runtime `laya_gguf.py` (~200 lines of Python; it tokenizes, pulls
hidden states from llama.cpp, then runs the decision head in PyTorch). Put the three
weight files next to this README and run from that directory:

```python
from pathlib import Path
from laya_gguf import LayaGGUF

HERE = Path(__file__).parent          # directory holding the three weight files

agent = LayaGGUF(
    HERE / "laya-multilingual-f16.gguf",
    HERE / "laya-multilingual-f16-head.gguf",
    HERE,                             # directory holding tokenizer.json
    base_url="http://127.0.0.1:8081",
)

result = agent.predict(
    {"body": "I cancelled on the 3rd but you still charged me $49 for April. Refund my card."},
    {
        "department": {
            "type": "choice",
            "instructions": "Which team should handle this `body`?",
            "criteria": {
                "billing": "invoices, charges, refunds",
                "technical": "software errors and bugs",
                "sales": "buying and upgrading",
            },
        },
        "wants_refund": {
            "type": "noul",
            "instructions": "Is the sender asking for their money back?",
            "criteria": {"false": "they are not asking for money",
                         "true": "they are asking for money to be returned"},
        },
    },
)

print(result["answers"]["department"]["choice"])       # billing
print(result["answers"]["wants_refund"]["noul"])       # 0.99
```

### The three question types

| type | output | use for |
|---|---|---|
| `choice` | probability per option + confidence | department routing, intent classification, topics |
| `score` | expected level on an ordinal rubric | urgency, frustration, severity |
| `noul` | calibrated P(true), 0.0-1.0 | yes/no judgements: spam, jailbreak, churn |

## About the tokenizer

The encoder GGUF records `tokenizer.ggml.pre = gemma4`, chosen for compatibility with
**stock llama.cpp**. mmBERT's tokenization scheme (Metaspace + U+2581 normalizer + BPE
over the whole text) is the same family as Gemma 4, so stock builds recognise it.

Known difference: on a stock build, **when you feed text directly to `llama-cli` /
`llama-tokenize`**, the leading space prefix is dropped and runs of spaces split as `▁▁`
rather than `▁`+`▁`.

**This does not affect Laya inference** — the runtime tokenizes with `tokenizer.json`
itself and sends the token id array to `/embeddings` (llama-server accepts a raw token
array), bypassing llama.cpp's tokenizer entirely.

Full text-tokenization parity would need a small llama.cpp patch (a new `mmbert`
pre-tokenizer plus a prepend-space step in the BPE path, 24 lines). The patch is not
included here; implement it yourself if you need it.

## Verification

All measured after conversion, not inferred:

| check | result |
|---|---|
| GGUF structure | 134 encoder tensors, correct `modern-bert` metadata |
| llama.cpp loads and runs the encoder | pass, emits 768-d hidden states |
| Encoder vs a PyTorch reference rebuilt from the checkpoint | max abs error **1.1** on values up to **41** (fp16 accumulation over 22 layers) |
| Sliding-window layers really use a 128-token window | local-pattern reference **1.1**, all-full-attention reference **30.4** |
| Head weights round-trip through GGUF | 36/36 tensors identical, 0 mismatches |
| Full pipeline vs official laya 0.3.4 (same token ids, 4 questions) | probability diff **≤ 0.0006**, argmax identical |
| Latency (CPU / Vulkan) | ~110 ms / ~80 ms per question |

## Capability boundaries (measured)

### Reliable

Three tickets whose signal is **stated outright and needs no inference** (charged after
cancelling and asking for a refund / app crashes on export / second delivery complaint
that explicitly says they will cancel), one classification plus one yes/no question each,
English and Chinese:

| question | English | Chinese |
|---|---|---|
| which team | billing ✓ (1.000) | billing ✓ (0.999) |
| asking for a refund | P=0.990 ✓ | P=0.998 ✓ |
| which team | technical ✓ (1.000) | technical ✓ (0.998) |
| reporting a software problem | P=0.979 ✓ | P=0.980 ✓ |
| which team | support ✓ (0.923) | support ✓ (0.986) |
| says they will cancel | P=0.983 ✓ | P=0.958 ✓ |

**English 6/6, Chinese 6/6, identical decisions.**

### Unreliable

Implied-meaning questions on the same tickets (not stated outright, requires inference):

| question | English | Chinese |
|---|---|---|
| possible churn ("rather not spend next quarter onboarding a different vendor") | 0.158 ✗ | 0.051 ✗ |
| disputing a charge | 0.004 ✗ | 0.094 ✗ |
| sender is unhappy (opens with "thanks") | - | 0.120 ✗ |

Layered probe (Chinese, question language matching the ticket):

| layer | score |
|---|---|
| literal (does the text mention CFO / 40% / the 28th / CSV) | **6/6** |
| explicit semantics (facts stated outright) | 3/4 |
| implied semantics (not stated, requires inference) | **1/4** |

### Four reproducible findings

1. **It reads Chinese, but neither language gets inference.** Full marks on the literal
   layer proves the whole chain works; the collapse is on inference, and English and
   Chinese fail equally. **The boundary is not language, it is whether inference is
   required.**
2. **The model is deterministic.** Repeated calls on the same input differ by `0.000000`.
3. **It is more sensitive to wording than to meaning.** Changing only the option wording
   moved one question's P from 0.043 to 0.930. Copying the question's keywords into the
   options rescues some (`needs_human` 0.227 -> 0.611) but not churn risk (0.218 -> 0.168).
4. **Confidence cannot be used as a safety net.** On easy questions it scores 0.92-1.00
   when correct; on hard ones it gave **0.791 (wrong)** for `sales` and **0.956 (wrong)**
   for `is_bug_report`. **It does not tell you when it is unsure.**

## Official positioning

These results match the upstream documentation: the base checkpoint is near chance on
specific workflows (typed-decisions 0.342 against a 0.318 random and 0.461 majority-class
baseline) — **below the majority-class baseline**. The value comes from fine-tuning on
your own data:

- fine-tuning moves the same benchmark from 0.362 to **0.766**
- temperature calibration moves ECE from 0.314 to 0.106

## Suitable / not suitable

**Suitable for:**

- explicit classification and yes/no judgements, in English or Chinese; auto-release at
  P >= 0.9 is safe
- high-volume, low-value-per-item triage that needs a consistent rule
- local deployment, no API cost, data never leaves your network
- a starting point for fine-tuning

**Not suitable for:**

- making business decisions directly (near chance in a new domain — **do not use
  out of the box**)
- anything needing "they did not say it but they meant it" inference
- gating on confidence (confidence does not drop when it is wrong)

## License

The original model is Apache 2.0 by Convai Innovations. This conversion only changes the
serialization format; the weights are unchanged.

## Uploading to ModelScope (repository maintainers)

The three weight files and this README **must be uploaded together**. Uploading the README
alone does not bring the weights along.

`modelscope upload` operates on a **whole folder**, not a single file, so point it at the
directory:

```bash
pip install modelscope
modelscope login --token <your SDK token>   # token: modelscope.cn/my/myaccesstoken

modelscope upload <username>/<repo-name> "D:\Ai\model\laya"
```

Verify afterwards with `modelscope download`:

```bash
modelscope download <username>/<repo-name> --local_dir ./check
ls -la ./check
```

You should see 6 files (3 weights + 2 .py files + README.md). If only README.md is present, the upload path
pointed at a file rather than the directory.

> Remove stray files first (for example an empty `新建 文本文档.txt`), or exclude them with
> `--exclude "*.txt"`, so the repository stays clean.
