---
title: metalingo-direct-metaphor
canonical_url: "https://www.modelscope.cn/models/TommyLeo/metalingo-direct-metaphor"
md_url: "https://www.modelscope.cn/models/TommyLeo/metalingo-direct-metaphor.md"
repository: TommyLeo/metalingo-direct-metaphor
last_updated: 2026-06-10
license: apache-2.0
pipeline_tag: token-classification
tasks:
  - token-classification
model_type:
  - deberta-v2
architectures:
  - DebertaV2ForTokenClassification
base_model:
  - microsoft/deberta-v3-large
base_model_relation: finetune
parameters: 434.0M
tensor_type:
  - F32
library_name:
  - pytorch
  - transformer
  - safetensors
frameworks:
  - Pytorch
language:
  - en
downloads: 54
stars: 0
tags:
  - metaphor-detection
  - direct-metaphor
  - simile
  - token-classification
  - mipvu
  - deberta-v3
---

# metalingo-direct-metaphor

> metalingo-direct-metaphor - TommyLeo 在 ModelScope 开源的模型。MetaLingo-Direct-Metaphor

TommyLeo/metalingo-direct-metaphor 是 ModelScope 魔搭社区上的 434.0M 参数token-classification模型，采用 apache-2.0 许可，基于 microsoft/deberta-v3-large 构建。

- **Repository**: TommyLeo/metalingo-direct-metaphor
- **License**: apache-2.0
- **Tasks**: token-classification
- **Parameters**: 434.0M
- **Base model**: microsoft/deberta-v3-large
- **Tags**: metaphor-detection, direct-metaphor, simile, token-classification, mipvu, deberta-v3
- **Downloads**: 54
- **Stars**: 0
- **Last updated**: 2026-06-10

Source: https://www.modelscope.cn/models/TommyLeo/metalingo-direct-metaphor

---

# MetaLingo-Direct-Metaphor

This model is a fine-tuned version of **Microsoft DeBERTa-v3-large** for **direct metaphor (simile) identification** in English at token level. It is trained on a combination of the BE06 corpus and the Open American National Corpus (OANC), both annotated under the MIPVU framework (Steen et al. 2010). The model simultaneously identifies two token roles within a direct metaphor: the **comparison signal word** (`mFlag`, e.g. *like*, *as*) and the **source-domain content word** (`mrw_lit`, e.g. *bell* in "a voice like a bell").

## Model description

- **Base model:** [microsoft/deberta-v3-large](https://huggingface.co/microsoft/deberta-v3-large)
- **Task:** Token classification (3-class: O / mFlag / mrw\_lit)
- **Training unit:** Sentences, word-level labels
- **Annotation standard:** MIPVU — direct metaphors only (`mFlag` + `mrw_lit` pairs)
- **Negative examples:** Hard negatives — sentences containing mFlag-lexicon words (*like*, *as*, *resemble*, …) but carrying no direct metaphor annotation

## Training & evaluation data

- **Datasets:**
  - **BE06** — Amsterdam BE06 corpus, covering written British English.
  - **OANC** — Open American National Corpus (\~17M tokens, 8 genres: face-to-face speech, telephone, fiction, journalism, letters, non-fiction, technical writing, travel guides).
- **Training annotations:** Both **BE06** and **OANC** training labels were generated by **DeepSeek-V4-Flash** (LLM) applying the MIPVU direct metaphor identification procedure — for each sentence, the LLM was prompted to locate `mFlag` (comparison signal) / `mrw_lit` (source-domain content word) pairs following Steen et al. (2010). No human re-annotation was performed on the training data.
- **Training split:** Positive sentences (mFlag/mrw\_lit pairs present) combined from both DeepSeek-annotated sources; hard negatives (sentences containing mFlag-lexicon words but no annotated direct metaphor) sampled at neg\_ratio = 3.
- **Validation set — VUAMC (human gold standard):** 110 positive sentences + 200 hard-negative sentences from the VU Amsterdam Metaphor Corpus, manually annotated under MIPVU (unchanged from v1). Because this set is independently human-annotated, it provides an out-of-distribution quality check on the DeepSeek-labelled training data.

| Source                              | Annotation method            | Positive sentences | Hard-negative sentences |
| ------------------------------------ | ----------------------------- | -----------------: | -----------------------: |
| BE06 (train)                         | DeepSeek-V4-Flash (MIPVU procedure) |               1,114 |                    4,563 |
| OANC (train)                         | DeepSeek-V4-Flash (MIPVU procedure) |               8,931 |         68,616 available |
| **Training total (neg\_ratio = 3)**  | —                              |          **10,045** |       **30,135 sampled** |
| VUAMC (validation)                   | Human gold standard            |                 110 |                      200 |

## Training hyperparameters

| Parameter           | Value                                                             |
| ------------------- | ----------------------------------------------------------------- |
| Epochs              | 4 (best checkpoint at step 8,792 ≈ epoch 3.5)                     |
| Learning rate       | 2e-5                                                              |
| LR scheduler        | Linear with warmup                                                |
| Warmup ratio        | 0.1                                                               |
| Batch size          | 16                                                                |
| Max sequence length | 256                                                               |
| Weight decay        | 0.01                                                              |
| Class weights       | O = 1.0, mFlag = 5.0, mrw\_lit = 5.0 (sqrt-inverse-freq, cap 5.0) |
| Evaluation interval | Every 1,256 steps (≈ half epoch)                                  |
| Final train loss    | 0.0812                                                            |

## Results

Evaluated on the held-out validation set (110 positive / 200 hard-negative sentences from VUAMC):

### Token-level

| Class        | Precision | Recall |    **F1** |
| ------------ | --------: | -----: | --------: |
| mFlag        |     74.82 |  72.73 | **73.76** |
| mrw\_lit     |     76.97 |  78.00 | **77.48** |
| **Combined** |         — |      — | **76.52** |

### Sentence-level

| Precision | Recall |    **F1** |
| --------: | -----: | --------: |
|     83.06 |  93.64 | **88.03** |

A sentence is predicted positive if at least one token is labelled `mFlag` or `mrw_lit`.

### Comparison with LLM baseline

Same validation set (110 pos + 200 hard-neg from VUAMC), zero-shot DeepSeek-V4-Flash vs. this fine-tuned model:

| | This model | DeepSeek-V4-Flash (zero-shot MIPVU) |
|---|---:|---:|
| sent F1 | **88.03%** | 87.50% |
| sent P | 83.06% | 85.96% |
| sent R | **93.64%** | 89.09% |
| mFlag F1 | 73.76% | ~95.9% (acc) |
| mrw_lit F1 | 77.48% | 86.0% |

Despite being trained entirely on **DeepSeek-V4-Flash-generated labels**, this model now **slightly exceeds the DeepSeek-V4-Flash zero-shot baseline on sentence-level F1** (88.03% vs 87.50%), at a fraction of the inference cost/latency. The remaining gap at token level reflects DeepSeek-V4-Flash's access to explicit MIPVU symbolic reasoning during annotation.

## Label dictionary

```json
{
  "0": "O",
  "1": "mFlag",
  "2": "mrw_lit"
}
```

Subwords are aligned to words via the tokenizer's `word_ids`; the **first subword** of each word is used for prediction.

## Usage example

```python
from transformers import AutoTokenizer, AutoModelForTokenClassification
import torch

model_path = "tommyleo2077/metalingo-direct-metaphor"  # or local path
tokenizer = AutoTokenizer.from_pretrained(model_path)
model = AutoModelForTokenClassification.from_pretrained(model_path)
model.eval()

words = ["Her", "laugh", "was", "like", "a", "bell", "."]
inputs = tokenizer(
    words,
    is_split_into_words=True,
    return_tensors="pt",
    truncation=True,
    max_length=256,
)
word_ids = inputs.word_ids(batch_index=0)

with torch.no_grad():
    logits = model(**inputs).logits
preds = logits.argmax(dim=-1)[0].tolist()

word_preds = {}
for i, wid in enumerate(word_ids):
    if wid is not None and wid not in word_preds:
        word_preds[wid] = preds[i]

id2label = model.config.id2label
for i, w in enumerate(words):
    print(f"{w}\t{id2label[word_preds.get(i, 0)]}")
```

Expected output:

```
Her     O
laugh   O
was     O
like    mFlag
a       O
bell    mrw_lit
.       O
```

## Citation

**Model author:** [Tommy Leo](https://huggingface.co/tommyleo2077) — <1683619168tl@gmail.com>

**Dataset (MIPVU):**

```bibtex
@book{steen2010method,
  title     = {A Method for Linguistic Metaphor Identification: From {MIP} to {MIPVU}},
  author    = {Steen, Gerard and Dorst, Aletta G. and Herrmann, J. Berenike and Kaal, Anna and Krennmayr, Tina and Pasma, Thea},
  year      = {2010},
  publisher = {John Benjamins}
}
```

**Base model:** [Microsoft DeBERTa](https://github.com/microsoft/DeBERTa)

**This model:**

```bibtex
@misc{leo2025metalingodirectmetaphor,
  title        = {metalingo-direct-metaphor: Direct Metaphor Identification with DeBERTa-v3-large},
  author       = {Leo, Tommy},
  year         = {2025},
  howpublished = {\url{https://huggingface.co/tommyleo2077/metalingo-direct-metaphor}},
  note         = {Contact: 1683619168tl@gmail.com}
}
```
