BERT-FT posted the strongest score on the dataset used for native testing, while Skill-ZS had the most even observed profile across three Chinese datasets for sentence-level metaphor identification. BERT-FT reached a native CMRE Test Macro-F1 of 91.76, while Skill-ZS had the highest external floor at 82.64 and the smallest three-dataset range, at 4.08 points. Macro-F1 is the unweighted average of the F1 scores for the metaphorical and non-metaphorical classes.
The figures distinguish between native and external comparisons rather than naming one condition as best on every measure. LLM-FT had the highest average Macro-F1 across the two external datasets, at 83.52, compared with 82.92 for Skill-ZS. The results describe different performance profiles on the observed data.
A split result across the test sets
The document is an arXiv preprint comparing four knowledge-supply regimes. BERT-FT and LLM-FT learned from CMRE labels, while LLM-ZS and Skill-ZS received no labeled examples. In the matched zero-shot comparison, the same base model, runtime, prompt shell, output format and evaluation rows were used, with the frozen Skill as the only condition-level addition.
The main cross-dataset summaries were the external mean, the lowest external score, the signed difference between native and external mean, and the range across the three datasets. Fine-tuned results were reported as means with sample standard deviations over three seeds. Zero-shot results were deterministic point estimates.
What was tested
The evaluation used existing research datasets and model-generated predictions, with no new personal or human-subject data collected. After exact and normalized-exact matches with CMRE Train or Dev were removed, the retained sets contained 6,794 CMRE Train rows, 850 CMRE Dev rows, 850 CMRE Test rows, 1,162 CCIME rows and 698 CMC rows. The same retained rows were used across all conditions.
The overlap filtering excluded 38 CCIME rows and two CMC rows. The datasets also used different annotation policies: CMRE retained released-positive similes as metaphorical, while CCIME assigned similes and hyperbole to the non-metaphorical class.
The frozen Skill translated principles about contextual meaning, basic meaning, contrast and comparison into six operations. Its wording was refined through seven iterations on 60 AI-assisted synthetic specification sentences. Those sentences were excluded from formal evaluation and used to revise the instruction rather than estimate performance.
The balance between precision and recall varied
In the matched zero-shot results, Skill-ZS had higher metaphor precision but lower recall than LLM-ZS on CMRE Test and CMC. On CMRE Test, precision was 88.58% versus 83.62%, while recall was 67.53% versus 79.29%. On CMC, precision was 88.48% versus 87.06%, while recall was 70.07% versus 90.88%. Precision is the share of metaphor predictions that were correct; recall is the share of actual metaphorical cases that the system found.
CCIME showed a different balance. Skill-ZS had metaphor precision of 77.56%, compared with 61.80% for LLM-ZS, while recall was similar at 90.83% and 91.01%. The Macro-F1 score for Skill-ZS was 15.90 points above LLM-ZS on that dataset. Across the three datasets, the class-level results varied by dataset rather than showing a uniform direction.
The aggregate decisions followed the same dataset-dependent pattern. Skill-ZS predicted the metaphorical class less often on every dataset. On CCIME, the predicted-metaphor rate was 57.14% for Skill-ZS and 71.86% for LLM-ZS; false positives were 149 and 319, respectively, while false negatives were 52 and 51. False negatives rose by 50 on CMRE Test and by 57 on CMC in the matched comparison.
A result with a narrow reach
The contrasts were reported descriptively, without significance claims, inferential tests or confidence intervals. Fine-tuned conditions supplied sample standard deviations across three seeds, but each zero-shot condition supplied one deterministic formal result, leaving no comparable between-run dispersion. The largest displayed dispersion was 2.77 points for LLM-FT on CCIME.
The evidence covers four prespecified conditions, one base LLM, one expert-informed Skill and three Chinese sentence-level datasets. CCIME and CMC provide two observed external settings. The observed mean, floor and range do not establish performance on unobserved genres, languages, annotation schemes or metaphor constructions.
The supplementary package is described as containing preprocessing records, prompts, the frozen Skill, runtime configurations, instance-level predictions, evaluation code and figure source where redistribution terms permit. The supplied text does not report a funding source or a conflict-of-interest statement.
Paper data and sources
Original title: Cross-Dataset Stability of Expert-Informed Skill Prompting and Fine-Tuning for Chinese Metaphor Identification
Authors: Yufeng Wu, Meichun Liu
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text