On a 0-to-100 scale, the unified Qwen3-8B question-answering model scored 62.9 for Muscle EM, which requires the active-muscle set to match exactly; 74.0 for geometric-value accuracy averaged across six feature types; and 65.9 for Direction EM, which requires the target-directed muscle set to match exactly. The reported standard deviations were 9.2, 0.2 and 4.7 points, respectively.
Those results changed sharply when the model was given a different mesh. In a mesh-shuffling control, the question and correct answer stayed the same, but the mesh paired with them was replaced. The 3D Muscle-Aware model then scored 2.2 for Muscle EM, 31.2 for geometric-value accuracy and 40.9 for Direction EM, with standard deviations of 0.5, 0.2 and 0.7 points.
A pipeline built from simulated facts
The framework starts with controlled simulator inputs and turns them into observable tongue geometry and reusable, structured fact records. Deterministic task generators then build questions and answers from those records. A language-naturalization step changes the surface wording while leaving the underlying facts unchanged.
The construction pool contained 295,157 screened activation configurations, of which 295,115 valid mesh configurations were retained. The retained corpus contained 891,156 naturalized records per language: 443,934 shape-description records, 443,859 physics-chain records and 3,363 dose-response records. A separate 600-record English pilot was excluded from all statistics.
For the primary sample-held-out protocol, the split used 265,611 training configurations, 14,752 validation configurations and 14,752 test configurations. The automatic evaluation used 400 anchor-balanced test meshes and 1,200 items, pairing each mesh with active-muscle recovery, geometric-value prediction and target-directed correction.
Specialist readouts raise the scores
Task-specific structured readouts performed better on the matched tasks than the unified decoder. They reached 88.7 for Muscle EM, 87.5 for geometric-value accuracy and 93.3 for Direction EM. When the meshes were shuffled, the same scores fell to 1.4, 41.6 and 1.8, respectively.
The authors interpret the large matched-versus-shuffled drop as evidence that answers were grounded in sample-specific geometry. That evidence concerns performance inside the constructed task, not whether a stored simulator activation is a unique physiological cause of a tongue shape.
A harder test, and a language extension
To probe transfer beyond the anchor inventory, researchers used a leakage-controlled test. It removed direct anchor samples, related interpolations and neighbors, source-mesh questions and prescriptive questions using held-out target anchors from both training stages. The evaluation used 500 direct-anchor meshes, balanced at 125 per category.
On that test, the full-inventory model scored 61.8 for Muscle EM, 47.8 for geometric-value accuracy and 50.1 for Direction EM. The held-out model scored 49.7, 44.8 and 49.4, retaining 80.4%, 93.7% and 98.6% of the full-inventory scores. The paper describes the transfer as substantial but incomplete.
The same construction and verification procedure was instantiated in Korean as well. Its scores were 87.8 for Muscle, 69.3 for Value and 48.9 for Direction, but the paper cautions that these figures are not a direct cross-language comparison.
Record-based checks also tested whether naturalized wording kept the facts intact. First-generation faithfulness passed for 87.2% of English outputs and 88.6% of Korean outputs. After regeneration, a full-corpus re-audit ran 3.2 million verification jobs per language and reported zero muscle-identity violations, zero altered gold values and zero span-integrity errors.
What remains untested
The resource remains tightly tied to its simulator. It uses a single Badin anatomy, a fixed mesh topology and no coupled jaw-lip motion. Its activation labels are privileged simulator-defined labels, not uniquely identifiable physiological causes.
In open-ended human ratings, the 3D Muscle-Aware model had a mean factual-accuracy score of 3.81, compared with 2.23 for GPT-5 Pro. GPT-5 Pro scored 3.50 for fluency versus 2.72 for the 3D model, but that difference was not statistically significant after Holm correction, with an adjusted p-value of 0.263.
The work is an arXiv preprint, version 4, dated 1 September 2026. It was partly supported by IITP grants funded by the Korea government, including support under MSIT.
Paper data and sources
Original title: A Simulator-Grounded Framework For Constructing Verifiable Muscle-Grounded QA From 3D Tongue Meshes (extended version)
Authors: Seungho Eum, Unsang Park
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-24
DOI: Not available
Original paper · Full text