Preprint

Pattern-Making AI Linked to Faster Work and Better Patterns

Preprint: A 20-person comparison reported higher task-success odds, shorter completion times and higher expert-rated artifact quality with TailorCoPilot than with baseline tools.

A pattern-making system was linked to better reported performance in a small user comparison. Participants using TailorCoPilot had higher odds of completing tasks successfully, shorter completion times and higher expert ratings for the patterns they produced than with baseline tools.

The study involved 20 people, including 10 novices and 10 advanced novices. Each completed 12 trials across two sessions, with simple, medium and complex pattern-making tasks. Researchers measured task success, completion time, expert-rated artifact quality and perceived workload.

A record of each accepted design step

TailorCoPilot uses TailorTrace, which records design states and the operations that connect them. The system keeps only validated states with their semantic operation sequences and rejects invalid commits. The result is a record in which each stored step is tied to a state the system has accepted as valid.

The training material consisted of 93 design traces. Fifty-six were textbook traces covering three to 10 states, while 37 came from live production and covered 10 to 30 states. The researchers split the data at the trace level, used stratified sampling and expert verification, and created approximately 2,500 translation pairs to fine-tune Qwen3-VL-8B.

The evaluation combined an interactive comparison of TailorCoPilot with baseline tools and a separate diagnostic comparison of three trace conditions. The diagnostic test examined a model using no traces, textbook traces, or textbook traces combined with expert-authored traces.

The advantage varied by task

Across the main comparison, the reported odds of task success were 2.31 times higher with TailorCoPilot than with baseline tools. The 95% confidence interval, an estimate of the uncertainty around that result, ran from 1.42 to 3.86, with p = .001. The reported advantage appeared across difficulty levels for novices and was wider for advanced novices on medium and complex tasks.

Completion times were also reported as shorter with TailorCoPilot. The size of the difference varied with task difficulty, with p < .001 for the completion-time comparison and p = .018 for the variation by difficulty. CAD remained competitive for advanced novices on some simple tasks.

Experts gave higher quality ratings to artifacts associated with TailorCoPilot, with p = .001. The clearest reported gains were in structural balance and conformity to the reference, with smaller but consistent gains in sewability. Agreement among the raters was strong, with an intraclass correlation coefficient, a measure of rating consistency, of 0.81.

Among novices, the average sewability rating was 3.82 with TailorCoPilot versus 2.02 with baseline tools. Structural-balance ratings averaged 3.97 versus 1.48, and reference-conformity ratings averaged 3.97 versus 1.35. Advanced novices also received higher TailorCoPilot ratings across all four reported dimensions, including overall quality.

Participants reported lower workload on the NASA-TLX scale with TailorCoPilot, with p < .001. Exit interviews indicated that they retained a sense of agency and ownership while revising their patterns.

The diagnostic conditions showed different results

The separate diagnostic study tested 15 complex pattern-making tasks under three conditions using the same references, base patterns and inference budgets. The model-only condition used a pre-trained Qwen3-VL-8B, the textbook condition used textbook traces, and the combined condition used textbook traces plus expert-authored traces.

The model-only condition recorded no successful tasks and received a mean expert rating of 1.3, with a standard deviation of 0.5. The textbook-trace condition succeeded on three of 15 tasks, or 20.0%, with a mean rating of 1.9 and a standard deviation of 0.8. The condition combining textbook and expert traces succeeded on 10 tasks, or 66.7%, and received the highest mean rating, 4.1, with a standard deviation of 0.6. The reported comparison between the textbook-only and combined conditions yielded W = 1 and p < .001.

A bounded test, not a production verdict

The authors describe the evaluation as short-term and based on a modest sample. They also identify limits in the system's scoped operator representation and its dataset of 93 traces with approximately 2,500 derived pairs. The study used unified default fabric-simulation parameters, leaving material-specific behavior and production variability for future work.

Those limitations mean the findings are best read as evidence from this bounded test, not as a general verdict about garment production or independent long-term skill.

The authors acknowledge support from Style3D Research, the Fundamental Research Funds for the Central Universities through grants CUSF-DH-T-2025007 and 2232026G-08, and the International Cooperation Fund of the Science and Technology Commission of Shanghai Municipality through grant 21130750100.

The document is labeled an arXiv version 1 record dated 26 Aug 2026.

Paper data and sources

Original title: TailorCoPilot: Enabling Agentic Pattern Making with Version-Controlled State Tracking
Authors: Yuexin Sun, Zhaohui Wang, Ruiyang Liu et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.