A preprint reports that MediSkill-Evo scored higher than the best-performing prior agent on a fixed Qwen FullChain benchmark. The reported relative changes were a 7.81% improvement in diagnosis accuracy, a 70.67% improvement in treatment-intent coverage and a 43.04% reduction in critical failures. These are benchmark comparisons, not independently clinician-adjudicated clinical outcomes.
The system organizes trajectory-derived knowledge into four typed banks: clinical, process, symbolic and visual. It does not fine-tune the underlying model. A Process-Constrained Preference Harness then grounds candidate actions in evidence and prioritizes safer decisions.
The paper asks how trajectory-derived knowledge with different epistemic roles can be published and used through type-dependent validation and decision authority.
What the benchmark tests found
With DeepSeek-V4-Flash, MediSkill-Evo led seven of eight non-diagnosis metrics. Relative improvements were 37.9% for treatment-intent coverage, 163.9% for evidence recall, 1,070.9% for required-history recall and 190.4% for gated interaction efficiency; relative reductions were 59.1% for unnecessary examinations and 29.8% for critical failures. The comparisons were automatic benchmark comparisons; no confidence intervals or p-values were reported, and there was no independent clinical adjudication.
Under controlled stress, the stress-process composite showed a 7.77% relative improvement and required-action completion a 12.41% relative improvement over the best-performing agent for each metric. The system also showed stronger recovery of patient facts, temporal evidence and triage red flags, with no controller-scored errors in the specified checks for unavailable evidence, treatment and triage safety.
In aggregate controlled-stress outcomes, MediSkill-Evo had the highest registered stress-process score at 83.79%, required-action recall at 88.30%, treatment-intent coverage at 82.40% and core score at 80.03%. Diagnosis accuracy was 93.89%. Automatic safety violations and critical failures were 1.11% and 2.78%, respectively.
The evaluation covered 300 MIMIC-IV-derived FullChain encounters, 180 hard-isolation conditions covering six process obligations and 100 multimodal NEJM image-diagnosis cases. The findings are limited to those benchmark cases and conditions and do not establish clinical safety or population-level generalization.
The test splits were fixed before evaluation, while reference diagnoses and evaluator targets were kept outside the Doctor-visible interaction. Within each block, comparisons used the same model backbone and environment; calls used temperature zero, and each case–configuration pair had one observed rollout.
The image comparison was not tool-only
On the 100 NEJM image-diagnosis cases, the MedSAM-enabled condition outperformed the strongest MemP or Reflexion result. Relative improvements were 2.6% in diagnosis accuracy, 35.6% in required-history recall, 49.3% in required-test recall and 19.0% in core score; unnecessary examinations saw a relative reduction of 48.4%.
The optional MedSAM comparison did not isolate the image tool alone: the learned Measurement Bank was condition-specific and evolved on the corresponding training condition.
In a matched component analysis, the complete system reached 76.00% diagnosis accuracy and 69.40% treatment-intent coverage, with critical failures at 13.00% and safety violations at 0.00%. Compared with no memory, relative improvements were 13.4% in diagnosis and 90.1% in treatment coverage, while critical failures saw a relative reduction of 56.7%.
The authors describe that analysis as hypothesis-generating. Because it used one frozen run per profile, it cannot establish that any individual bank was necessary or causally beneficial.
A boundary case shows the risk behind the image result. Visual annotations were overinterpreted, and an unverified posterior-fossa finding entered the differential and escalation plan. A correct automatic label did not validate the image interpretation or show that MedSAM improved the case.
What remains untested
Treatment-intent, safety, critical-failure and semantic-stress measures were produced by automatic evaluators rather than independent clinician adjudication. They support comparisons within the benchmarks, but do not establish construct calibration, clinical certification or prospective validity.
The authors say the findings do not establish clinical safety, population-level generalization, judge construct validity or causal credit for individual components. They call for controlled mechanism comparisons and independent clinical calibration.
The supplied manuscript is identified as arXiv:2608.23397v2. Its artifact release includes orchestration, comparator configurations, evaluator and table scripts, manifests, hashes and reconstruction instructions. It excludes credentials, raw MIMIC-derived cases, learned artifacts that failed leakage review and NEJM images; authorized users can rebuild restricted inputs from source indices and validators.
Paper data and sources
Original title: MediSkill-Evo: Process-Constrained Self-Evolution for Evidence-Grounded Clinical Interaction
Authors: Ruoyu Wu, Shenfu Xie, Yinqian Sun et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-24
DOI: Not available
Original paper · Full text