The AGENT-O assessment points to a clear weakness in how health-oriented AI-agent work is documented: papers described evaluation and benchmark procedures more completely than runtime architecture, meaning how systems are put together and operate, governance and safety, or provenance and reproducibility, the trail showing where information comes from and whether work can be retraced. The pattern emerged in a model-assisted scoring exercise covering 279 records. The average completeness score was 63.7 out of 100, and the median was 67.5.
Among papers for which a dimension applied, incomplete reporting was highest for runtime/architecture at 84.6%, followed by governance/safety at 82.8% and provenance/reproducibility at 78.1%. The corresponding incomplete rates were 25.8% for evaluation and 29.8% for benchmark-process alignment. The authors characterize that contrast as an evaluation-specification gap: studies spelled out how tests and benchmarks were run more fully than how the agents themselves were structured, governed and described in reproducible detail.
All 279 papers received a score, with no failed or unscored cases. Seventy-two papers, or 25.8%, were in the 50 to 64.9 band, while 130 papers, or 46.6%, were in the 65 to 79.9 band. These are reporting scores, not measures of whether an agent performed well.
A common format for describing agents
AGENT-O was developed as a modular ontology framework for health-oriented AI-agent systems. An ontology is a structured vocabulary for recording the parts and relationships of a system. The framework pairs that vocabulary with a semantic Agent Card and a workflow for assessing publication-level reporting completeness. Its evaluation combined ontology inventory, OWL-RL reasoning, three SHACL constraint suites, 12 SPARQL competency queries, three cases and model-assisted assessment across five dimensions.
The score was weighted and scaled to 100, using only dimensions that applied to each paper. A present item received 1.0, a partial item 0.5 and a missing item 0. The weights were 25% each for runtime/architecture and evaluation, 20% each for provenance/reproducibility and governance/safety, and 10% for benchmark-process alignment.
The underlying corpus contained 278 extracted paper documents plus one prespecified AgentArena case, for 279 records in all. It came from two review-associated inventories rather than an independent systematic search. The collection was therefore not intended to exhaustively cover the health-oriented AI-agent literature.
The scoring labels were not adjudicated by multiple human reviewers. The results are model-assisted estimates from the selected corpus, not a human-validated reference standard. That limits how confidently the reported percentages can be extended beyond the material that was assessed.
Checks on the release
The evaluated release contained 1,962 RDF triples and 1,922 Protégé axioms, formal statements used in the ontology. Of those axioms, 687 were logical and 506 were declarations. The release included 252 active classes, 198 active object properties and 51 active datatype properties across six modules.
Release checks found 25 Turtle files that parsed successfully, 183 external mappings and three application profiles covering architecture, governance and reporting.
Automated checks found no listed validation failures, including parse failures, deprecated predecessor namespaces, missing labels, missing domains or ranges for active properties, or undefined AGENT-O alignment terms. The three SHACL suites conformed without findings on their corresponding example graphs, and all 12 competency queries returned prespecified evidence.
Three cases, different levels of detail
The three case descriptions could all be represented in AGENT-O, but their reporting completeness differed. AgentArena was partial on provenance/reproducibility and governance/safety. MedAgent-Pro was more complete on several dimensions but remained partial on governance/safety. The multi-agent consensus case satisfied all five dimensions in the curated annotation.
What the score does not tell us
The paper draws a firm boundary around the score. It measures publication content, not agent performance, scientific correctness, clinical utility, fairness, safety or deployment readiness.
A more complete paper does not show that its underlying agent is effective, clinically useful or ready for deployment. The work is a reporting framework and a publication-level gap assessment, not a performance evaluation.
Within that remit, the central finding is a mismatch in the kinds of detail papers provide. Evaluation and benchmark procedures were reported more completely than runtime architecture, governance and safety, and provenance and reproducibility. AGENT-O is presented as a way to represent those elements semantically and identify reporting gaps.
The work was supported by the National Institutes of Health through grants R01AG083039 and R01AG084236. The authors declared no competing interests, and the paper reports that an AGENT-O public-release repository is available.
Paper data and sources
Original title: AGENT-O: A Semantic Agent Card Framework for Interoperable and Governed Healthcare AI Agents
Authors: Pengze Li, Cui Tao
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-28
DOI: Not available
Original paper · Full text