AI systems that build and revise deliverables are being judged more by what they finish than by how they get there, according to a review of 259 works. The review found that evidence was strongest for delivered artifacts and bounded executions, while construction trajectories, use outcomes and validity beyond tested cases were covered less consistently. Only three protocols assessed a system property beyond task capability, and each covered one property.
What counted as an agentic system
The review uses a specific test for what it calls agentic artifact creation. An AI system must materially construct or revise a deliverable, carry artifact or process state across decisions, and use at least one intermediate observation to redirect later artifact-related work. Direct generators, observation-independent workflows and systems used only to consume or evaluate artifacts were screened out.
The survey searched arXiv, Google Scholar, Semantic Scholar, ACM Digital Library and IEEE Xplore from January 2023 through August 20, 2026. Citation tracing and targeted venue audits supplemented the searches, followed by deduplication and screening against the definition. The retained corpus contained 230 systems meeting the definition and 29 construction benchmarks.
The process behind the product
To explain how these systems operate, the survey links three functions: an Operational Representation, a Construction Policy and Runtime Verification. In ordinary terms, the system carries a working representation of the artifact or process, decides what construction work to do, and checks what happens during execution. Those checks can redirect later construction decisions, making the workflow stateful and revisable.
The resulting map is deliberately broad. The literature is organized into six artifact families and 16 analytical profiles. Application evidence is grouped into six contexts: creative production, brand communication, educational support, professional work, scientific research and engineering design. The review says context determines the objectives, the bundle of artifacts being built and who holds decision authority.
A harder evaluation problem
That breadth makes evaluation harder to reduce to a single score. The survey separates three targets: the delivered artifact, the construction trajectory and the agentic construction system itself. It also says an evaluation protocol must fix four groups of choices: the task set, the metrics, the run configuration and the analysis plan. The targets are related, but they are not interchangeable, so a score for the finished artifact cannot by itself describe the path that produced it or the system that controlled it.
An example from the survey shows why task choice matters. In the Design Arena figures, Kimi K3 occupied the top percentile on all five Models Arena tasks, yet fell to the 54.3rd percentile on native Android in Agents Arena. GPT-5.6 Sol ranged from the 87.5th percentile on slides to the 34.2nd percentile on mobile apps. These are relative leaderboard percentiles from a public snapshot, so they show task-dependent performance rather than a uniform ranking across tasks.
Revision introduces another problem because later edits can be accompanied by losses in previously covered content or citation quality. The survey cites the Mr. Dre benchmark as reporting regression of 16% to 27% during later revisions. That is benchmark evidence summarized by the review, not a causal estimate produced by the review itself.
Design rules for a moving target
The cross-family synthesis points to a recurring trade-off. Decomposing an artifact can lower local complexity, but it adds coordination and reassembly costs. Learned judges may also contribute little independent evidence when they share the generator's preferences or blind spots.
In response, the authors propose four design principles: Externalize Commitments, Define Control Boundaries, Make Feedback Actionable and Revalidate Affected State. Together, the principles connect commitments, responsibility, actionable feedback and revalidation.
A map, not a verdict
The review's evidence also reflects what could be documented. The authors did not retain database-specific query exports, per-source yields or labels recorded before reconciliation, although released files document the resulting coding and limitations. The paper is an arXiv version 1 preprint dated 28 August 2026.
The study's central message is that a finished deliverable is only one evaluation target, and the review's evidence remains uneven across the rest. For systems that revise work over time, the missing picture is the one between the first construction decision and the final result.
Paper data and sources
Original title: Agentic Artifact Creation: Systems, Evaluation, Principles, and Opportunities
Authors: Tianfu Wang, Zhezheng Hao, Xilin Xia et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-28
DOI: Not available
Original paper · Full text