A new preprint on autonomous driving argues that the field should stop treating a low prediction error as a stand-in for competent driving. The more meaningful test, it says, is whether a system’s learned structure, training signals and evaluation methods work together to support safe, reliable and reproducible planning.
The document is an arXiv version-1 preprint dated 20 Aug 2026. It is a structured narrative review of public research, benchmark evidence and released resources in end-to-end autonomous driving, rather than a new driving experiment.
From steering imitation to inspectable plans
The review presents the field as moving away from systems that mainly imitate low-level controls and toward approaches whose representations, supervision and tests make planning behavior easier to inspect and compare.
To organize that change, the survey groups methods along four lines: what the system takes in, what it predicts or plans, what learning signal guides it, and how its performance is evaluated. That structure is meant to keep architectural choices connected to the evidence used to judge them.
For each work retained in the review, the authors recorded details such as its input representation, planning output, supervision signal, evaluation protocol, benchmark evidence and public-resource status when those details were available. Eligible work could introduce an influential end-to-end formulation, predict planning-relevant outputs, contribute a planning benchmark, add world-model or vision-language mechanisms, release public resources, or examine safety, interpretability, robustness or evaluation.
The review does not describe that progression as a clean march toward a single winning design. Instead, it identifies a recurring trade-off: one generation may address covariate shift, recovery or another weakness while exposing a different problem, including learner-expert asymmetry, reliance on intermediate labels or maps, uncertainty in world models, or the difficulty of translating text into control.
Why the score on paper can mislead
A central concern is the gap between open-loop and closed-loop evaluation. Displacement-based open-loop measures can be useful diagnostics, the review says, but they do not by themselves show that a vehicle will drive safely or behave in a way people would recognize as aligned with the road situation.
The survey therefore calls for results to be compared only within aligned protocols. Open-loop results, non-reactive tests, reactive closed-loop driving and preference-aware assessments answer different questions, so claims from one category should not be presented as if they were evidence from another.
The review says cross-benchmark evidence suggests that open-loop scores designed with safety in mind may track closed-loop driving scores more closely than simple average or final displacement errors. Even that relationship is not dependable enough to remove the need for careful testing: rankings can reverse between evaluations, and individual submetrics can stop distinguishing systems once they reach a high level.
That warning also changes how benchmark tables should be read. A result is inseparable from the sensor setup, controller, metric implementation and other protocol details behind it. The survey’s recommendation is not to discard open-loop measures, but to treat them as one kind of evidence and keep their limits visible.
More reasoning does not settle the safety question
World models are presented as potentially useful when they help a system imagine how a scene may develop and choose among possible trajectories. But the review says the decisive evidence should be whether those models select safe, goal-consistent and interaction-aware paths, including in rare scenarios and conditions that differ from the data used to build them. Realistic-looking generated scenes, on their own, are not enough.
The same caution applies to vision-language-action systems, which connect visual understanding and language-based reasoning with actions. Strong reasoning in a vision-language model does not automatically establish safe control. Relevant evaluations need to show a connection from language-grounded reasoning to trajectory choice, recovery after problems, controllability and decisions in long-tail situations.
Across the public literature it reviews, the survey says safety is still usually reported through empirical improvements rather than formal guarantees. Public evidence is dominated by benchmark performance, not by deployable safety certificates. That leaves an important distinction between a system that scores well in a test and one for which operational safety has been established.
A map of the evidence, not a verdict on a car
The paper’s contribution is consequently a framework for judging claims, not a universal ranking of architectures. It does not show that one representation, supervision method or evaluation scheme is best in every setting. Its message is that planning quality should be considered alongside safety, route compliance, interaction and reproducibility, with each conclusion matched to the protocol that produced it.
The review also has a built-in blind spot. Because it emphasizes public research with enough technical detail for citation and comparison, it may underrepresent proprietary systems and parts of safety work that are less visible in public papers, including operational validation, fleet monitoring, industrial-scale simulation and management of the conditions in which a system is designed to operate.
The unresolved issues it highlights are practical as well as technical: how to build benchmarks that capture realistic sensors, reactive interaction, rare events, human preferences and scalable testing; how to calibrate uncertainty and multimodal plans; how to connect language reasoning to precise actions; and how to validate world-model reliability when unusual events matter most. The review also points to runtime safety assurance, computing limits and reproducibility as continuing barriers.
For readers trying to interpret autonomous-driving results, the review offers a simple rule: ask what the evaluation actually measures before treating a score as evidence of driving competence. The work provides a conceptual map of methods and tests, but it does not deliver a new measured safety estimate or demonstrate that any system is ready for operational deployment.
The acknowledgment lists support from the Science and Technology Development Fund of Macau, the University of Macau Research Services and Knowledge Transfer Office, and several regional, national and provincial science programs.
Paper data and sources
Original title: Planning-Oriented End-to-End Autonomous Driving: Architectures, Evaluation, and Emerging Paradigms
Authors: Yanchen Guan, Xingcheng Liu, Bin Rao et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text