A computer simulation of a freight corridor reported a clear trade-off in its main comparison. The Cognitive scenario recorded 98.27 completed trips per hour, against 77.38 for the fully manual Baseline, while its congestion index was 0.15 rather than 0.42. Yet Cognitive also had the highest total CO2 emissions proxy of the three scenarios. The figures are model outputs, and the emissions measure is a proxy rather than a direct measurement.
Three control tiers in the model
The paper is labeled a preprint and identified as arXiv:2608.25193v1. It evaluates a computer framework that combines vehicle-to-everything communication, or V2X, with reinforcement learning and multi-agent reinforcement learning to model adaptive platoon formation and charging coordination across Baseline, Assisted and Cognitive scenarios.
Baseline used 100% manual trucks without V2X. Assisted used 70% assisted and 30% manual trucks, with rule-based platooning and an assignment rule that sent trucks to the least-loaded charging station. Cognitive used 70% autonomous and 30% assisted trucks, a trained corridor policy, and five station agents using multi-agent reinforcement learning for pricing and charging priority.
Each scenario run used a 1,500-truck simulated fleet, so 1,500 was the fleet size per run rather than a pooled count across scenarios. The corridor contained 20 road segments, five charging stations and four terminals. Each episode covered 125 steps, or 10 simulated hours. Scenario comparisons used five replications per scenario, and disruption tests used five replications per condition.
The carbon result was less favorable
Relative to Baseline, the paper reports energy use per kilometer 7.5% lower for Assisted and 7.9% lower for Cognitive. It reports Cognitive's travel-time measure as 7.2% lower than Baseline. These results were averaged over five replications per scenario, but the paper does not define the accompanying plus-or-minus terms as standard deviations, standard errors or confidence intervals.
Total CO2 proxy emissions were 101,436.47 kilograms in Baseline, 102,404.27 kilograms in Assisted and 117,673.60 kilograms in Cognitive. Per-kilometer emissions were reported as comparable across the scenarios. The calculation used an assumed grid carbon intensity of 0.4 kilograms of CO2 per kilowatt-hour, so the totals are model outputs rather than direct emissions measurements.
Stress tests within the model
Demand testing varied origin-terminal demand multipliers from 0.25 to 2.0, with five replications for each condition. Cognitive travel time stayed near 1.34 hours, ranging from 1.31 to 1.36 hours. Its stranded-truck rate ranged from 0.001 to 0.064, compared with 0.019 to 0.242 for Baseline, and its congestion advantage widened at high demand.
Charging capacity was varied from five to 25 ports in increments of five. Cognitive throughput remained between 98 and 101 trips per hour. Baseline throughput rose from 73.5 trips per hour at five ports to 77.4 at 15 ports, while Assisted rose from 79.0 to 82.7. Queue waiting time fell as more ports were added.
Disruption tests were zero-shot, meaning the controllers were not retrained after accidents, weather events, station failures or demand spikes. The robustness score compared disrupted performance with baseline performance. Cognitive's scores were 1.00 for accidents, 1.00 for weather, 0.99 for station failure and 1.31 for demand spikes. Assisted scored 1.02 for station failure and 0.00 for accidents, where its reported travel time rose by 155%. On the study's measure, Cognitive had the strongest reported results for accidents, weather and demand spikes, while Assisted was strongest for station failure.
What the model leaves open
The learning setup used 125-step episodes at 0.08 hours per step, a replay buffer holding 25,000 transitions, a 200-by-125-step training horizon, and learning-rate decay of 0.98 with a floor of 10^-4.
Across 200 Cognitive training episodes, combined reward ranged from 442 to 2,038 and peaked at 2,038 in episode 77. The multi-agent learning component moved from -1,378 in the first episode to -386 in the 200th, a change the paper reports as a 72% improvement. Energy stayed near 1.11 kilowatt-hours per kilometer, while throughput stabilized near 1,876 completed trips per episode, with a range of 1,757 to 2,005. These are training-curve observations, not held-out or field-validation results.
In a separate benchmark of 10,000 forward passes, a corridor policy action took 3.8 microseconds, the five station agents' two communication rounds took 42.9 microseconds, and combined decision latency was 47 microseconds per simulation step. The paper compared that latency with a five-minute control interval, or 300,000 milliseconds, and described it as four orders of magnitude below the interval.
The authors say the current network's unidirectional design limits transferability. They call for tests on bidirectional and multipath topologies, broader policy-transfer tests, reward ablations and extended training, comparisons with optimization benchmarks, additional replications with paired statistical tests, and calibration to real freight-demand data.
Funding and disclosures
The work was supported in part by Panama's National Secretariat of Science, Technology, and Research through the IFARHU-SENACYT Scholarship Program. The authors state that the views are their own, do not necessarily reflect Amazon.com, Inc. or its affiliates, and that the work was conducted independently of their institutional roles.
Paper data and sources
Original title: Simulating Cognitive Smart Freight Corridors with Agent-Based Models and Reinforcement Learning
Authors: Madelaine Martinez-Ferguson, Chun Wang, Mustafa Can Camur, Xueping Li
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-25
DOI: Not available
Original paper · Full text