Preprint

Offline AI shows gains in simulated multi-drone fleet tests

Preprint: An arXiv version-1 study dated 26 Aug 2026 reports higher return, communication and sensing measures, plus lower collision-risk counts in modeled multi-UAV tests.

An offline learning method for teams of unmanned aerial vehicles reported a 29.3% higher episode return than TD3+BC in the random-mobility evaluation. It also reported higher communication sum rate, sensing pass rate and sensing margin, while the collision-risk count was 54.2% lower. The method, STAR-CRDT, is part of AERIS, a fixed-log approach designed to improve policies while keeping execution distributed. The document is an arXiv version-1 preprint dated 26 Aug 2026.

Learning from a fixed record

AERIS uses centralized training with decentralized execution. During training, the system uses logged global information, while each UAV acts from its own local history. STAR-CRDT makes support-aware local action corrections and passes only trusted improvements into the decentralized actor. In plain language, the method is designed to change decisions the record can support while preserving local decision-making during execution.

TD3 was used to construct the behavior log, and all reported AERIS policy improvement came from offline learning. The authors say no public flight-log dataset with joint communication, sensing and safety labels was available, so they collected a fixed behavior dataset before offline training. The reported setup used three UAVs, six users and three targets. The log contained 1,000 episodes and 200k transitions, while random-mobility evaluation used 50 held-out episodes. Fleet-scale tests varied the number of UAVs from three to six.

The gains were uneven

The study defined episode return as the sum of per-slot reward over the episode horizon. In the random-mobility comparison, STAR-CRDT's return was 29.3% higher than TD3+BC. Communication sum rate was 3.4% higher, sensing pass rate was 4.8% higher and sensing margin was 69.1% higher. Collision-risk count was 54.2% lower.

An ablation comparison pointed to a difference between the method's correction stages. CRDT stayed closer to the behavior log but left a higher violation cost. The reported explanation was that critic regularization alone did not identify which local correction would be globally beneficial. STAR-CRDT combines support-aware rectification with trust-gated distillation, passing trusted improvements to the decentralized actor.

The comparison widened, but stayed modeled

The pattern continued in zero-shot system-scale transfer. STAR-CRDT ranked first at every tested fleet size from three to six UAVs. At six UAVs, its return was 39.3% higher than TD3+BC and 60.8% higher than CRR, while it retained the best sensing pass rate and sensing margin.

The researchers also tested road-map deployment protocols alongside random mobility. Across six scene-scale settings, STAR-CRDT had the best return in every setting. It was best in 24 of 30 reported metric rows and in the top two in 29. With six UAVs, return was 36.7% higher in the dense-grid scene and 132.1% higher in the sparse-road scene than the strongest baseline.

A guarantee tied to the data

The paper's formal guarantee is limited to the offline-support distribution. It is an offline-support policy-improvement guarantee for the trust-gated correction, and the proof does not claim global optimality outside the dataset. The guarantee depends on stated conditions involving local action rankings, regularity and distillation error.

The evidence comes from the reported random-mobility and road-map evaluation protocols. The road-map transfer covered the reported fleet settings, and no physical UAV flight experiment is reported in the supplied material. These results therefore do not validate operation in physical flights.

Baseline comparisons used the same dataset, training budget, model capacity where applicable and evaluation episodes. Questions remain about whether the reported gains hold with different behavior-log quality, mobility distributions, geographies or larger fleets, and how the guarantee behaves when its local-support or critic assumptions fail.

Paper data and sources

Original title: AERIS: Offline Policy Improvement for Multi-UAV Integrated Sensing and Communication
Authors: Ziyuan Wang, Yifan Sui, Wei Wei et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.