Preprint

Preprint reports stronger cooperative-driving results with lower communication cost

G-MARK is reported to outperform V2V-GoT on eight of nine benchmark entries, with the largest gains on tasks requiring explicit cooperative evidence.

An arXiv preprint reports that G-MARK outperformed the V2V-GoT baseline on eight of nine benchmark entries, with the largest gains on tasks requiring explicit cooperative evidence.

In occlusion reasoning, G-MARK reached an F1@0.5m score of 0.428, compared with 0.301 for V2V-GoT, a reported 42.2% improvement. F1 is a localization-aware score based on precision and recall, the measures used for the study’s object and visibility tasks.

Hidden-object discovery produced an F1@0.5m score of 0.494 versus 0.440, a reported 12.3% improvement. On Notable Objects, G-MARK scored 0.586 versus 0.525 for V2V-GoT; on Planning Awareness, it scored 0.614 versus 0.608. The reported improvements were 11.6% and 0.9%, respectively.

A structured record of shared evidence

G-MARK constructs a provenance-aware knowledge graph from processed multi-agent scene artifacts for downstream reasoning. In practical terms, the graph links observations with information about where the evidence came from.

The pipeline creates local evidence graphs, conservatively associates compatible observations across connected vehicles and retains weak unmatched candidates for later reasoning.

The evaluation uses V2V4Real and V2V-GoT-QA. The official split contains approximately 110,000 training questions and 31,000 validation questions; reported results use the validation split.

Motion gains, with one exception

Motion results also favored G-MARK. On Object Motion Q5, its L2 average was 3.822 versus 8.050 for V2V-GoT, a reported 52.5% improvement; on Q7, it was 3.822 versus 7.610, a reported 49.8% improvement. L2 is an error measure, so lower values indicate closer predictions.

Agent Motion accuracy was 0.905 versus 0.874 for V2V-GoT, a reported 3.5% improvement. For Control Settings, Action L1—the normalized action error used for control—was 0.076 versus 0.088, a reported 13.1% improvement.

Future-trajectory forecasting was near parity: G-MARK’s average L2 error was 2.710 metres, compared with 2.620 metres for V2V-GoT.

In the future-trajectory communication comparison, G-MARK used 0.0159 MB per sample—about 25.2 times lower than intermediate-fusion methods and 25.6 times below V2V-GoT, according to the report.

What the component tests showed

In a fixed-head, construction-level ablation, removing partner observations produced Invisible Objects F1 of 0.000, versus 0.494 for full G-MARK.

The no-provenance variant produced Invisible Objects F1 of 0.396 versus 0.494 for full G-MARK, and Control Settings Action L1 of 0.152 versus 0.076. Replacing the knowledge graph with an unstructured object-level representation produced 0.443 versus 0.494 on Invisible Objects and 0.089 versus 0.076 on Action L1.

A benchmark result, not a road test

The results come from the validation split and from processed perception artifacts rather than raw camera or LiDAR streams, so they do not assess upstream perception quality.

This is a benchmark performance comparison, not evidence of improved road safety or readiness for live multi-vehicle deployment. The report gives no confidence intervals, standard errors or hypothesis tests for the differences.

The primary V2V-GoT comparison relies on reported baseline results; the available analysis does not establish identical reruns or matched implementation details. The document is an arXiv preprint dated 20 Aug 2026.

Paper data and sources

Original title: G-MARK: Grounded Multi-Agent Reasoning for Cooperative Driving via Knowledge Graphs
Authors: Bhavya Gupta, Onat Gungor, Tajana Rosing
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.