Graph detection scored higher than sequence decoding on the official CubiCasa5K test, but the comparison did not produce one winner across the study's tested conditions. On 296 corrected scans, graph detection's wall F1, the paper's wall-accuracy score, was 0.818 at a tolerance of 0.05, against 0.790 for sequence decoding. At tolerance 0.015, the scores were 0.787 and 0.736. The graph readout's margins were 2.7 and 5.1 percentage points. Paired 95% bootstrap intervals for those margins ran from 1.4 to 4.0 points and 3.7 to 6.4 points.
The arXiv document is a version 1 preprint dated 26 Aug 2026 by He Zhang, identified as an independent researcher. It asks when autoregressive geometry emission, which produces elements in sequence, or dense heatmap graph detection is the better readout for rasterized floorplan vectorization. In the main comparison, both readouts used the same trained wall-first network. The graph readout had no trained parameter, and its thresholds were calibrated on each domain's validation split.
The lead depended on plan size
Within the CubiCasa5K scans, sequence decoding led by 3.7 percentage points on plans with five to 22 walls. Graph detection led by 2.4 points on plans with 23 to 31 walls and by 8.7 points on plans with 32 to 70 walls. The ordering also changed across image domains: sequence decoding led by 8.2 points on an open-plan synthetic tier and by 5.3 points on finetuned ResPlan data. The paper reports that neither ink nor plan size explained this cross-domain reversal.
The wider evaluation used corrected CubiCasa5K scans, procedurally generated synthetic tiers and the clean vector-render ResPlan-FP benchmark. CubiCasa5K's official split has 3,328 training plans, 318 validation plans and 296 test plans.
The opening heatmap gave a separate result
The study also read an unused opening heatmap without retraining. On a 287-plan evaluation roster, a family-classification heuristic produced opening F1 of 0.644, compared with 0.246 for sequence-emitted openings. Ignoring the family label gave 0.799. The paper describes the first comparison as a 2.6-fold improvement. These figures came from an internal matcher stricter than fpeval and are intended for comparison only with one another.
Room-centric output came with a coverage limit
In a matched-data, matched-recipe comparison on 287 plans, room-centric emission followed by reconciliation had no double walls, wall precision around 0.9 and watertightness scores from 0.92 to 0.94. Its edit cost was about 30% lower than native wall-first emission, and its wall F1 was 3.1 percentage points higher. The paired 95% interval for that lead was 1.3 to 4.9 points.
The coverage limit was specific: walls that bound no room cannot be recovered from room output. In the cited comparison, room-output recall was 0.30 versus 0.66 for the wall-first result, accounting for about four points of aggregate recall. The paper did not indicate a corresponding difference in topological quality.
The edit-cost metric uses a matching-derived weighted script over MOVE, TYPE, CREATE, DELETE and CONVERT operations. The operation magnitudes were set by hand, but the reported ranking did not change anywhere in the tested weight region satisfying the stated ordering.
Fusion outscored its room-centric base
On the 287-plan CubiCasa5K fusion roster, final deterministic output fusion scored 0.853 wall F1 at tolerance 0.05 and 0.811 at 0.015. The room-centric system alone scored 0.781 and 0.750. The reported gains were 7.1 and 6.2 percentage points.
The fused output had an edit-cost score of 73.1, compared with 79.3 for the room-centric output. The paired 95% intervals for the wall-F1 differences ran from 5.9 to 8.3 points and from 5.1 to 7.2 points. The interval for the edit-cost reduction ran from 4.9 to 7.6 edit-cost units.
Input-side conditioning did not beat the unconditioned baseline under a prespecified stop rule. The baseline scored 0.790 and 0.736 at tolerances 0.05 and 0.015. Room-outline, fused-draft and symbolic-prefix deployment forms scored 0.762 and 0.708, 0.768 and 0.720, and 0.756 and 0.696, respectively. Ground-truth-content controls differed from zero-content controls by no more than 0.3 points.
The clean-render benchmark showed different score patterns
ResPlan-FP contains 16,998 retained plans after two over-capacity rejections. Its frozen splits contain 14,998 training plans, 1,000 validation plans and 1,000 test plans, created with seed 42. The benchmark includes id-hash fingerprints, a published rendering protocol, evaluation code and three baseline tracks.
Scores on ResPlan-FP varied across the reported tracks. In zero-shot evaluation, fusion scored 0.688 wall F1, wall-first sequence decoding scored 0.640 and graph readout with frozen thresholds scored 0.621; the reported graph score with calibration was 0.677. Finetuned room-centric and wall-first sequence systems scored 0.928 and 0.968. On a 200-plan zero-shot subset, the language model scored 0.811 at tolerance 0.05 and 0.472 at 0.015.
A conditional result
The overall result is conditional. Graph detection had the higher wall F1 on the corrected CubiCasa5K test, while sequence decoding led on the smallest plans and in the open-plan synthetic tier and finetuned ResPlan comparison. The study does not identify a single readout that wins in every tested condition. Its research question instead treats readout, representation, reconciliation and fusion as choices linked to accuracy and correction effort.
The paper states that its code, benchmark and corrected annotations are available through the listed GitHub repository. No funding statement is reported. The document is an arXiv version 1 preprint dated 26 Aug 2026 and identifies He Zhang as an independent researcher.
Paper data and sources
Original title: When Should a Network Emit Geometry, and When Should It Detect It? Readout, Reconciliation, and Representation in Floorplan Vectorization
Authors: He Zhang
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text