Preprint

New York audit flags gaps in predictive lead-line records

An arXiv preprint says New York City's model-cleared records showed no lead or Unknown entries, but the estimate is not a direct accuracy test.

A statewide audit of New York's public lead-service-line records found seven localities where predictive-model classifications showing only one material conflicted with the same locality's physical verification. Six of those cases went beyond what the study said could be explained by ordinary sampling, although the audit did not treat a uniform output by itself as evidence of misconduct.

The most detailed finding came from New York City. The predictive model was listed as the basis for 43,215 addresses, and every one was recorded as Known Other; the database recorded neither lead nor Unknown in that group. By contrast, the city recorded Unknown for 121,779 addresses and lead for 120,692 under its other classifications. A statistical calculation based on zero events put the 95% upper bound for the model group's recorded lead rate at 0.0085% - a bound on the filed record, not an estimate of the model's accuracy.

A pattern that drew scrutiny

The study screened 153 reporting units covering 221,090 model-classified addresses. It first looked for localities with at least 100 model-classification records but only one recorded value, then checked whether those localities had at least 200 physically verified lines with a lead rate of at least 1%. Seventy-five localities met the single-value screen, covering 125,990 addresses, or 57% of the screened model-classified lines. Sixty-eight were consistent with verification or did not have enough information to test; seven were contradicted.

One example was East Rochester, where 472 model-classified addresses were all marked Known Other. Among 557 verified lines in the locality, 9.69% were found to be lead, a rate that would correspond to 46 findings; the study reported a zero-event probability of 10 to the minus 21 if that local rate applied to the model group.

The audit compared two public snapshots of the state inventory: one dated 11 August 2026 with 4,618,115 rows and an earlier capture dated 22 June 2025 with 3,747,025 rows. In the archived New York City data, the public-side verification method and material matched customer-side values on 100.0% of rows. The earlier snapshot had 817,375 city addresses with blank public columns and 43,440 model classifications visible only on the customer side; 14 months later, those addresses carried the same public value.

New York City's older buildings complicate the picture

The city's model-cleared addresses were newer than the physically verified population: their median construction year was 1984, compared with 1930, and 77% were in buildings constructed after 1960, compared with 24% of the physically verified group. But after comparisons were made within construction eras, records-based lead rates ranged from 4.32% to 31.85%, physical-verification rates from 1.45% to 14.45%, and the predictive-model rate remained zero in every era.

Among the model-cleared addresses, 7,782, or 18.0%, were in buildings built before 1940. Physical verification in that age group found lead at 12.93%, based on 98,749 addresses, so applying that rate implied roughly 1,000 Known Other lines. The installation-or-replacement-date field was blank for all of those model-cleared addresses, and building age is only a proxy for the age of a service line.

Because only 66 address keys statewide, and just one in New York City, had both a model record and a physical record, the study did not report a direct accuracy figure. Instead, six era-aware estimating approaches produced 1,152 to 1,461 expected lead lines among the 43,215 model-cleared addresses, or 2.7% to 3.4%, conditional on the study's identifying assumption; the authors summarized that as roughly 1,150 to 1,450 lines.

That range is a spread between specifications, not a confidence interval. A 400-draw bootstrap put the sampling error for a single estimator at roughly 3%, while the analysis identified residual bias from nonrandom physical verification and conditional exchangeability as the larger sources of uncertainty.

What the audit can - and cannot - say

The findings describe consistency in published records, not the internal behavior or accuracy of a vendor's model. Physical checks were not randomly selected, and excavation and field inspection were different methods. The study therefore treats the comparison as a population-level audit rather than a paired prediction-versus-excavation accuracy test.

The result was sensitive to how physical verification was defined. Excavation found lead on 10.13% of lines versus 8.48% for field inspection, while Unknown material appeared on 2.60% versus 0.01%. Using excavation alone left the same seven localities contradicted, raised the pre-1940 physical lead rate to 13.35%, and widened the estimated range to roughly 850 to 1,325 lines.

As a separate check, an open gradient-boosted-tree baseline using coordinates and building type showed moderate spatial discrimination between lead and non-lead records in 165,063 physically verified New York City addresses, including 15,154 lead addresses. Its no-ZIP cross-validation AUC ranged from 0.639 to 0.736, but it was not the audited vendor model and does not establish that model's accuracy.

The preprint's authors call the combination of uniform predictive outputs and contradictory utility verification an audit red flag. They say the expected-count estimate depends on assumptions and is not proof that every Known Other line is lead or that every model-cleared address is misclassified. The paper says its inputs are public and that code, source files and checksums accompany the submission.

Paper data and sources

Original title: Auditing Recorded Predictive Lead Service-Line Classifications Against Physical Verification: A Statewide Study of New York
Authors: Muhammad Sarmad Sohail
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published after independent verification and editorial approval.