Preprint

Targeted Checks Cut AI Errors From Outdated Memory

An arXiv preprint found that directing one of two verification checks to a critical provenance path sharply improved decisions in controlled model tests.

In controlled tests, directing one of an agent’s two verification checks to the record behind an outdated inherited constraint raised current-record-consistent decisions from 34 of 150 to 145 of 150 in the primary run — a 74.0-point increase. Here, “current-record-consistent” means the model’s second-turn action matched the archive’s current record, rather than the stale rule it had inherited.

The study asks what happens when a constraint that was once valid is carried forward after a newer record supersedes it. It tests whether the resulting mistake can be avoided without adding to a scarce verification budget. The policies were assigned at a fixed budget of two records: native allocation, one slot redirected to the critical provenance path, or one slot redirected to a random non-critical record.

The verification gap

One clue appeared before the final decision. When the constraint was stated in memory, the target provenance path was named in 181 of 900 primary episodes, or 20.1%, and in 208 of 900 fresh-wording replication episodes, or 23.1%. When the constraint was removed from the wording, those rates were 66.9% and 72.9%. In these runs, the relevant trail was selected much less often while the constraint remained in memory.

Under native allocation in the superseded-world episodes, the models followed stale memory rather than the archive’s current record in 116 of 150 primary episodes — 77.3%. The same pattern appeared in 112 of 150 fresh-wording replication episodes and 112 of 150 procurement-world episodes, both 74.7%. These were deterministic decisions: each was classified by whether the second-turn action matched the current record.

The policy effect held across the runs

That gap was large when the critical path was forced into the two-record allocation. In the primary run, current-record-consistent decisions rose from 34 of 150 under native allocation to 145 of 150 under forced-critical allocation, a 74.0-point increase, with a 95% interval from 68.0 to 80.0 points. In the fresh-wording replication, the corresponding figures were 38 of 150 and 147 of 150, a 72.7-point difference with an interval from 66.7 to 78.7 points.

The held-out procurement run was less dramatic but still substantial: 38 of 150 native decisions matched the current record, compared with 130 of 150 when the critical path was forced, for a 61.3-point difference. After an inconsistency in that scenario was corrected, the held-out robustness run produced 36 of 150 versus 146 of 150, a 73.3-point effect; native allocation still followed stale memory in 76.0% of episodes.

To gauge uncertainty, the analysis repeatedly resampled episodes within each model and policy 4,000 times to build the 95% ranges, and used a model-stratified test as corroboration. When the fetched record confirmed the stated constraint in valid-world episodes, forced-critical allocation changed the outcome by only 0.7, 2.0, 0.7 and 0.0 points across the four runs. Those small changes stood in sharp contrast to the effects in the superseded-world runs.

A result with a narrow frame

The original held-out procurement scenario carried a timing inconsistency. Its situation text said the contract expired in three days, while the source record said onboarding took six weeks and an earlier turn-1 text said 14 days — implying 11 rather than three days remaining. A corrected held-out robustness replication then produced 146 of 150 current-record-consistent decisions under forced-critical allocation versus 36 of 150 under native allocation, a 73.3-point effect.

The confirmatory material consisted of 5,400 episodes across six language models: 1,800 primary episodes, 1,800 fresh-wording replication episodes, 900 original held-out episodes and 900 corrected held-out robustness episodes. The fixed design compared policies within a two-record verification budget, so the reported effects describe that test setup.

The work is identified as an arXiv preprint, version 1, dated 26 August 2026. It separates pre-run specification commitments from a post-run OSF archive and explicitly says that the archive is not a preregistration. The paper also says all 5,400 episode files, frozen specifications and manifests, analysis scripts, independent recomputation scripts and the generator for the reported numbers were released.

Paper data and sources

Original title: When Stale Constraints Go Unchecked: Budgeted Verification Failures in Inherited Agent Memory
Authors: Kazuki Nakayashiki
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.