An AI method designed to change a specific fact inside a language model while limiting unrelated changes reported near-perfect edit success and the lowest measured output-distribution drift in one benchmark comparison. The results also showed that spreading an edit to rephrased prompts can come at the cost of preserving unrelated predictions.
In the main CounterFact test on Llama3-8B-Instruct, KLOD recorded 99.70 for Reliability, the study's measure of whether the requested edit succeeds, 47.37 for Generalization, which measures transfer to rephrased prompts, and 44.75 for Locality, which measures whether unrelated prompts retain the pre-edit model's predictions. It also reported 60.77 for Capability and an overall Score of 63.15.
The balance changed with the model and dataset
On CounterFact with Qwen2.5-7B-Instruct, the corresponding scores were 99.60 for Reliability, 68.10 for Generalization, 29.45 for Locality, 50.48 for Capability and 61.91 overall. The different balance between the two backbones shows why the reported operating point cannot automatically be treated as universal.
ZsRE produced a more favorable combination of edit transfer and unrelated-prompt preservation. KLOD reported Reliability of 99.89, Generalization of 86.85 and Locality of 86.47 on Llama3, and 99.92, 87.92 and 78.50 on Qwen. Capability was 61.42 and 56.80, with overall Scores of 83.66 and 80.79. The reported results were high for both Generalization and Locality on this dataset.
The primary evaluation used the same fixed 3,000 edit requests from ZsRE and 3,000 from CounterFact for every method. It tested Llama3-8B-Instruct and Qwen2.5-7B-Instruct and compared KLOD with seven other editing methods. The main run used a target-probability threshold of 0.85.
Keeping the change close to its target
KLOD's recipe is aimed at containing the change. During localized fine-tuning, it places a ceiling on how much probability the new target can gain. At positions where the target is being learned, it preserves the probability distribution of all other possible tokens, excluding the target itself. In the prompt prefix leading toward the target, it preserves the full next-token distribution.
An additional analysis used KL divergence, a measure of how much a probability distribution shifts. In the Llama3 CounterFact comparison, KLOD had the lowest measured KL. At target positions, its Non-target KL was 0.147 for rewrites and 0.579 for rephrases. Prompt KL was 0.031 for rewrites, 0.150 for rephrases and 0.135 for locality prompts. The authors associate these lower values with suppressing unnecessary distributional drift rather than simply weakening the edit.
The preservation terms mattered in the component tests. With the full objective, CounterFact Locality was 44.75. Removing the prefix-preservation term while retaining the target-position non-target term left 41.30, but removing the latter dropped Locality to 9.15. Removing both terms produced 2.02, while the standard cross-entropy comparator reached 1.87. Reliability stayed above 99% in every variant.
The result also held in a more closely matched comparison. At matched Generalization, KLOD's position-wise logit-odds hinge recorded Locality of 44.75, versus 37.50 for sequence-level thresholded cross-entropy. Reliability was nearly identical, at 99.70 versus 99.80, while Generalization was 47.37 versus 46.20.
A control with a cost
The target threshold acts as a control, but not a free improvement. In the Llama3 CounterFact setting, alpha 0.85 produced Generalization of 47.37 and Locality of 44.75. At alpha 1.0, Generalization rose to 62.35, while Locality fell to 19.48 and Capability to 57.65. The reported curve also showed low Reliability below 0.4, while most edits succeeded above 0.5.
The Locality result was not confined to one random seed. Across seeds 1, 42 and 100, KLOD's mean Locality was 46.51 on CounterFact, with a standard deviation of 1.98, and 86.38 on ZsRE, with a standard deviation of 0.33. The reported baselines stayed near 2 on CounterFact and roughly 51 on ZsRE, and the lower CounterFact Generalization profile persisted.
When edits were applied sequentially at larger scale, Locality fell as the edit count grew, but KLOD remained above the reported comparators. On Llama3 ZsRE, its Locality was over 85 at 3,000 edits and around 60 at 150,000. At 150,000 edits, LocFT-BF was roughly 30 and UltraEdit roughly 10, while KLOD's Reliability remained high and its Generalization competitive. The scaling values were approximate, with no exact intervals reported.
What the tests leave open
The headline scores came from token-level accuracy under teacher forcing, while a smaller free-running check tested what happened when the model generated its response on its own. That check covered 500 CounterFact cases using greedy decoding and exact-match scoring. At alpha 0.85, Reliability exact match was 99.60, Generalization exact match was 48.00 and Capability was 60.77. At alpha 1.0, Reliability stayed at 99.60, Generalization rose to 61.60 and Capability fell to 57.65.
On 3,000 WikiBigEdit cases with Llama3-8B-Instruct, KLOD reported Reliability of 99.75, Generalization of 93.11, Locality of 56.30 and Capability of 61.30. The additional evaluation reproduced the high edit-success pattern while retaining the threshold-dependent balance between Generalization and Locality.
The results do not amount to a blanket guarantee. CounterFact Locality remained limited in the 3,000-edit setting, and the large-scale run still saw KLOD Locality fall from over 85 at 3,000 edits to around 60 at 150,000. The evidence is concentrated on the tested datasets, models and editing protocols, and token-level measures do not establish preservation of broad model behavior during generation. The free-running results reported Reliability, Generalization and Capability, but not a free-running Locality score.
The work is an arXiv v1 preprint dated 28 August 2026. It states that code is available on GitHub and reports support from a National Research Foundation of Korea grant funded by the Government of Korea's MSIT, grant No. 2022R1A2C1005316.
Paper data and sources
Original title: KLOD: Locality-Preserving Knowledge Editing via Non-Target Distribution Preservation
Authors: Hojun Jeong, Gyunyeop Kim, Sangwoo Kang
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-28
DOI: Not available
Original paper · Full text