The central finding is a limit, not a performance gain. When a supervised-learning system receives fresh observations from independently drawn units, with no observed identity or linkage, a world in which all units share one response pattern can produce exactly the same observable data as a world in which units respond differently. The analysis says that indistinguishability holds for every sample size. In that setting, every randomized test has total error of 1, while the lowest possible worst-case error is one-half.
The unit is part of the model’s question
The document is an arXiv version 1 preprint, with the supplied version line dated 25 Aug 2026. It is a formal theoretical analysis over abstract populations of persistent units and linked events, with no empirical study reported. Its central proposal is to make the persistent referent behind linked events, the enduring entity those records belong to, an explicit unit primitive declared by the task.
That move forces apart two ideas that are easy to conflate. A homogeneous unit-response world means the probability rule for responses is the same across units. A learner may also appear unit-insensitive simply because it has left unit information out. Those are different explanations, and row-level fit alone cannot choose between them.
To represent the unit explicitly, the proposed learner uses a contextual unit token and one shared form for the response law. Structured differences between units enter through a declared simple relation in that token, with a linear predictor serving as the running example. But this is an interface, not a guarantee that training will recover a unique representation. If token dimension or the readout is left unrestricted, the factorization can become a reparameterization, and joint training need not identify either module.
What the formal results can reveal
The framework distinguishes how the token is obtained. Direct access removes uncertainty about which unit is involved. The paper calls the second route unit abduction: forming a same-typed contextual token from factual evidence when no resolver identifies the unit. If identity is unresolved, the exact target is a mixture of the response laws for the possible units, whereas the learner builds a composition in token space. The two agree only when the tokenizer and shared response form realize the relevant targets.
That distinction gives the paper a way to describe the value of identity information. Under its log-loss framework, perfect unit knowledge is an oracle ceiling. Factual evidence exposes only the portion of unit information that is relevant to the response; any response-relevant information still hidden at the evidence cutoff remains inaccessible in that setup. The result separates information that the available evidence can support from information that would require access to the unit itself.
For deployed predictions, the analysis represents excess error with a KL discrepancy, a comparison between probability assignments. It bounds that discrepancy using two sources: error in the system’s belief about the unit and error in the fixed-unit response law after averaging over that belief. A separate stability result says the effect of unit-belief error at a given query is scaled by how different the true fixed-unit laws are at that query. If those laws agree there, a change in belief does not change the response mixture.
Linked observations provide information that isolated rows do not. In a restricted binary witness, trusted pairs of observations from the same unit, assumed conditionally independent once that unit is fixed, separate the alternatives through different agreement probabilities. The broader linked-pair result is more limited: it identifies a variance or covariance moment, not the full mixing law and not the realized unit. The separation depends on trusted linkage and the conditional-independence assumption.
Evaluation changes with the unit
The distinction also changes what counts as a fair evaluation. Row-weighted and unit-weighted empirical risks agree for every possible pattern of unit losses only when the observed units have equal record multiplicity. A record-wise split can mix new events from units already represented in training with cases from units absent from training. Claims about new-unit generalization therefore require an explicitly unit-disjoint test.
Because it reports no empirical study, the preprint offers no predictive-performance comparison on real datasets. It also does not establish that a learned token recovers identity or is unique. The shared-token design organizes the problem, but the factorization itself does not establish learnability or module identifiability.
The paper’s practical message is to declare the unit, the evidence available to infer it and the deployment question together. A system evaluated on new events for known units is answering a different question from one tested on units it has never seen, and the framework makes that distinction explicit.
Paper data and sources
Original title: Toward Machine Learning with the Unit as a Primitive: Learning from Unit-Linked Events
Authors: Heyang Gong
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-25
DOI: Not available
Original paper · Full text