A preprint review of neural-computation research finds a mismatch between the variety of ways models process information and the narrower set of ways they are trained. In an audited set of 32 architecture–learning configurations, every forward state-dynamics class was represented, but Global gradient appeared in 16 configurations, compared with 7 for Approximate or implicit gradient.
Credit assignment—the route by which a learning signal reaches a system’s internal parts—is the review’s term for the backward side of that divide. The authors call the mismatch the forward–backward disconnect, and say the largest demonstrated scales in the surveyed evidence are concentrated in global or closely gradient-derived error-propagation methods.
The ledger shows the split
The review counted a concrete architecture paired with a concrete learning method, and each pairing needed support from primary evidence. The taxonomy recorded state-dynamics structure, credit-assignment mechanism and separate biological-grounding ratings for the forward operation and the learning rule.
The search covered ACM Digital Library, IEEE Xplore, Scopus, arXiv and Google Scholar, with citation chaining; the last comprehensive search was in June 2026.
The audited ledger contained 32 configurations: 6 Static, 14 Discrete-time, 6 Continuous-time, 4 Implicit and 2 Hybrid event-driven. Those counts describe the selected representative set, not the prevalence of these approaches across the full literature.
Looking inside the backward categories, ordinary reverse-mode backpropagation alone accounted for 10 of the 32 configurations—more than the entire Approximate/implicit category at 7. That makes the imbalance visible at the level of individual mechanisms, not just broad groupings.
Scale brings the trade-off into view
The survey’s evidence-bounded scaling synthesis points in the same direction. The largest demonstrated scales were associated with global or closely gradient-derived and error-propagation mechanisms, while strictly local and mechanistic-plasticity rules remained substantially less scalable in the audited evidence.
On one cited ResNet-18 ImageNet comparison, PFA reached 68.46% top-1 accuracy, PFA-o 69.30% and backpropagation, or BP, 69.69%. The results put the approximate-gradient methods close to backpropagation on that benchmark.
Spiking networks reveal a second divide
Spiking networks make the distinction between forward behavior and learning behavior especially clear. Surrogate-gradient and STDP-trained spiking networks share the same LIF forward operator—the rule used to generate their spikes—but occupy Weak versus Mechanistic plasticity tiers on the learning axis.
The review cites QKFormer as a high-scale spiking example: 85.65% top-1 accuracy on ImageNet-1k with 64.96 million parameters, using spiking Q–K token attention. But the example used surrogate-gradient training, so it adds to the evidence for scalable spiking computation without resolving the question of scalable local learning.
A separate local-plasticity example was an unsupervised STDP-trained LIF network that reached 95% MNIST accuracy with 6,400 excitatory neurons. The survey says local spiking rules have not shown the same combined scale, task breadth and accuracy in its audited evidence.
The divide also appears between training and deployment. Large spiking systems are typically trained off-chip on GPUs with surrogate gradients and deployed on neuromorphic hardware for inference; the survey identifies no competitive local-learning bridge at deep-network scale.
The evidence has a narrow frame
The counts should not be read as a census of neural-computation research. They are descriptive of an audited representative set, with coverage conditioned on the databases used, the scope rules, the search cutoff and the classification criteria.
Nor does the review establish that concentrated credit assignment causes the scaling pattern. Its scale conclusion is an evidence-bounded narrative across selected studies rather than a pooled comparison, and no confidence intervals or inferential uncertainty estimates were reported.
That leaves the review’s central problem in plain terms: varied ways of computing have outpaced the scalable learning mechanisms demonstrated in the audited evidence.
Paper data and sources
Original title: The Forward-Backward Disconnect: State Dynamics, Credit Assignment, and Biological Grounding in Neural Computation
Authors: Hadi Al Mubasher, Mariette Awad
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text