In its most complete benchmark, the study reports that PI-SAP had lower mean solution error than NTK-SAP at four of five tested Gray–Scott pruning levels. PI-SAP ranks parameters by their sensitivity to the PDE residual, while NTK-SAP uses output-side training dynamics. At 50% pruning, however, NTK-SAP was slightly lower. Across the other equation tests, the preferred rule changed with sparsity, network width and the convection parameter, so the reported comparisons do not identify one rule as best in every setting.
The reported comparison holds the PirateNet architecture, eligible parameter set and mask-enforcement procedure constant, leaving the saliency calculation as the stated difference. The study asks which notion of weight importance should be preserved when pruning a physics-informed neural network, or PINN. NTK-SAP targets output-side training dynamics, whereas PI-SAP assigns saliency from sensitivity of the PDE residual. Here, saliency is the parameter-ranking score used to identify important weights.
Gray–Scott provided the clearest comparison
Gray–Scott was the most complete benchmark; complex Ginzburg–Landau, Burgers’ and convection were included for cross-PDE validation. The Gray–Scott models used PirateNet with three residual adaptive blocks, width 256, Swish activation, 256-dimensional Fourier features, Adam optimization, GradNorm weighting and causal time marching over 10 temporal windows. The comparison covered pruning levels of 10%, 30%, 50%, 70% and 90%.
PI-SAP’s mean solution error was 7.5343 × 10^-3 versus 1.1439 × 10^-2 for NTK-SAP at 10% pruning, and 1.7555 × 10^-2 versus 1.7896 × 10^-2 at 30%. The ordering held at 70%, with 1.9430 × 10^-1 for PI-SAP versus 2.7291 × 10^-1 for NTK-SAP, and at 90%, with 3.7866 × 10^-1 versus 4.3493 × 10^-1. At 50%, NTK-SAP was lower, at 4.3067 × 10^-2 versus 4.4727 × 10^-2 for PI-SAP. Relative to NTK-SAP, the reported changes for PI-SAP across 10%, 30%, 50%, 70% and 90% were -34.1%, -1.9%, +3.9%, -28.8% and -12.9%.
The residual comparison was more consistent. PI-SAP had lower mean Gray–Scott PDE residual at every pruning level, with the largest reductions at 10% and 50%. Reported runtimes stayed between 16.4 and 18.0 hours, with no systematic wall-clock advantage from masking.
At 70% pruning, a separate terminal high-frequency check favored PI-SAP over NTK-SAP for both predicted fields. The error for u was 0.3591 versus 0.7238, and for v it was 0.4132 versus 0.7749. Dense PirateNet values were 0.0227 for u and 0.0277 for v, so PI-SAP’s edge was relative to NTK-SAP rather than to the dense baseline.
The score used changed the result
An additional Gray–Scott diagnostic changed the ordering again. At 70% pruning and 600 K total steps, original NTK-SAP had the lowest pruned field errors, at 0.1786 for u and 0.3229 for v. PINN-block SAP had the lowest pruned residual losses, 1.605 × 10^-4 for ru and 5.671 × 10^-5 for rv, but the highest field errors, 0.2249 for u and 0.4021 for v. Conditioning-aware NTK-SAP was worse than original NTK-SAP on both field errors, at 0.1966 and 0.3516 compared with 0.1786 and 0.3229.
The diagnostic put solution accuracy, residual loss and conditioning on different sides of the ranking. The method with the lowest field error was not the method with the lowest residual losses.
Other equations produced a different pattern
For the complex Ginzburg–Landau equation, both pruning methods had lower error than dense at 10% pruning. NTK-SAP had lower error at 30% and 50%, while PI-SAP had lower error at 70% and 90%. At 70%, the reported mean errors were 5.8078 × 10^-2 for PI-SAP versus 6.5615 × 10^-2 for NTK-SAP, a relative difference of about 11.5%. At 90%, the figures were 1.7434 × 10^-1 versus 1.8836 × 10^-1, a difference of about 7.4%.
On Burgers’ equation, the best-pruned error was below the dense error at every tested width: 3.7532 × 10^-3 versus 1.0849 × 10^-2 at width 128, 5.4160 × 10^-3 versus 3.1485 × 10^-2 at width 256, and 2.5324 × 10^-3 versus 3.0051 × 10^-1 at width 512. PI-SAP was selected at widths 128 and 256, both at 30% pruning, while NTK-SAP was selected at width 512 at 90% pruning. These results were averaged over seeds 0 through 4, but only the best pruned result for each width was reported.
The convection comparison was similarly dependent on setting. NTK-SAP was selected at β = 1 with width 512, β = 5 with width 128, and β = 10 with width 256. At width 128, NTK-SAP also had lower mean relative L2 error than dense at β = 15, 2.5010 × 10^-2 versus 5.3298 × 10^-2 at 70% sparsity, and at β = 20, 7.9886 × 10^-2 versus 1.3302 × 10^-1 at 50% sparsity.
What remains uncertain
Seed coverage differed across the benchmarks. Burgers’ and convection results were averaged over five seeds, whereas Gray–Scott and complex Ginzburg–Landau were treated primarily as completed benchmark runs rather than large multi-seed sweeps.
The PINN-block diagnostics used small-batch approximations rather than full training-set NTKs, and their coefficients were not extensively tuned. The authors propose jointly conditioning solution and residual blocks in future objectives.
The work is identified as arXiv:2608.25564v1, dated 26 Aug 2026. The document reports author affiliations but no funding source or conflict-of-interest statement.
Taken together, the reported tests describe a setting-dependent result. PI-SAP had the more consistent Gray–Scott residual ranking, but solution accuracy changed at 50% pruning and the other equation and diagnostic tests showed different rankings. No universal pruning rule emerges from these reported benchmarks.
Paper data and sources
Original title: Physics-Informed Foresight Pruning for Sparse PINN Solvers of Nonlinear PDEs
Authors: Ahmad Ishaque Karimi, Uvini Balasuriya Mudiyanselage, Kookjin Lee
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text