Preprint

Sparse sampling linked to slower functional linear regression rates

Preprint: A pooling-ridge estimator is paired with matching bounds and tested in simulations and two datasets.

The central result is theoretical: in scalar-on-function linear regression, the reported prediction-risk rate, or how prediction error falls as data grow, has two contributions. One comes from the number of subjects, while the other comes from the total number of discrete observations collected from predictor curves. Matching upper and lower bounds support a minimax-optimal rate under the paper's assumptions. In this framework, sparse sampling corresponds to a slower rate than sufficiently dense sampling.

The work covers two forms of functional linear regression. In scalar-on-function regression, one curve is used to predict a scalar outcome. In function-on-function regression, one curve is used to predict another. In both cases, the data are discrete and noisy, and the theory spans sparse-to-dense sampling designs.

Pooling the observations

To work with these measurements, the proposed pooling-ridge framework pools discrete observations across subjects to construct unbiased estimates of the operators used in reproducing kernel Hilbert space, or RKHS, regression. RKHS regression is a regularized framework for fitting flexible functions. The estimated operators are then plugged into a unified estimator.

Where sparse and dense designs meet

Matching upper and lower bounds support what statisticians call a minimax-optimal rate: the proposed convergence rate is matched from above and below over the paper's distribution class, rather than being only a report on one estimator. The result is asymptotic, meaning it concerns behavior as sample sizes grow, and it is not a finite-sample confidence guarantee.

A single transition separates the scalar problem's sparse and dense regimes. With sparse predictor sampling, the total number of discrete observations is the slower part of the convergence rate. When sampling becomes sufficiently dense, the sample-size rate takes over, and adding more observations per subject no longer changes the asymptotic order. The boundary depends on predictor sampling frequency and the model's smoothness parameters.

An aligned toy case adds a less intuitive wrinkle. When the relevant eigenfunctions are assumed to line up, increasing predictor smoothness pushes the transition to a higher sampling frequency. The paper calls this the curse of smoothness, but also ties the illustration to the alignment assumption, so it may not describe more general settings.

When both curves are sampled

Function-on-function regression tracks two sampling frequencies, one for the predictor curve and one for the response curve. Its reported prediction-rate structure has separate contributions from sample size, predictor discretization and response discretization, meaning how coarsely each curve is observed. Matching bounds support a minimax-optimal structure. Logarithmic factors can be removed only under the stated inequalities, and boundary cases may retain them. As the two frequencies vary, the theory allows up to three phase transitions.

What the tests showed

The simulations broadly followed the theoretical picture. In scalar-on-function experiments, the proposed estimator was reported to have the lowest errors across the settings studied, with its clearest advantage when sampling was sparse. Differences from the FullyRKHS benchmark narrowed as sampling frequency rose. Increasing the frequency improved error, but the gains diminished at higher frequencies.

In cosine-design function-on-function simulations, the proposed method outperformed five benchmark alternatives. Results using a Legendre design in the supplement were described as qualitatively similar. These comparisons depend on the simulated processes, kernels, truncation and tuning choices.

Two applications

In a wheat application, the data came from 100 individuals whose near-infrared spectra were recorded at 701 wavelengths, spaced 2 nanometres apart from 1100 to 2500 nanometres. Each evaluation used a random split of 80 training subjects and 20 testing subjects. The training data were resparsified to 5, 10 or 20 measurements per subject, and the exercise was repeated 100 times.

At all three sparsity levels, the proposed method had the lowest reported empirical mean prediction error: 0.8029 at five measurements per subject, 0.6761 at 10 and 0.5920 at 20. The comparison reports repeated empirical means and standard deviations, but no significance tests or confidence intervals.

The other application used anthropometric data from 197 children in the CONTENT cohort. Predictor BMI-Z trajectories had 5 to 17 observations, with a median of 14, while response trajectories had 2 to 11 observations, with a median of 7.

Across 100 repeated CONTENT evaluations, the proposed method had the lowest reported value on every listed summary measure: mean empirical error of 0.3536, median 0.2754, standard deviation 0.1972, first quartile 0.1868 and third quartile 0.4899. The evaluation used repeated internal train/test partitions rather than external validation.

A result bounded by its assumptions

The empirical evidence is limited to the specified simulations and two application datasets, with the real-data evaluations using the split and resparsification procedures described above. The theoretical conclusions depend on the paper's stated assumptions. The reported comparisons therefore remain tied to the settings studied.

The front matter says the manuscript was submitted to the Annals of Statistics. Its supplement is described as containing proofs of the main theoretical results, technical lemmas and additional simulation results.

Paper data and sources

Original title: Functional linear regression from sparse to dense designs: a pooling-ridge method and minimax optimality
Authors: Shunxing Yan, Fang Yao
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.