An artificial-intelligence model has flagged 14,426 likely young stars in a catalog drawn from about 1.3 million APOGEE sources. But its reported recovery rate varied dramatically by stellar temperature, from almost 100% for M dwarfs to about 0% for O-type stars.
Of the new candidates, 12,621 were ABYSS-targeted and 1,805 were not. The study used a convolutional neural network to look for stars younger than 40 million years in APOGEE spectra, then applied a separate classifier to optical BOSS spectra.
Both data sets came from SDSS-V. APOGEE supplied near-infrared H-band spectra with resolving power around 22,500, while BOSS supplied optical spectra at around 1,800. ABYSS aimed to observe around 200,000 candidate young stars.
From spectra to training labels
The initial labeling scheme marked ABYSS-targeted sources as likely young and the remaining APOGEE sources as field stars. The labels were then refined with effective-temperature and surface-gravity cuts, isochrone ages, and spatial or kinematic clustering.
The final APOGEE set contained 19,498 labeled likely young stars and 889,094 labeled field stars. Those labels were the yardstick for training and evaluation, rather than an independently verified census of every star's age.
The model used median-normalized spectra, four residual-network blocks, adaptive average pooling and six fully connected layers. A fixed random subset of 12,800 spectra was withheld as a development set. Newer sources from Astra version 0.8 were reserved for final evaluation. Training used a learning rate of 10−5 and early stopping after two epochs; no systematic hyperparameter optimization or cross-validation was performed.
Against the adopted labels, the APOGEE classifier recovered 13,663 of 19,498 labeled likely young stars. That is 70% recall, the share of labeled young stars the model recovered. Its reported precision, the share of flagged sources matching the adopted young-star label, was about 90%. About 45% of false positives fell in parameter regions originally excluded during training.
Performance changes with temperature
Performance changed sharply when the results were split by effective temperature, the standard measure of a star's surface temperature. Recall fell from almost 100% for M dwarfs to about 0% for O-type stars, with a local minimum near 5,000 kelvin. Precision was 50% among the coolest stars, reached nearly 100% at 3,000 K, and declined for hotter stars.
Most of the sample was estimated to be younger than 10 million years, while stars reaching about 40 million years appeared only among lower-mass stars. The ages came from Sagitta, so they are estimates within the study's analysis, not independent age measurements for every source.
The reported near-zero recall for O-type stars means the catalog cannot be treated as a complete inventory of young stars across all temperatures. It is better understood as a temperature-dependent selection, with the strongest recovery reported for cooler stars.
The clues in the infrared
The study also examined what the network might be seeing in the near-infrared data. It built empirical field-star templates by grouping spectra by temperature, surface gravity and metallicity, smoothing and stacking them, then comparing them with young-star spectra after rotational broadening. For stars cooler than 6,300 K, the typical best-fit projected rotation was about 20 kilometers per second, and the comparison identified 57 large residual features.
The authors considered enhanced magnetic activity in young stars a more plausible explanation for the residual pattern than a single chemical species or a simple abundance difference. But the line identities were incomplete, and the broadening correction was described as a first-order treatment. The residual pattern is therefore a tentative interpretation of the spectral signal, not a settled physical explanation.
BOSS sees less
In the separate BOSS analysis, the classifier identified 23,154 likely young stars among 1.26 million sources. It recorded 2,096 false positives, about 91% precision, 24,288 false negatives and 46% recall. Maximum recall among the coolest stars was about 80%.
That precision was comparable to APOGEE's reported figure, but BOSS was less complete within the same labeling framework. The study discusses BOSS's lower resolving power and weaker optical spot contribution as factors associated with its lower recall.
What the result does not settle
The largest caveat is the labels themselves. Likely-young and field-star categories were built from targeting and selection rules, so the reported precision and recall measure agreement with that scheme, not true sensitivity, specificity or real-world contamination against independent ground truth.
Future validation will need to test the catalog against an independent, well-characterized young-star sample, especially for hotter stars and dissolving moving groups. The study also leaves open whether more physically informed models, different calibration methods or alternative class-balance strategies could improve completeness while controlling contamination, and what the 57 residual features physically represent.
Paper data and sources
Original title: ABYSS. IV. Identifying signatures of stellar youth in APOGEE spectra
Authors: Valentina Bonilla Villalobos, Marina Kounkel, Joseph Mullen et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-25
DOI: 10.3847/1538-3881/ae9d5f
Original paper · Full text