A year-long field campaign in a subarctic boreal forest found substantial degradation in state-of-the-art autonomous navigation methods under the studied conditions. Simple proprioceptive odometry, a baseline that tracks the robot’s own motion, was comparatively robust.
The study evaluated 64 km of sensor data with nine odometry, localization and mapping methods. It included six repeated trajectories, mainly on wide forest roads with good sky visibility, and covered seasonal temperature shifts of 60 °C.
The simplest estimate held up best
Proprioceptive odometry had a combined mean Sequentially Aligned Relative Trajectory Error, or SARTE, of 3.6%. SARTE was the study’s measure of trajectory error. The highest reported trajectory values were 5.6% on Green and 5.4% on Magenta.
Exteroceptive odometry, which relies on outside observations such as visual, lidar or radar data, generally performed worse than the proprioceptive baseline, with average degradation of 10 percentage points. Lidar methods nevertheless improved over baseline drift across most trajectories.
Radar Teach and Repeat was the most fragile exteroceptive odometry method in the evaluation. It degraded on every trajectory, reached 22 percentage points of degradation on Orange and had a 17% failure rate.
The visual-odometry system DROID-SLAM also degraded on all six trajectories. Its mean SARTE was above 13%, and its 5.1% failure rate included a failure associated with a night deployment.
More processing did not guarantee a better result
Loop closure, the recognition that a robot has returned to a known place, and pose-graph optimization, or PGO, improved state estimation for only two of the six methods tested in that comparison. For 2Fast-2Lamaa, SARTE fell by 0.2 percentage points to 1.4%.
Navtech-Radar-SLAM identified 9,860 loop closures across the other 48 runs, most of which were false positives. Its performance dropped by 44.2 percentage points compared with ORORA, illustrating the risk of treating a changing forest scene as a familiar place.
Seasonal changes reshaped the map problem
The seasonal results differed by sensor type. Proprioceptive median error was below 0.9% in summer and rose to 1.8% in both autumn and winter. WILN, by contrast, kept its median error below 0.6% across the three analyzed seasons.
Visual performance became especially poor when winter imagery offered few features to track. For cuVSLAM, Relative Traveled Distance Error, or RTDE, exceeded 7% when the system processed only winter images with an average of fewer than 50 detected features.
Using a map made in one season to localize the robot in another produced a mixed picture. 2Fast-2Lamaa had the best cross-season prior-map performance among the tested methods. WILN completed 94% of autumn-winter localization trajectories, but its inter-winter Localization Failure Ratio, or LFR, was 32%, and winter query data produced 44.5% error when matched against a summer map.
ORB-SLAM3 completed only one control trajectory entirely in localization mode and was highly sensitive to scenes that changed between seasons. In a lidar Teach-and-Repeat trial, a run using a 113-day-old teach trajectory failed in an area where snowbanks reached 3 m. The authors reported that filtering lidar returns under 1.5 m above the sensor left enough features for localization.
A demanding test with clear limits
The systems were evaluated offline, with sensor frames processed sequentially to reduce data loss from network lag, buffer overflows and limited computing resources. For its reference path, the study used ground truth reconstructed from three GNSS antennas and a static reference antenna with post-processed kinematic corrections.
A run counted as a failure only when its estimated trajectory ended before 95% of the ground-truth duration. Severe state-estimation degradation without early termination was not counted as failure, so the reported failure rates do not capture every serious loss of navigation quality.
The evidence remains narrow: it comes from one robot and one subarctic deployment site, with some seasonal analyses based on selected trajectories and no spring analysis because data were limited. The study reports no formal inferential statistical model or uncertainty estimates, and most method parameters were retained from the original systems.
The results therefore describe how the tested systems behaved in this field campaign, rather than establishing that any method is universally robust across subarctic environments or that seasonal change alone caused each error.
Paper data and sources
Original title: One year in a forest: Analyzing the challenges of autonomous navigation in subarctic environments
Authors: Matěj Boxan, Nicolas Lauzon, Veronica Vannini et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-27
DOI: Not available
Original paper · Full text