A new arXiv preprint reports that the study’s best vision-based setup recorded 71.5% accuracy when predicting stated choices among five travel modes. That was only slightly above the best text-only zero-shot result, 69.9%, and a random-forest benchmark, 69.6%. The authors stress that these are descriptive differences, not evidence that the vision setup was statistically superior.
The study evaluated a multi-agent workflow linking AI-assisted behavioral data collection with travel-choice modeling and large-language-model prediction. Three researcher-supervised agents—Data Collection, Data Processing and Data Modeling—exchanged standardized data objects through explicit interfaces.
A survey built around five weather scenes
Researchers used a chatbot to administer an image-augmented stated-preference survey. It presented five fixed weather scenarios and five mode alternatives, with every respondent seeing the same fixed image for each scenario.
The sample included 92 students from one university. Across the five scenarios, there were 460 potential respondent–scenario observations; six unparseable or missing choices were excluded, leaving 454 valid observations.
The models were evaluated on the same 454 observations. One task required a five-class choice, while the other grouped choices into active versus non-active travel. Random guessing would score 20.0% on the five-class task and 50.0% on the binary task.
Reported choices varied with the weather
The stated choices showed a marked change across the weather scenarios. Public transit’s share rose from 17.6% under Sunny to 45.1% under Snowy. In the seasonal travel profile, cycling accounted for 13.0% of summer primary commutes and 2.2% of winter primary commutes, while public transit accounted for 20.7% and 43.5%, respectively.
A multinomial logit model, used to estimate associations among several travel choices, also linked adverse scenarios with less cycling and more public-transit and driving choices. With Walking and Sunny as reference categories, cycling coefficients were negative under Rainy, Foggy/cold and Snowy conditions. The expanded model fit better than the baseline, with a likelihood-ratio statistic of 53.92, 8 degrees of freedom and p < 0.0001. These results describe associations in stated choices, not evidence that weather caused changes in actual travel.
How the prediction systems compared
Among the conventional benchmarks, random forest had the highest reported accuracy: 69.6% for five classes and 88.8% for the binary task. Logistic regression reached 60.2% and 85.1%, while the multinomial logit model reached 44.7% and 81.2%, respectively.
The language-model comparison covered nine locally run models, ranging from 2 billion to 35 billion parameters. The researchers varied Expert and Role-Play framing and Base versus Richer Context, then added persona information, few-shot examples and vision in further configurations.
The strongest text-only zero-shot result came from Gemma 4:12B under Expert framing with Richer Context: 69.9% five-class accuracy and 73.5% binary accuracy. The study reports those figures with ±3.9 and ±4.2, but the analysis provided for review does not specify what those quantities measure.
Prompt details changed the comparison
In a reported Gemma 3:4B comparison, five-class accuracy was 41.3% with Expert framing and Base Context, compared with 64.2% with Expert framing and Richer Context. Expert framing generally outperformed Role-Play across the tested configurations.
Persona information was most useful when direct travel history was absent; its added contribution was limited when Richer Context already included habitual seasonal modes. Few-shot prompting was associated with higher five-class accuracy for most models, especially smaller ones, but performance generally stabilized after approximately ten examples, with later gains limited or inconsistent.
Why the result remains preliminary
The study’s evidential base was narrow: 92 students from one university, stated choices rather than observed trips, and one fixed generated image for each weather condition. Because weather and image identity were not separately varied, uncontrolled visual features may have influenced both respondents and models.
The authors also caution that exploring multiple model and prompt configurations on the same sample may have contributed to the highest reported accuracy. The workflow was not experimentally compared with a conventional research workflow, so the study does not demonstrate lower cost or processing effort.
The document is an arXiv preprint, version 1, dated 20 Aug 2026. Its results are therefore best understood as an early methods comparison of stated travel choices within this sample.
Paper data and sources
Original title: An Agentic Approach for Active Data Collection, Travel Behavior Modeling, and Weather-Sensitive Demand Prediction
Authors: Narges Ahmadi, Yubo Jiao, Jônatas Augusto Manzolli et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text