AI agents are not immune to a hotel's place in a search list, but rank matters less to them than it does to people in the benchmark used in the experiment. Moving a hotel higher made it more likely to be inspected. Yet rank did not affect final booking in every model, and at the highest reasoning setting tested, the inspection curve no longer had a statistically significant U-shape and final-choice position effects were no longer statistically significant for both tested models.
Researchers randomized the presentation order of 100 hotels in each session. That design separated a hotel's placement from the other listing information in the test and made it possible to examine two stages separately: whether the agent inspected a hotel and whether it submitted that hotel as the final booking. A position effect simply means a change in inspection or choice associated with where the listing appeared.
The agents searched beyond the first results
The main baseline comprised 2,000 independent AI-agent sessions across four major LLMs from Google and Anthropic, with 500 sessions for each model. The agents worked in a simulated Manhattan hotel-booking environment. They could inspect listings sequentially, submit a hotel booking, or terminate without booking, representing an outside option.
The AI agents inspected between 1.63 and 5.83 hotels per session, compared with 1.12 for the human benchmark. They booked in 100% of the reported baseline sessions. The human benchmark converted in 66% of sessions and took the outside option in 34%.
Rank shaped inspection in a smaller way
Rank mattered most at the inspection stage. In the randomized AI sessions, moving a listing down 10 ranks was reported to lower its inspection probability by 0.2 to 0.5 percentage points. The comparable human estimate was a 1.9-point fall, making the AI effect four to ten times smaller. All reported position coefficients had p<0.001.
Rank did not produce a smooth top-to-bottom decline. When researchers allowed the relationship to bend rather than forcing a straight line, inspection was lowest between ranks 68 and 74 before rising again for three of the four models. The straight-line term was negative and significant for all four LLMs, while the upward-curving term was significant for three. That left a dip in inspection in the middle-to-lower part of the list rather than a simple preference for the first result.
Booking choices split by model
The inspection pattern did not carry through uniformly to final bookings. For Flash and Pro, position was statistically indistinguishable from zero, with p=0.338 and p=0.606. Sonnet and Flash Lite had negative, statistically significant effects, with p=0.018 and p<0.001. The estimated effects ranged from negative 0.000008 for Pro to negative 0.000080 for Flash Lite.
Even with these differences, choices converged. Every LLM concentrated between 89.6% and 100% of its choices in the same five hotels. The shared modal choice was citizenM New York Times Square. The analysis described it as undominated on price and review score, meaning no other listing plainly beat it on both measures. It captured 78.2% of pooled choices, or 1,563 of 2,000 sessions.
More reasoning, weaker position effects
Researchers varied reasoning effort in a follow-up with seven cells, 500 replications in each and 3,500 in total. At the highest effort level, the inspection curve lost its U-shaped curvature, and final-choice position effects were no longer statistically significant for both tested LLMs. The result links higher tested effort with weaker position patterns, while the earlier model differences show that the response was not uniform.
The test has clear limits
These results are limited to the tested sessions and models. The order was randomized for the 100 hotels in each AI-agent session, but the human numbers are a benchmark comparison, so the study does not show that the same magnitude would hold in every search setting. Nor does it show that rank affects final booking uniformly: the final-choice estimate was significant for Sonnet and Flash Lite but not for Flash and Pro. The U-shaped inspection pattern appeared in three of the four models, not all four.
The practical message is narrower than a simple rank-wins story. Placement can influence inspection, but the effect is smaller than in the human benchmark, and final booking varies by model and tested reasoning effort. In this experiment, the models still converged on the same small group of hotels. Whether listing attributes matter more than rank in broader real-world settings remains an open question.
Paper data and sources
Original title: Does Rank Still Matter? Position Bias When AI Agents Shop on Our Behalf
Authors: Davood Wadi, Yu Ma
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-24
DOI: Not available
Original paper · Full text