The most striking result from a Shenzhen neighborhood simulation was the gap in modeled travel time between resident profiles. Aunt Chen completed all 20 of her activity-mobility records, yet her mean travel time was 25.1 minutes and only 10.0% of her trips were within 15 minutes. Across her activities, mean times ranged from 9.8 to 31.3 minutes. At the same time, no scheduled activity failed or required substitution in the presented scenario. The figures come from this modeled scenario.
From nearby facilities to lived routines
The prototype was designed to address a gap between static spatial accessibility measures and resident-specific livability evaluation. It simulates household activity-mobility schedules and uses event histories for evidence-grounded post-simulation interviews. The aim was to make resident-specific differences visible through modeled schedules and their resulting events.
It connects four modules: spatial setup and knowledge-graph construction, household and persona configuration, iterative scheduling and network materialization, and inspection and interviews. The knowledge graph is a structured map of local places used to retrieve facilities, while event histories support the later livability assessment.
How the routes were built
For each residence, the knowledge graph kept up to 10 facilities per category within 3 km. It assigned approximate walking time using 70 metres a minute, while the complete road network supplied the final routes and travel times. The model therefore narrowed nearby options before calculating journeys on the network.
Schedules went through a repair-and-review loop. If issues remained unresolved, they triggered targeted revision by the large language model and re-entry into the repair and diagnostic process. A schedule was accepted only when no unresolved issues remained. If the maximum review rounds were exceeded, the system recorded a scheduling failure rather than inserting a rule-generated replacement.
Accepted schedules were then materialized by snapping trip locations to nearby road-network nodes, calculating distance-weighted shortest paths, assigning modes through rules, and deriving network distance and travel time from mode-specific reference speeds.
A bounded Shenzhen test
The demonstration focused on Hong Leong Technology Park in Shenzhen. It included five households built from predefined profiles and one user-defined household, making six households and 11 resident agents in all. Residences were randomly sampled within the neighborhood boundary. The evidence is consequently bounded to that neighborhood and those households.
Travel and care did not look the same
Across the 11 agents, the simulation generated 172 unique event records. No activity failed or required substitution in the presented scenario. The authors characterize this as an environment with routable facility options, but say the run is more informative about travel burden and household coordination than about severe service deprivation.
Travel burden differed across the modeled profiles. Qiang Wu had the highest mean travel time, at 30.7 minutes; Xiao Hong had the lowest, at 17.9 minutes. Aunt Chen had 10.0% of trips within 15 minutes, compared with 36.4% for both Mr. Gao and Xiao Hong. The study presents these as descriptive between-profile differences: the figures also reflect persona, schedules, household responsibilities, and residential micro-location, and no inferential comparison was reported.
Care requirements were visible in the schedule as well. Every simulated event involving the infant and the school-age child was accompanied by another household member. Qiang Wu participated in 27 accompanied events and Fang Wu in 16. In the model, this illustrated dependence on another household member’s availability.
The ratings were synthetic
The model-generated livability averages were 6.0 for daily convenience, 5.0 for healthcare access, 6.4 for travel burden, 6.4 for activity feasibility, 5.2 for trip chaining, 6.2 for household coordination, and 6.1 for recreation or social life. In this scenario, healthcare access and trip chaining received lower averages than travel burden and activity feasibility.
Open-ended suggestions also varied by profile. Older or mobility-constrained agents emphasized closer healthcare, parks, and accessible walking. Households with children emphasized childcare-supportive facilities, nearby recreation, and convenient daily services. Younger adults more often requested parcel lockers, community services, and leisure or social destinations.
Those ratings and suggestions were structured interpretations of simulated event histories, not survey observations or validated measures of real residents’ perceptions.
A result that needs more testing
That distinction is important because the model was not calibrated against observed trajectories. Schedule generation and interview outputs were sensitive to the selected language model, prompts, and review settings. The authors identify repeated runs and sensitivity analysis as necessary steps for testing the stability of the outputs.
The zero-failure result should therefore be read only within the scenario that produced it. The result may change with different spatial data, residence locations, or scheduling assumptions, and the evidence remains limited to the demonstrated neighborhood and six-household scenario. The prototype can expose scenario-specific differences in modeled access and household coordination, but the run does not establish universal livability.
Paper data and sources
Original title: Spatial-Knowledge-Graph-Grounded LLM Agents for Neighborhood Livability Evaluation
Authors: Haiyan Hao
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text