论文
arXiv
Trajectory
Mobility
LLM
Agent
UrbanTraffic
GeoSimulation
GeospatialFlow
中文标题
当合理不等于真实:评估基于大语言模型的urban simulation中人类移动行为的真实性
English Title
When Plausible Is Not Realistic: Evaluating Human Mobility in LLM-Based Urban Simulation
Gustavo H. Santos, Aline Carneiro Viana, Thiago H. Silva
发布时间
2026/6/12 03:03:52
来源类型
preprint
语言
en
摘要
中文对照

基于大语言模型(LLM)的生成式智能体正日益被用于城市模拟器,但其是否能再现经验上真实的人类移动模式,抑或仅生成表面合理的移动叙事,仍不明确。我们提出一种验证框架,用于将LLM驱动的城市模拟器中生成式智能体的移动行为与真实世界移动数据进行比对评估。该框架涵盖移动定律、时间节律、网络模体、语义活动转移以及行为移动画像等维度。我们利用大巴黎地区与上海的数据集,在多个移动真实性维度上评估AgentSociety与CitySim。分析表明,叙事合理性与经验移动真实性之间存在显著差距。尽管这些模拟器能在一定程度上捕捉高层语义活动分布,却难以再现核心时空约束,包括真实的行程长度分布、起讫点(OD)流、停留时长及转移动态。我们进一步发现,真实的移动多样性在默认提示配置下不稳定,可能需依赖显式的画像感知初始化。为支持可复现的评估,我们还贡献了一套可扩展、开源的LLM驱动基础设施,支持区域尺度地图生成、可观测性增强的模拟、移动度量计算及交通模拟。本研究强调了对LLM驱动城市模拟器开展严格经验验证的必要性,并提供了构建更真实、更可复现城市模拟系统的实用工具。

English Original

LLM-based generative agents are increasingly used in urban simulators, yet it remains unclear whether they reproduce empirically realistic human mobility patterns or merely generate plausible mobility narratives. We introduce a validation framework for evaluating the mobility of generative agents of LLM-based urban simulators against real-world mobility data. For this, we use mobility laws, temporal rhythms, network motifs, semantic activity transitions, and behavioral mobility profiles. Using datasets from the Greater Paris region and Shanghai, we evaluate AgentSociety and CitySim across multiple dimensions of mobility realism. Our analysis reveals a substantial gap between narrative plausibility and empirical mobility realism. Although the simulators capture some high-level semantic activity distributions, they struggle to reproduce core spatial and temporal constraints, including realistic trip-length distributions, origin-destination flows, dwell times, and transition dynamics. We further observe that realistic mobility diversity is unstable across default prompting configurations and may require explicit profile-aware initialization. To support reproducible evaluation, we also contribute scalable and open LLM-driven infrastructure for regional-scale map generation, observability-enhanced simulation, mobility-metric computation, and traffic simulation. Our findings highlight the need for rigorous empirical validation of LLM-based urban simulators and provide practical tools for building more realistic and reproducible urban simulation systems.

元数据
arXiv2606.13835v1
来源arXiv
类型论文
抽取状态raw
关键词
Trajectory
Mobility
LLM
Agent
UrbanTraffic
GeoSimulation
GeospatialFlow
cs.CL
cs.AI
cs.MA