论文
arXiv
SpatialIntelligence
Multimodal
GeoMultimodal
GeoSimulation
中文标题
WildCity:面向渲染、仿真与空间智能的现实世界城市级测试平台
English Title
WildCity: A Real-World City-Scale Testbed for Rendering, Simulation, and Spatial Intelligence
Xiangyu Han, Mengyu Yang, Jiaqi Li, Bowen Chang, Ziyu Chen, Hexu Zhao, Rahul Kumar Agrawal, Anthony Rodriguez, Fiona Hua, Marco Pavone, Chen Feng, Yiming Li
发布时间
2026/7/8 06:19:51
来源类型
preprint
语言
en
摘要
中文对照

人类能够导航陌生城市,并逐步构建覆盖数十平方公里的连贯空间心理地图。AI 能否在同等尺度上构建空间表征?尽管近期的基础模型已在场景重建与具身智能方面取得进展,但扩展至整座城市仍是一个开放性挑战,主要原因在于缺乏城市尺度的数据。为弥合这一差距,我们提出 WildCity——一个由自动驾驶车队在复杂城市环境中采集的真实世界多模态数据集。该数据集包含 18 条轨迹,平均每条长度达 83.7 公里,并保留了野外感知的核心挑战,例如动态物体、光照变化及不完美的相机位姿。我们进一步构建了面向城市的重建基线,并将重建环境转换为闭环仿真器。除数据集与基线外,我们系统性地分析了构建仿真就绪型城市数字孪生体的关键挑战:可扩展性、外推能力与不确定性。最终,WildCity 旨在推动城市尺度渲染的发展,并更广泛地促进 AI 在空间感知、记忆与推理能力上的进步,使其达到与人类认知相当的尺度。项目主页:https://han-xiangyu.github.io/Wild-City/

English Original

Humans can navigate an unfamiliar city and gradually form a coherent spatial mental map spanning tens of square kilometers. Can AI build spatial representations at a comparable scale? Although recent foundation models have advanced scene reconstruction and embodied intelligence, scaling to entire cities remains an open challenge, primarily due to the lack of city-scale data. To bridge the gap, we introduce WildCity, a real-world multimodal dataset collected by autonomous fleets traversing complex urban environments. Our dataset includes 18 trajectories, each averaging 83.7 kilometers in length, and preserves the core challenges of in-the-wild perception, e.g., dynamic objects, lighting variations, and imperfect camera poses. We further establish an urban-tailored reconstruction baseline and convert the reconstructed environments into a closed-loop simulator. Beyond the dataset and baseline, we systematically analyze the key challenges on the path to simulation-ready urban digital twins: scalability, extrapolation, and uncertainty. Ultimately, WildCity aims to catalyze progress not only in city-scale rendering, but more broadly in the pursuit of AI that can perceive, remember, and reason across space at a scale comparable to human cognition. Project page: https://han-xiangyu.github.io/Wild-City/

元数据
arXiv2607.06838v1
来源arXiv
类型论文
抽取状态raw
关键词
SpatialIntelligence
Multimodal
GeoMultimodal
GeoSimulation
cs.CV