多模态大语言模型(MLLM)能够解析街景图像,但城市主体性取决于此类局部感知证据在智能体开始移动后是否仍具实用性。本文探究当前MLLM智能体在复杂的真实尺度城市中,将局部城市感知转化为可靠行为的能力边界。我们提出UrbanGround——首个基于全境三维地理空间数据构建的香港物理约束式仿真沙盒,使该问题得以实证检验。UrbanGround支持第一人称视角的闭环交互,并提供交互式地图用于导航。智能体可直接进入三维城市环境,以第一人称视角进行探索。我们的分析围绕空间问题的演进,通过三个研究问题展开:首先检验智能体能否通过主动观察充分实现局部场景的视觉定位,从而回答空间问题;其次考察该定位能力能否支撑导航任务,尤其当目的地距离更远、指示更模糊时;最后评估所生成的行为在路径可用性变化及行人运动干扰下是否具备鲁棒性。当前主流MLLM智能体通常在视觉识别与短程空间推理方面展现出有效的原子能力,但在方向定位与行人感知式运动方面仍不可靠。其核心缺陷显现于延展性探索过程中:局部能力无法有效组合为持续的目标导向行为,且错误随探索进程累积而缺乏有效修正机制。我们期望UrbanGround能推动更广泛的研究,以评估当前MLLM智能体在复杂、开放的城市环境中可靠探索能力的边界。
Multimodal large language models (MLLMs) can interpret a street view, but urban agency depends on whether such local evidence remains useful after the agent starts to move. In this paper, we investigate how far current MLLM agents can turn local urban perception into reliable action in a complicated real-scale city. We propose UrbanGround, the first sandbox to make this question testable in a physically constrained replica of Hong Kong built from territory-wide 3D geospatial data. UrbanGround supports closed-loop interaction from a first-person view and provides an interactive map for navigation. Agents can directly enter the 3D city and explore from a first-person view. Our analysis follows the growth of the spatial problem through three research questions. We first test whether an agent can ground a local scene well enough to answer spatial questions after active observation. Then we ask whether that grounding supports navigation as destinations become farther away and less explicit. Finally, we examine whether the resulting behavior survives changes in route availability and pedestrian motion. Contemporary MLLM agents usually show useful atomic abilities in visual recognition and short-range spatial reasoning, while orientation and pedestrian-aware movement remain unreliable. Their central failure emerges over extended exploration, where local abilities do not compose into sustained goal-directed behavior and errors accumulate without effective correction. We hope UrbanGround will support broader study of how far current MLLM agents can explore reliably in complex, open-ended urban environments.