在密集城市环境中,机器人与可穿戴平台上的精确视觉定位仍具挑战性。现有方法通常依赖GPS实现绝对定位,但在城市峡谷中,由于多径传播效应,GPS信号常发生退化。因此,标准方案(如视觉里程计)会随时间累积不可控的漂移;而地图匹配技术则难以获取其所需的可靠GPS先验信息,且计算开销过大,难以满足实时边缘端执行需求。为应对上述局限,我们提出Spotter——一种鲁棒、实时的视觉定位框架,以建筑物立面作为可靠的全局地理参考源,同时保留对可用GPS信号的融合能力。在离线阶段,Spotter处理Google街景全景图像,通过语义分割提取立面区域,并结合多视角立体深度与测绘数据构建紧凑的度量数据库;在运行时,查询图像经由级联式检索与几何验证流程完成匹配,从而恢复细粒度的全局相机位姿。我们在巴塞罗那多个城区采集的可穿戴智能眼镜行人序列新数据集上对Spotter进行了基准测试。实验结果表明,Spotter优于基于里程计的基线方法,在定位精度上可媲美当前最先进的基于地图的方法,同时帧率显著更高。
Accurate visual localization on robotic and wearable platforms remains challenging in dense urban environments. Existing methodologies typically rely on GPS for absolute positioning, yet GPS signals frequently degrade in urban canyons due to multipath propagation. Consequently, standard solutions like visual odometry suffer from unmitigated drift over time, while map-matching techniques struggle to acquire the reliable GPS priors they need, on top of being too computationally heavy for real-time edge execution. To address these limitations, we propose Spotter, a robuts and real-time visual localization framework that uses building facades as a reliable source of global geo-reference, while retaining the capability to integrate GPS signals when available. In an offline stage, Spotter processes Google Street View panoramas by semantically segmenting facades and pairing multi-view stereo depth with cartographic data to build a compact metric database. At runtime, query images are matched via a cascaded retrieval and geometric verification pipeline to recover fine-grained global camera localization. We benchmark Spotter on a newly collected dataset of pedestrian sequences acquired with wearable smart glasses across several districts of Barcelona. Experimental results show that Spotter outperforms odometry-based baselines and achieves localization accuracy comparable to state-of-the-art map-based methods while operating at significantly higher frame rates.