诸如AlphaEarth Foundation等地理空间基础模型可生成紧凑且全球一致的地球表面表征,能有效迁移到多种下游任务。然而,由于这些模型主要基于地球观测影像进行训练,其嵌入主要捕获物理与光谱特征,而对人类活动与城市功能的编码则较弱。为弥补这一局限,我们提出BEACON——一种三模态对比学习框架,旨在对齐城市空间的三种互补视角:来自AlphaEarth(AE)嵌入的物理表征、来自兴趣点(POI)文本的语义表征,以及来自小时级POI访问量的人类行为表征;同时保持部署后的表征仍为纯图像形式。以休斯顿大都市区为案例研究区域,我们在九项下游任务(包括七项回归任务与两项分类任务)上评估了BEACON框架的性能,并以六种基线方法(原始坐标、Space2Vec、SatCLIP、TESSERA、Clay与AlphaEarth)为对照,在五组随机种子下采用冻结的线性探针与MLP探针进行评估。在线性探针设置下,BEACON在肥胖患病率、不良心理健康状况与家庭中位收入三项指标上的相对R²较AlphaEarth分别提升最高达43%、34%与22%,同时在物理与环境变量预测方面仍保持竞争力。上述结果凸显了将语义与行为信号融入地理空间基础模型的价值,从而将其适用范围从物理地球观测拓展至以人为中心的城市分析。
Geospatial foundation models such as the AlphaEarth Foundation produce compact and globally consistent representations of the Earth's surface that transfer effectively to a wide range of downstream tasks. However, because these models are trained primarily on Earth-observation imagery, their embeddings mainly capture physical and spectral characteristics while encoding human activity and urban function only weakly. To address this limitation, we propose BEACON, a tri-modal contrastive learning framework that aligns three complementary views of urban space: physical representations from AE embeddings, semantic representations from point-of-interest (POI) text, and human behavioral representations from hourly POI visitation, while keeping the deployed representation image-only. Using the Houston Metropolitan Area as a case study area, we evaluated the performance of the BEACON framework on nine downstream tasks, including seven regression and two classification tasks against six baselines (raw coordinates, Space2Vec, SatCLIP, TESSERA, Clay and AlphaEarth), using frozen linear and MLP probes over five seeds. Under a linear probe, BEACON improves relative R^2 over AlphaEarth by up to 43% for obesity prevalence, 34% for poor mental health, and 22% for median household income, while remaining competitive in the prediction of physical and environmental variables. These findings highlight the value of augmenting geospatial foundation models with semantic and behavioral signals, extending their applicability from physical Earth observation to human-centered urban analytics.