地球观测(Earth Observation, EO)已从根本上改变了对环境过程和人类活动的全球尺度监测。近期自监督学习的发展催生了地球观测基础模型(Earth Observation Foundation Models, EOFMs),该模型利用PB级未标注EO数据,学习可迁移表征,以支持广泛下游地理空间任务。尽管取得进展,当前EOFMs仍主要局限于栅格模态,忽视了OpenStreetMap与Overture等公开可获取矢量数据源中蕴含的丰富结构化信息。矢量数据以显式且紧凑的方式表征地理实体,涵盖几何、拓扑及语义关系,提供了影像本身常难以明确表达或无法获取的关键上下文信号。因此,栅格与矢量数据代表了地理空间的互补视角:栅格数据捕获连续的物理与光谱模式,而矢量数据则编码离散对象及其关系结构,且往往更侧重于人类系统(如社会或人口统计学数据),而非纯物理系统。然而,现有地理空间表征学习范式将这两种模态割裂处理,依赖不完善且常具信息损失的转换方法来弥合二者。本文作为观点论文,呼吁转向一种新范式——联合空间表征学习(Spatial Representation Learning, SRL),即在统一嵌入空间中整合栅格感知与基于矢量的推理。我们基于多模态地理空间学习的新兴进展,阐述其概念基础、技术挑战以及融合异构空间数据源的潜在路径,并主张此类融合对于构建下一代地理空间基础模型至关重要。
Earth Observation (EO) has fundamentally transformed the monitoring of environmental processes and human activities up to planetary scale. Recent advances in self-supervised learning have given rise to Earth Observation Foundation Models (EOFMs), which leverage petabyte-scale unlabeled EO data to learn transferable representations across a wide range of downstream geospatial tasks. Despite these advances, current EOFMs remain largely confined to raster modalities, overlooking the rich, structured information encoded in openly-accessible vector data sources such as OpenStreetMap and Overture. Vector data provides explicit and compact representations of geographic entities, including geometry, topology, and semantic relationships, offering critical contextual signals that are often ambiguous or inaccessible in imagery alone. Raster and vector data thus represent complementary views of geographic space: raster data captures continuous physical and spectral patterns, while vector data encodes discrete objects and their relational structure and often represents more of the human rather than the physical systems (e.g. social or demographic data). However, existing geospatial representation learning paradigms treat these modalities in isolation, relying on imperfect and often lossy transformations to bridge them. This perspective paper calls for a paradigm shift toward joint Spatial Representation Learning (SRL) in an unified embedding space that integrate raster perception with vector-based reasoning. Building on emerging efforts in multimodal geospatial learning, we highlight conceptual foundations, technical challenges, and promising directions for aligning heterogeneous spatial data sources. We contend that such integration is essential for developing next-generation geospatial AI systems capable of more accurate, interpretable, and semantically grounded understanding of the Earth.