地理空间基础模型(Geospatial Foundation Models, GFMs)的基准测试日益依赖聚合得分进行模型排序,但此类排名掩盖了模型差异的成因:性能差距在多大程度上源于架构设计、解码器容量,又在多大程度上源于特定用例的偶然因素?本研究通过受控实验,对欧洲航天局(ESA)Φ-lab 下开发的两款设计理念迥异的 GFM——THOR 与 TerraMind——展开系统性比较。THOR 采用计算自适应架构,支持可变图像块(patch)尺寸,并在原始分辨率下统一融合 Sentinel-1、Sentinel-2 与 Sentinel-3 数据;TerraMind 则是一种多模态生成式 GFM,以双尺度 token/pixel 目标进行预训练,支持任意模态到任意模态的跨模态生成(Thinking-in-Modalities),可在推理阶段推断缺失传感器数据。本研究未报告单一排行榜,而是考察二者在十个涵盖分割与回归任务的多样化应用场景(包括气候灾害响应、甲烷泄漏检测、积雪监测及海冰制图)中,沿若干关键维度的实际差异:图像块尺寸、解码器复杂度、微调策略、输入模态及模型规模。结果表明:架构设计选择——尤其是图像块尺寸与解码器类型——所解释的性能方差大于模型身份本身;两类模型体现了互补的投资策略(TerraMind 侧重预训练阶段的模型规模,THOR 侧重推理阶段的 token 化灵活性);正确解读结果需以数据集层级的特征刻画为前提。最终呈现的并非单一优胜者,而是一组可迁移的假设及一种诊断性消融方法论,预期可推广至 THOR 与 TerraMind 之外的未来 GFM。
Benchmarks for Geospatial Foundation Models (GFMs) increasingly rank models by aggregate score, but such rankings obscure why models differ: how much of the gap is architecture, how much is decoder capacity, and how much is a use-case-specific artefact? This study addresses that gap through a controlled comparison of two GFMs developed under European Space Agency's $Φ$-lab with contrasting design philosophies: THOR, which introduces a compute-adaptive architecture supporting variable patch sizes and unifies Sentinel-1, -2, and -3 data at their native resolutions; and TerraMind, a multimodal generative GFM pretrained with a dual-scale token/pixel objective that enables any-to-any cross-modal generation (Thinking-in-Modalities) to infer missing sensors at inference time. Rather than reporting a single leaderboard, we investigate the axes along which the two architectures actually differ - patch size, decoder complexity, finetuning regime, input modality, and model scale - across ten use cases spanning segmentation and regression in diverse domains, including climate disaster response, methane leak detection, snow monitoring, or sea ice mapping. We find that architectural design choices - patch size and decoder type in particular - explain more performance variance than model identity itself, that the two models embody complementary investment strategies (pretraining-time scale for TerraMind versus inference-time tokenisation for THOR), and that correctly interpreting results requires dataset-level characterisation. The resulting picture is not a single winner but a set of hypotheses and a diagnostic ablation methodology that we expect to generalise to future GFMs beyond THOR and TerraMind.