地理空间基础模型(GeoFMs)通常仅依据其在标准基准条件下的准确率(通过平均排名)进行排序与选择。本文指出,该评估范式过于狭窄:面向关键地球观测(EO)任务的实际部署,亟需引入额外分析维度,尤其是校准性(calibration),即模型置信度与其预测正确性之间的一致性。我们在16个冻结编码器、4个分类数据集与5个分割数据集、以及两条正交压力轴上开展实验,结果表明:所有编码器均随扰动强度增加而性能下降,且其相对排名亦随之变化。在4个分类基准上,地球观测预训练(EO-pretrained)与ImageNet预训练编码器在干净数据上的准确率与校准性均无显著差异;EO预训练并未比ImageNet预训练提供更强的分布偏移鲁棒性。在分布偏移下,GeoFMs在每一扰动等级及每一类扰动家族中均较ImageNet预训练编码器表现出更严重的过度自信倾向。中心化核对齐(CKA)分析揭示了这一现象的表征根源:EO预训练嵌入在扰动下表征刚性更强(变化更小),但任务信息损失程度相当,且仍维持过度自信状态。我们测试了三种常用不确定性量化方法,发现温度缩放(temperature scaling)与深度集成(deep ensembles)均无法缓解该退化现象;而高斯过程探针(Gaussian-process probe)仅在严重云覆盖扰动下将预期校准误差(ECE)约降低一半,却以干净数据上ECE翻三倍为代价。在选择性预测(selective prediction)实验中,我们发现基于置信度的拒绝机制无法有效规避高置信度错误预测。因此,我们主张:基准测试的排名与评估应覆盖多种条件与指标,以更全面地衡量模型研发进展,并缩小与真实世界部署场景之间的差距。
Geospatial Foundation Models (GeoFMs) are most commonly ranked and selected by accuracy on standard benchmark conditions via averaged ranks. We show that this protocol is too narrow: the promised deployment in critical EO tasks requires further angles of analysis, mainly calibration, the agreement between a model's confidence and its correctness. Across 16 frozen encoders, four classification and five segmentation datasets, and two orthogonal stress axes, every encoder degrades as corruption intensifies, and the ranking changes as well. Across the four classification benchmarks, EO-pretrained and ImageNet-pretrained encoders are indistinguishable on clean accuracy and clean calibration, and EO pretraining provides no more stability under shift than ImageNet pretraining. Under shift the GeoFMs drift further into overconfidence than the ImageNet-pretrained encoders, at every grade and in every corruption family. A centered kernel alignment (CKA) analysis ties this to representational rigidity: EO-pretrained embeddings move less under corruption while losing just as much task information and remaining overconfident. We apply three commonly explored uncertainty quantification methods and find that temperature scaling and deep ensembles cannot counteract the degradation, while a Gaussian-process probe roughly halves ECE under severe cloud only by tripling it on clean data. In selective prediction experiments, we find that confidence-based abstention cannot defer around confidently wrong predictions, and advocate that benchmark rankings and evaluations should therefore operate across a multitude of conditions and metrics to more holistically evaluate model development progress and close the gap to real world deployment scenarios.