论文
arXiv
GeoAI
GIS
GeoLargeModel
GeoFoundationModel
中文标题
尚无公认的地空间基础模型(GFM)技术前沿
English Title
No One Knows the State of the Art in Geospatial Foundation Models
Isaac Corley, Nils Lehmann, Caleb Robinson, Gabriel Tseng, Anthony Fuller, Hamed Alemohammad, Evan Shelhamer, Jennifer Marcus, Hannah Kerner
发布时间
2026/5/13 03:29:51
来源类型
preprint
语言
en
摘要
中文对照

地空间基础模型(Geospatial Foundation Models, GFMs)被提出作为适用于灾害响应、土地覆被制图、粮食安全监测及其他高风险地球观测任务的通用骨干模型。然而,现有已发表的相关研究未能为评审者或用户提供足够信息,以判断何种模型适配特定任务。我们认为,目前尚无人能明确界定地空间基础模型的技术前沿。尽管相关方法可能具备实用价值,但GFM领域文献在评估标准、训练与测试协议、公开模型权重以及预训练控制等方面缺乏统一规范,致使模型间无法有效比较或排序。在对152篇论文开展的系统性审计中,我们发现:同一模型、基准与协议下,跨论文报告结果存在至少10分差异的情形达46例;在126篇可提取预训练数据的论文中,94篇所采用的配置未被其他任何论文复用;39%的GFM论文未公开模型权重。此类社区标准缺失问题具备可解性。我们提出六项具体建议:以命名许可证方式发布模型权重;建立共享的核心评估集;明确标注基线结果为“直接复用”或“独立重跑”;强制报告结果方差;构建统一的评估框架(evaluation harness);分离并独立控制数据、架构与算法三类变量。这些缺口源于协作机制失灵,而非任何单一实验室之过;本文作者亦如GFM领域诸多研究者一样,曾无意中促成此类问题。我们不仅旨在指出问题,更致力于提供切实可行的步骤,推动社区就GFM创新路径达成共识。

English Original

Geospatial foundation models (GFMs) have been proposed as generalizable backbones for disaster response, land-cover mapping, food-security monitoring, and other high-stakes Earth-observation tasks. Yet the published work about these models does not give reviewers or users enough information to tell which model fits a given task. We argue that nobody knows what the current state of the art is in geospatial foundation models. The methods may be useful, but the GFM literature does not standardize evaluations, training and testing protocols, released weights, or pretraining controls well enough for anyone to compare or rank them. In a 152-paper audit, we find 46 cross-paper disagreements of at least 10 points for the same model, benchmark, and protocol; 94/126 papers with extractable pretraining data use a configuration no other paper uses; and 39% of GFM papers release no model weights. This lack of community standards can be solved. We propose six concrete expectations: named-license weight release, shared core evaluations, copied-versus-rerun baseline annotations, variance reporting, one shared evaluation harness, and data-vs-architecture-vs-algorithm controls. These gaps are a coordination failure, not a fault of any individual lab; the authors of this paper, like many others in the GFM community, have contributed to them. Rather than just critiquing the community, we aim to provide concrete steps toward a shared understanding of how to innovate GFMs.

元数据
arXiv2605.12678v2
来源arXiv
类型论文
抽取状态raw
关键词
GeoAI
GIS
GeoLargeModel
GeoFoundationModel
cs.CV
cs.CY