人类世的一个核心挑战在于将物理地球系统与人类社会建模为一个耦合系统,然而目前尚无学习所得的表征能够覆盖二者在观测尺度上的全部广度。我们认为其根本障碍在于几何结构差异:物理地球以忽略政治边界的连续场形式被测量,而社会数据则按行政单元(如国家)上报。现有地球系统基础模型适配前一种几何结构;将其与后一种几何结构耦合时,不得不依赖有损的边界平均操作。本文提出 TerraNova——一种在原始几何结构下训练的基础模型,输入包含 1,024 类物理与社会记录:其中 512 类为格网化地球系统场,另 512 类为国家级指标。专用编码器分别表征位置、国家、时间和任务;跨模态 Transformer 将其融合为统一的时空状态;超网络则为每个查询生成专属解码器,其证据性输出头返回预测分布。该表征通过两项对比学习目标实现耦合:一是基于人口加权的国家与其境内地理坐标的对齐;二是与预训练的、承载图像语义的地理空间嵌入对齐。经该解码器读出后,该表征在性能上可媲美专用于地理空间任务的编码器,同时拓展了后者未涵盖的维度(时间、海洋及不确定性),并支持国家级别的建模能力。冻结主干网络可在消费级硬件上数分钟内完成从稀疏观测重建稠密场,并快速适配至未见变量。
A defining problem of the Anthropocene is to model the physical Earth and human societies as one coupled system, yet no learned representation spans their observational breadth. We argue the obstacle is geometric: the physical Earth is measured as continuous fields that ignore political borders, whereas societies are reported for administrative units. Earth-system foundation models serve the first geometry; coupling it to the second has required lossy averaging over borders. We introduce TerraNova, a foundation model trained on 1,024 physical and societal records in their native geometries: 512 gridded Earth-system fields and 512 national indicators. Dedicated encoders represent location, country, time and task, cross-modal transformers fuse them into a shared spatiotemporal state, and a hypernetwork generates a per-query decoder whose evidential head returns a predictive distribution. Two contrastive objectives couple the representation: a population-weighted alignment between each country and coordinates in its territory, and one to pretrained geospatial embeddings carrying image-derived semantics. Read out through that decoder, the representation is competitive with purpose-built geospatial encoders while spanning axes they do not represent (time, oceans and uncertainty) and supporting country-level capabilities. The frozen backbone reconstructs dense fields from sparse observations and adapts to unseen variables in minutes on consumer hardware.