卫星基础模型为通勤起讫地(OD)生成任务提供了全球可获取的普查数据替代方案,但尚无研究在统一下游任务流程中系统比较不同编码器范式。我们在相同的WeDAN图扩散框架下,对四种卫星视觉编码器——语言监督型(RemoteCLIP)、自监督型(DINOv3)及地理感知型(SatCLIP、AlphaEarth)——开展消融实验,覆盖1925个美国县、325个英国行政区及14个全球城市,并在五个随机种子下进行评估。主要发现有三:第一,语言监督特征在分布内性能最强(RemoteCLIP的CPC为0.602),而地理感知编码器零样本迁移更稳健:AlphaEarth在英国行政区上的CPC较RemoteCLIP提升33%;第二,预训练语料规模本身并不充分:DINOv3虽在更大规模卫星语料上训练,其分布内CPC仍比RemoteCLIP低0.091,且在全球范围内坍塌至CPC 0.022;第三,所有编码器均无法有效迁移至全球城市(RemoteCLIP最佳CPC为0.122,DINOv3为0.022),证实跨洲OD生成仍是开放问题。此外,我们厘清了人口普查噪声参数$η$的语义:其排序在跨洲评估中发生反转,该差异对正确解读既有结果至关重要。训练脚本与评估日志将开源发布。
Satellite foundation models offer a globally available alternative to census data for commuting origin-destination (OD) generation, yet no study has systematically compared encoder paradigms within a single downstream pipeline. We ablate four satellite vision encoders: language-supervised (RemoteCLIP), self-supervised (DINOv3), and geographically grounded (SatCLIP, AlphaEarth) within an identical WeDAN graph diffusion framework across 1,925 US counties, 325 UK districts, and 14 global cities under five random seeds. Three main findings emerge. First, language-supervised features achieve the strongest in-distribution performance (RemoteCLIP CPC 0.602), while geographically grounded encoders transfer more reliably zero-shot: AlphaEarth improves CPC by 33% over RemoteCLIP on UK districts. Second, pretraining corpus scale alone is insufficient: DINOv3, trained on a substantially larger satellite corpus, underperforms RemoteCLIP by 0.091 CPC in-distribution and collapses to CPC 0.022 globally. Third, no encoder transfers usefully to global cities (best CPC 0.122 for RemoteCLIP, 0.022 for DINOv3), confirming cross-continental OD generation remains an open problem. We additionally clarify the semantics of the census noise parameter $η$, whose ordering reverses under cross-continental evaluation, a distinction critical to correctly interpreting prior results. Training scripts and evaluation logs will be released.