论文
arXiv
RemoteSensing
EarthObservation
LLM
Multimodal
GeoMultimodal
中文标题
更少,更多:一种采用简洁方案的大规模遥感视觉-语言模型
English Title
More with Less: a Large Scale Remote Sensing VLM with a Simple Recipe
Stefan Maria Ailuro, Mario Markov, Mohammad Mahdi, Luc Van Gool, Danda Pani Paudel
发布时间
2026/7/17 21:25:44
来源类型
preprint
语言
en
摘要
中文对照

遥感视觉-语言模型(Remote Sensing Vision-Language Models, VLMs)日益被期望支持对地球观测(Earth Observation, EO)数据及多种任务的开放式推理。该领域近期的多数进展依赖于面向遥感任务定制的架构设计,常引入新的编码器、对齐模块或任务特定的融合机制。本文质疑此类架构特化的必要性。我们证明,一个通用型视觉-语言模型只要在足够大规模、多样化数据与任务上进行训练,即可在具有挑战性的遥感基准测试中达到具备竞争力甚至当前最优(state-of-the-art)的性能。本模型采用单一语言策略,可直接以文本形式作答,亦可调用定位工具执行分割与空间定位(grounding)。为训练这种异构行为,我们构建了一个多任务强化学习框架,并采用覆盖多项选择型视觉问答(VQA)、自由形式视觉问答、图像描述生成(captioning)、目标检测与分割等任务的自适应任务奖励机制,输入类型涵盖广泛。本方法在一系列基准测试中均取得具备竞争力的结果,包括高分辨率、多时相、多模态与多视角任务。此外,随着训练数据规模扩大,实验显示模型在分布内与分布外的大多数任务上均呈现持续提升趋势,且性能增益与各任务所对应数据的多样性呈正相关。这些发现表明,对于遥感视觉-语言模型而言,数据规模的重要性高于架构创新。

English Original

Remote sensing vision-language models are increasingly expected to support open-ended reasoning over Earth Observation data and a variety of tasks. Most recent progress in this area has been driven by remote-sensing-specific architectural designs, often introducing new encoders, alignment modules, or task-specific fusion mechanisms. In this work, we challenge the necessity of such architectural specialization. We show that a generally capable vision-language model can achieve competitive or state-of-the-art performance at challenging remote sensing benchmarks, provided that it is trained at sufficient scale across diverse data and tasks. Our model uses a single language policy that can either answer directly in text or invoke a localization tool for segmentation and grounding. To train this heterogeneous behaviour, we employ a multi-task reinforcement learning framework with adaptive task rewards covering multiple-choice VQA, free-form VQA, captioning, detection, and segmentation across a large variety of input types. Our approach achieves competitive results across a broad set of benchmarks, including high-resolution, multi-temporal, multi-modal and multi-view tasks. Further, as training data scales, our experiments show consistent improvements across most tasks both in and out of distribution, which correlate with per-task data diversity. These findings suggest that, for remote sensing VLMs, data scale is more important than architectural novelty.

元数据
arXiv2607.15942v1
来源arXiv
类型论文
抽取状态raw
关键词
RemoteSensing
EarthObservation
LLM
Multimodal
GeoMultimodal
cs.CV
cs.LG