论文
arXiv
RemoteSensing
EarthObservation
Trajectory
Mobility
LLM
Multimodal
GeoMultimodal
中文标题
VPRef:面向指代遥感图像分割的跨域基准
English Title
VPRef: A Cross-Domain Benchmark for Referring Remote Sensing Image Segmentation
Quanwei Liu, Tao Huang, Jiaqi Yang, Wei Xiang
发布时间
2026/9/15 09:18:06
来源类型
preprint
语言
en
摘要
中文对照

视觉-语言模型的快速进展推动了指代遥感图像分割(RRSIS)在地球观测领域的前沿发展。然而,在实际部署中,受耦合双重漂移范式的影响,模型性能出现严重退化:一方面是由跨空间分辨率失配和光谱变化引起的视觉域漂移,另一方面是由不受约束且粒度多变的用户输入导致的文本逻辑漂移。为缓解这些瓶颈,本文建立了首个跨域 RRSIS 基准数据集,命名为 Vaihingen-Potsdam Referring (VPRef) 数据集,包含 46,972 个语言-图像-标注三元组,并按三级语言层次结构组织。基于该基准,我们开发了一种以 Segment Anything Model (SAM3) 为核心、采用低秩适应(LoRA)技术的定制化参数高效域自适应基线。我们的框架通过伪标签驱动的自训练来抵消视觉分布差异,并通过随机多粒度文本提示混合来解决文本逻辑漂移问题。重要的是,消融变体中经验指标的分布表明,跨模态语义鲁棒性增强与视觉域对齐之间存在潜在的解耦关系,证明了语言方差驱动细粒度语义不变性,而伪标签传播主导宏观尺度的空间网格对齐。广泛的基准测试显示,所提出的框架仅修改了基础参数足迹的 1.08%,却实现了优越的跨域分割边界,为未来的多模态遥感域自适应研究确立了稳健的基线。数据集和代码将在 https://github.com/quanweiliu/VPRef 提供。

English Original

Rapid advancements in vision-language models have propelled Referring Remote Sensing Image Segmentation (RRSIS) to the forefront of Earth observation. However, practical deployments suffer severe performance degradation under a coupled dual-drift paradigm: visual domain drift from cross-spatial-resolution mismatches and spectral variations, alongside textual logic drift from unconstrained, variable user-input granularities. To mitigate these bottlenecks, this paper establishes the first cross-domain RRSIS benchmark, designated as the Vaihingen-Potsdam Referring (VPRef) dataset, comprising 46,972 language-image-annotation triplets organized into a three-tier linguistic hierarchy. Building upon this benchmark, we develop a tailored parameter-efficient domain adaptation baseline anchored on the Segment Anything Model (SAM3) via Low-Rank Adaptation (LoRA). Our framework counteracts visual distribution discrepancies through pseudo-label-driven self-training and addresses textual logic drift via random multi-granularity text prompt mixing. Crucially, the distribution of empirical metrics across ablative variants suggests a potential decoupling between cross-modal semantic robustification and visual domain alignment, demonstrating that linguistic variance drives fine-grained semantic invariance while pseudo-label propagation governs macro-scale spatial grid alignment. Extensive benchmarks demonstrate the proposed framework achieves superior cross-domain segmentation boundaries while modifying merely 1.08\% of the foundational parameter footprint, establishing a robust baseline for future multi-modal remote sensing domain adaptation research. The dataset and code will be available at https://github.com/quanweiliu/VPRef.

我的阅读记录

正在加载阅读记录…

元数据
arXiv2609.16486v1
来源arXiv
类型论文
抽取状态raw
关键词
RemoteSensing
EarthObservation
Trajectory
Mobility
LLM
Multimodal
GeoMultimodal
cs.CV