遥感视觉-语言模型已推动地球观测发展,但现有大规模视觉-语言资源仍以RGB为中心,导致互补的红外信息未被充分挖掘。红外观测可提供独特的强度结构、物体边界及光照不变线索,从而补充传统RGB影像,然而大规模RGB-红外-文本资源依然稀缺。我们提出FusionRS,这是首个面向受控双模态遥感视觉-语言学习的大规模RGB-红外-文本数据集。该数据集包含600,000组空间对齐的图像对,由将多样化的公开RGB遥感影像转换为红外风格图像生成。每组图像对均保留常规场景描述文本,其中精选子集额外提供45,913条红外感知型描述文本,涵盖可观测的强度、对比度、纹理与结构特征,同时保持场景语义一致性。我们训练了类CLIP模型以实现RGB-红外-文本对齐,并采用混合任务条件化字幕监督方式适配生成式视觉-语言模型。评估涵盖跨模态检索、缩放与监督消融实验、传感器实采数据迁移、以及严格留出的图像描述与视觉问答任务。相较于仅使用RGB或非红外感知设置,FusionRS显著提升了RGB-红外风格对齐性能及红外到文本的检索效果。消融实验表明,红外感知型描述文本可提升任务条件化红外描述质量,验证了模态特异性监督的价值。FusionRS为受控的RGB-红外遥感视觉-语言学习提供了可扩展的基础支撑。
Remote sensing vision-language models have advanced Earth observation, but available large-scale vision-language resources remain RGB-centered, leaving complementary infrared information underexplored. Infrared observations provide distinctive intensity structures, object boundaries, and illumination-invariant cues that complement conventional RGB imagery, yet large-scale RGB-infrared-text resources remain scarce. We introduce FusionRS, the first large-scale RGB-infrared-style-text dataset for controlled dual-modal remote sensing vision-language learning. It contains 600,000 spatially aligned pairs created by translating diverse public RGB remote sensing images into infrared-style counterparts. Each pair retains a conventional scene caption, and a curated subset adds 45,913 IR-aware captions describing observable intensity, contrast, texture, and structure while preserving scene semantics. We train CLIP-style models for RGB-infrared-style-text alignment and adapt a generative vision-language model with mixed task-conditioned caption supervision. Evaluation covers cross-modal retrieval, scaling and supervision ablations, sensor-captured transfer, and strictly held-out captioning and VQA. FusionRS substantially improves RGB-infrared-style alignment and infrared-to-text retrieval over RGB-only and non-IR-aware settings. Ablations show that IR-aware captions improve task-conditioned infrared description, demonstrating the value of modality-specific supervision. FusionRS provides a scalable foundation for controlled RGB-infrared remote sensing vision-language learning.