自主地球观测(Earth Observation, EO)智能体正从被动感知转向复杂、多步骤任务执行。然而,当前将规划与执行集成于单一模型的架构在动态EO场景中常面临组合爆炸与推理错误等挑战。为应对这些问题,我们提出轻量级多模态元规划框架(Lightweight Multimodal Meta-Planner, LMMP)。LMMP引入双感知机制,使战略规划同时锚定于多模态图像特征与高层任务语义。关键在于,我们构建了元任务库(Meta Task Library),将遥感领域专家知识直接注入工作流,从而标准化领域逻辑并确保规划具备物理可行性。此外,我们设计两阶段训练流程:首先通过专家蒸馏的监督微调(Supervised Fine-Tuning)初始化元规划器,再基于执行反馈采用直接偏好优化(Direct Preference Optimization)进行精调。在源自EarthBench与ThinkGeo的数据集上开展的大量实验表明,LMMP显著提升了工具调用准确率与任务成功率。该框架还展现出优异的“即插即用”通用性,能在先前未见的EO任务中持续提升多种执行器主干网络的性能。
Autonomous Earth Observation (EO) agents are transitioning from passive perception to complex, multi-step task execution. However, current architectures that integrate planning and execution within a single model often struggle with combinatorial complexity and reasoning errors in dynamic EO scenarios. To resolve these challenges, we propose the Lightweight Multimodal Meta-Planner (LMMP) framework. LMMP incorporates a dual-awareness mechanism that grounds strategic plans in both multimodal image features and high-level task semantics. Crucially, we introduce a Meta Task Library to inject remote sensing expert knowledge directly into the workflow, which standardizes domain logic and ensures plans are physically feasible. We further implement a two-stage training pipeline, initializing the Meta-Planner via expert-distilled Supervised Fine-Tuning and refining it through Direct Preference Optimization based on execution feedback. Extensive experiments on a dataset derived from EarthBench and ThinkGeo demonstrate that LMMP significantly improves tool-calling accuracy and task success rates. Moreover, the framework exhibits strong ``plug-and-play'' versatility, consistently enhancing the performance of diverse executor backbones across previously unseen EO missions.