← Publications

Journal of Radars

A Survey on Earth Observation Multimodal Large Language Models: Framework, Core Technologies, and Future Perspectives

A comprehensive survey of Earth observation multimodal large language models, covering architectures, training strategies, benchmark tasks, and future research directions.

Wenjia Xu, Ruiqing Yu, Minghao Xue, Xueyi Wang, Yuanben Zhang, Zhiwei Wei, Zhe Zhang, Mugen Peng, Yirong Wu

February 1, 2026Remote Sensing Multimodal Large Language ModelsEarth ObservationSurveyPaper中文版DOI

多模态对地观测大模型:架构、关键技术和未来展望

Journal of Radars, 2026, 15(1): 361–386. DOI: 10.12000/JR25088

Multimodal Large Language Models (MLLMs) have advanced rapidly, and Earth observation is one of the domains where that progress is most visible. By building bridging mechanisms between large language models and vision models and training the two jointly, Earth observation MLLMs (EO-MLLMs) deeply integrate optical imagery, Synthetic Aperture Radar (SAR) imagery, and text. The result is a paradigm shift in intelligent Earth observation interpretation — from shallow semantic matching toward higher-level understanding grounded in world knowledge.

This survey reviews that shift systematically.

What the survey covers

The figure above traces the development of EO-MLLMs and EO-Agents from 2023 onward, marking where each representative model appeared.