Publication GeoChrono: Benchmarking and Rethinking Long-Term Temporal Understanding in Remote Sensing We introduce ChronoBench, a multidimensional benchmark that decomposes long-term remote sensing understanding into four progressive cognitive levels, and GeoChrono, an MLLM that traces, memorizes, and reasons about long-term geographic evolution. Research
Four research directions: multi-modal LLMs, agents, intelligence for LEO communication and network, and Earth observation applications.
Multimodal LLMs
Studying the joint modeling of vision, language, and knowledge to explore how multimodal large language models understand complex scenes, connect information from multiple sources, and reason.
Latest Publications
View all publications
Publication GeoChrono: Benchmarking and Rethinking Long-Term Temporal Understanding in Remote Sensing We introduce ChronoBench, a multidimensional benchmark that decomposes long-term remote sensing understanding into four progressive cognitive levels, and GeoChrono, an MLLM that traces, memorizes, and reasons about long-term geographic evolution.
Publication Self-in-Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence We introduce SIS-Bench, a large-scale benchmark for evaluating self-awareness and spatial cognition in UAV vision-language models through real-world aerial video understanding.
Publication Structural Pruning of Large Vision Language Models: A Comprehensive Study on Pruning Dynamics, Recovery, and Data Efficiency A comprehensive study of compressing LVLMs by structurally pruning the language backbone and recovering with lightweight finetuning and distillation, characterizing pruning dynamics and data efficiency. Latest News
View all news
News ACM MM 2026 | GeoChrono: Benchmarking and Rethinking Long-Term Temporal Understanding in Remote Sensing ChronoBench breaks long-term remote sensing understanding into four cognitive levels and 12 sub-tasks, pinpointing long-term memory as the key bottleneck for MLLMs. GeoChrono reaches 78.34% overall accuracy, more than 20 points above the best commercial model.
News ACM MM 2026 | Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence SIS-Bench evaluates spatial cognition and self-awareness of multimodal LLMs for UAV embodied intelligence across 1,646 real UAV videos and 4,856 QA pairs, while SIS-Motion probes whether explicit optical-flow motion cues help.
News Explainer | From Prompt to Policy: How Reinforcement Learning Teaches LLMs to Call Tools A walk through the 2025 work on reinforcement learning for tool-integrated reasoning, from single-tool agents such as ReTool to multi-tool orchestration — and why prompting and supervised fine-tuning alone do not produce a reliable calling strategy.
Agents
Studying how agents understand tasks, make plans, use tools, and collaborate, and exploring methods for completing complex tasks in open and dynamic environments.
Latest Publications
View all publications
Publication HiRS-Agent: A Hierarchical Multi-Agent System for Reliable Long-Horizon Remote Sensing Task Solving We propose HiRS-Agent, a hierarchical multi-agent system with RS-specialized execution and verification-guided control for reliable long-horizon remote sensing task solving.
Publication RS-Agent: Automating Remote Sensing Tasks through Intelligent Agent A domain-adapted agent that connects user intent to professional remote sensing workflows through a central controller, a dynamic toolkit, a solution space of expert guidance, and a domain knowledge space.
Publication Self-in-Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence We introduce SIS-Bench, a large-scale benchmark for evaluating self-awareness and spatial cognition in UAV vision-language models through real-world aerial video understanding. Latest News
View all news
News ACM MM 2026 Oral | HiRS-Agent: A Hierarchical Multi-Agent System for Reliable Long-Horizon Remote Sensing Task Solving HiRS-Agent puts a Manager over three Specialists so a long remote sensing workflow is planned, routed, executed and checked step by step. Verification at every step lifts long-horizon task accuracy from 15.73% to 43.95%. Accepted as an Oral at ACM Multimedia 2026.
News ACM MM 2026 | Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence SIS-Bench evaluates spatial cognition and self-awareness of multimodal LLMs for UAV embodied intelligence across 1,646 real UAV videos and 4,856 QA pairs, while SIS-Motion probes whether explicit optical-flow motion cues help.
News SCIS | RS-Agent: Automating Remote Sensing Tasks through Intelligent Agent RS-Agent turns a large language model from a passive image-understanding tool into an agent that plans and executes remote sensing tasks, reaching over 95% planning accuracy across 9 datasets and 18 task types. Accepted by Science China Information Sciences.
LEO Communications and Network Intelligence
Studying theories and methods for coordinating sensing, communication, computing, and decision-making under changing connectivity and resource constraints in low-Earth-orbit satellite networks.
Latest Publications
View all publicationsLatest News
View all news
Earth Observation Applications
Integrating satellite observations, UAV data, and geospatial information from multiple sources to investigate methods for characterizing land-surface environments, analyzing changes, and reconstructing spatial structure, supporting the interpretation and application of Earth observation data.
Latest Publications
View all publications
Publication HiRS-Agent: A Hierarchical Multi-Agent System for Reliable Long-Horizon Remote Sensing Task Solving We propose HiRS-Agent, a hierarchical multi-agent system with RS-specialized execution and verification-guided control for reliable long-horizon remote sensing task solving.
Publication RS-Agent: Automating Remote Sensing Tasks through Intelligent Agent A domain-adapted agent that connects user intent to professional remote sensing workflows through a central controller, a dynamic toolkit, a solution space of expert guidance, and a domain knowledge space.
Publication GeoChrono: Benchmarking and Rethinking Long-Term Temporal Understanding in Remote Sensing We introduce ChronoBench, a multidimensional benchmark that decomposes long-term remote sensing understanding into four progressive cognitive levels, and GeoChrono, an MLLM that traces, memorizes, and reasons about long-term geographic evolution. Latest News
View all news
News ACM MM 2026 Oral | HiRS-Agent: A Hierarchical Multi-Agent System for Reliable Long-Horizon Remote Sensing Task Solving HiRS-Agent puts a Manager over three Specialists so a long remote sensing workflow is planned, routed, executed and checked step by step. Verification at every step lifts long-horizon task accuracy from 15.73% to 43.95%. Accepted as an Oral at ACM Multimedia 2026.
News ACM MM 2026 | GeoChrono: Benchmarking and Rethinking Long-Term Temporal Understanding in Remote Sensing ChronoBench breaks long-term remote sensing understanding into four cognitive levels and 12 sub-tasks, pinpointing long-term memory as the key bottleneck for MLLMs. GeoChrono reaches 78.34% overall accuracy, more than 20 points above the best commercial model.
News SCIS | RS-Agent: Automating Remote Sensing Tasks through Intelligent Agent RS-Agent turns a large language model from a passive image-understanding tool into an agent that plans and executes remote sensing tasks, reaching over 95% planning accuracy across 9 datasets and 18 task types. Accepted by Science China Information Sciences.