Publications

ACM MM 2026

HiRS-Agent: A Hierarchical Multi-Agent System for Reliable Long-Horizon Remote Sensing Task Solving

We propose HiRS-Agent, a hierarchical multi-agent system with RS-specialized execution and verification-guided control for reliable long-horizon remote sensing task solving.

ACM MM 2026
SCIS

RS-Agent: Automating Remote Sensing Tasks through Intelligent Agent

A domain-adapted agent that connects user intent to professional remote sensing workflows through a central controller, a dynamic toolkit, a solution space of expert guidance, and a domain knowledge space.

Science China Information Sciences
ACM MM 2026

GeoChrono: Benchmarking and Rethinking Long-Term Temporal Understanding in Remote Sensing

We introduce ChronoBench, a multidimensional benchmark that decomposes long-term remote sensing understanding into four progressive cognitive levels, and GeoChrono, an MLLM that traces, memorizes, and reasons about long-term geographic evolution.

ACM MM 2026
ACM MM 2026

Self-in-Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence

We introduce SIS-Bench, a large-scale benchmark for evaluating self-awareness and spatial cognition in UAV vision-language models through real-world aerial video understanding.

ACM MM 2026
IJCV

Structural Pruning of Large Vision Language Models: A Comprehensive Study on Pruning Dynamics, Recovery, and Data Efficiency

A comprehensive study of compressing LVLMs by structurally pruning the language backbone and recovering with lightweight finetuning and distillation, characterizing pruning dynamics and data efficiency.

International Journal of Computer Vision
ICASSP 2026

BTCChat: Advancing Remote Sensing Bi-Temporal Change Captioning with Multimodal Large Language Model

A multi-temporal MLLM for bi-temporal change captioning that replaces naive image-pair concatenation with a dedicated change extraction module, while retaining single-image interpretation ability.

ICASSP 2026
Journal of Radars

A Survey on Earth Observation Multimodal Large Language Models: Framework, Core Technologies, and Future Perspectives

A comprehensive survey of Earth observation multimodal large language models, covering architectures, training strategies, benchmark tasks, and future research directions.

Journal of Radars
ICML 2026

Dissecting Multimodal In-Context Learning: Modality Asymmetries and Circuit Dynamics in Modern Transformers

Controlled experiments on small transformers reveal an asymmetry in multimodal in-context learning: with a high-diversity primary modality, surprisingly low data complexity in the secondary modality is enough for ICL to emerge.

ICML 2026
GCPR 2025

Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models

A study of layerwise versus widthwise structural pruning on the language backbone of MLLMs, paired with supervised finetuning and knowledge distillation for recovery on a small fraction of the data.

DAGM GCPR 2025
TVT

Online Specific Emitter Identification via Collision-Alleviated Signal Hash

We propose a novel hash-based model, Collision-Alleviated Signal Hash (CASH), providing a unified approach for addressing both online few-shot and generalized zero-shot learning tasks.

IEEE Transactions on Vehicular Technology
IJCV

Group-based Distinctive Image Captioning with Memory Difference Encoding and Attention

A group-based differential captioning method that encodes the difference between an image and its similar-image group in memory, so the generated caption foregrounds what is actually unique.

International Journal of Computer Vision
GIS

Multi-Row Labeling with Semantic Analysis: A Case Study on Chinese POIs

A multi-row map labeling algorithm that adds linguistic and semantic analysis to Chinese POI label preprocessing, segmentation and placement, so long descriptive names stay readable on the map.

Transactions in GIS
TGRS

TS-SatMVSNet: Slope Aware Height Estimation for Large-Scale Earth Terrain Multi-View Stereo

A terrain-aware multi-view stereo network that injects slope information into height estimation, addressing the terrain characteristics that generic learning-based MVS frameworks overlook.

IEEE Transactions on Geoscience and Remote Sensing
Dataset

UAV-VisLoc: A Large-Scale Dataset for UAV Visual Localization

A large-scale, cross-region dataset for UAV visual localization, pairing down-looking UAV imagery with ortho satellite maps so drones can recover their coordinates when GNSS is unavailable.

arXiv preprint
ISPRS

Edge Aware Depth Inference for Large-Scale Aerial Building Multi-View Stereo

An edge-aware multi-view stereo method for large-scale aerial building reconstruction, targeting the depth discontinuities at building boundaries that generic MVS pipelines tend to smooth away.

ISPRS Journal of Photogrammetry and Remote Sensing
Pattern Recognition

ARAI-MVSNet: A Multi-View Stereo Depth Estimation Network with Adaptive Depth Range and Depth Interval

A multi-stage coarse-to-fine MVS framework that adaptively predicts the all-pixel depth range and refines depth intervals per stage, instead of using a fixed range with equal partitions.

Pattern Recognition
ISPRS

Deep Semantic-Visual Alignment for Zero-Shot Remote Sensing Image Scene Classification

A zero-shot scene classification model for remote sensing that automatically collects visually detectable attributes and aligns them with background context, rather than relying on hand-labelled attributes.

ISPRS Journal of Photogrammetry and Remote Sensing
ECCV 2022

Distinctive Image Captioning via CLIP Guided Group Optimization

We use CLIP to quantify caption distinctiveness and propose a group-optimization training strategy that widens the embedding gap between a target image and its similar-image group.

ECCV 2022 Workshops
IJCV

Attribute Prototype Network for Any-Shot Learning

An extension of the attribute prototype network to the any-shot setting, showing that attribute-localized image representations help both zero-shot and few-shot classification of novel classes.

International Journal of Computer Vision
CVPR 2022

VGSE: Visually-Grounded Semantic Embeddings for Zero-Shot Learning

VGSE discovers semantic embeddings with discriminative visual properties without any human annotation, by clustering local image regions from seen classes and relating those clusters to unseen classes.

CVPR 2022
TPAMI

On Distinctive Image Captioning via Comparing and Reweighting

A journal extension of CIDErBtw showing that human annotations differ in distinctiveness, and that reweighting training captions accordingly improves both distinctiveness and standard captioning accuracy.

IEEE Transactions on Pattern Analysis and Machine Intelligence
ACM MM 2021

Group-based Distinctive Image Captioning with Memory Attention

GdisCap compares each image against a group of similar images and uses memory attention to highlight the unique objects and relations, producing captions that separate near-duplicate images.

ACM MM 2021
NeurIPS 2020

Attribute Prototype Network for Zero-Shot Learning

A zero-shot representation learning framework that jointly learns global visual-semantic embeddings and local features through an attribute prototype network that regresses and decorrelates attributes.

NeurIPS 2020
ECCV 2020

Compare and Reweight: Distinctive Image Captioning Using Similar Images Sets

We introduce between-set CIDEr (CIDErBtw), a metric for caption distinctiveness, and weighted training strategies that push captioning models to describe what makes each image unique.

ECCV 2020
ECCV 2020

Generating Visual and Semantic Explanations with Multi-Task Network

A multi-task network that produces textual explanations for its predictions and grounds the mentioned attributes visually, so the two modalities of explanation reinforce each other.

ECCV 2020
JSTSP

Where Is the Model Looking At? Concentrate and Explain the Network Attention

An Explainable Attribute-based Multi-task (EAT) framework that concentrates model attention on discriminative image regions and grounds attribute-based textual explanations back onto the image.

IEEE Journal of Selected Topics in Signal Processing