Publications
ACM MM 2026 HiRS-Agent: A Hierarchical Multi-Agent System for Reliable Long-Horizon Remote Sensing Task Solving
We propose HiRS-Agent, a hierarchical multi-agent system with RS-specialized execution and verification-guided control for reliable long-horizon remote sensing task solving.
SCIS RS-Agent: Automating Remote Sensing Tasks through Intelligent Agent
A domain-adapted agent that connects user intent to professional remote sensing workflows through a central controller, a dynamic toolkit, a solution space of expert guidance, and a domain knowledge space.
ACM MM 2026 GeoChrono: Benchmarking and Rethinking Long-Term Temporal Understanding in Remote Sensing
We introduce ChronoBench, a multidimensional benchmark that decomposes long-term remote sensing understanding into four progressive cognitive levels, and GeoChrono, an MLLM that traces, memorizes, and reasons about long-term geographic evolution.
ACM MM 2026 Self-in-Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence
We introduce SIS-Bench, a large-scale benchmark for evaluating self-awareness and spatial cognition in UAV vision-language models through real-world aerial video understanding.
IJCV Structural Pruning of Large Vision Language Models: A Comprehensive Study on Pruning Dynamics, Recovery, and Data Efficiency
A comprehensive study of compressing LVLMs by structurally pruning the language backbone and recovering with lightweight finetuning and distillation, characterizing pruning dynamics and data efficiency.
ICASSP 2026 BTCChat: Advancing Remote Sensing Bi-Temporal Change Captioning with Multimodal Large Language Model
A multi-temporal MLLM for bi-temporal change captioning that replaces naive image-pair concatenation with a dedicated change extraction module, while retaining single-image interpretation ability.
Journal of Radars A Survey on Earth Observation Multimodal Large Language Models: Framework, Core Technologies, and Future Perspectives
A comprehensive survey of Earth observation multimodal large language models, covering architectures, training strategies, benchmark tasks, and future research directions.
ICML 2026 Dissecting Multimodal In-Context Learning: Modality Asymmetries and Circuit Dynamics in Modern Transformers
Controlled experiments on small transformers reveal an asymmetry in multimodal in-context learning: with a high-diversity primary modality, surprisingly low data complexity in the secondary modality is enough for ICL to emerge.
GCPR 2025 Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models
A study of layerwise versus widthwise structural pruning on the language backbone of MLLMs, paired with supervised finetuning and knowledge distillation for recovery on a small fraction of the data.
TVT Online Specific Emitter Identification via Collision-Alleviated Signal Hash
We propose a novel hash-based model, Collision-Alleviated Signal Hash (CASH), providing a unified approach for addressing both online few-shot and generalized zero-shot learning tasks.
IJCV Group-based Distinctive Image Captioning with Memory Difference Encoding and Attention
A group-based differential captioning method that encodes the difference between an image and its similar-image group in memory, so the generated caption foregrounds what is actually unique.
GIS Multi-Row Labeling with Semantic Analysis: A Case Study on Chinese POIs
A multi-row map labeling algorithm that adds linguistic and semantic analysis to Chinese POI label preprocessing, segmentation and placement, so long descriptive names stay readable on the map.
TGRS TS-SatMVSNet: Slope Aware Height Estimation for Large-Scale Earth Terrain Multi-View Stereo
A terrain-aware multi-view stereo network that injects slope information into height estimation, addressing the terrain characteristics that generic learning-based MVS frameworks overlook.
Dataset UAV-VisLoc: A Large-Scale Dataset for UAV Visual Localization
A large-scale, cross-region dataset for UAV visual localization, pairing down-looking UAV imagery with ortho satellite maps so drones can recover their coordinates when GNSS is unavailable.
ISPRS Edge Aware Depth Inference for Large-Scale Aerial Building Multi-View Stereo
An edge-aware multi-view stereo method for large-scale aerial building reconstruction, targeting the depth discontinuities at building boundaries that generic MVS pipelines tend to smooth away.
Pattern Recognition ARAI-MVSNet: A Multi-View Stereo Depth Estimation Network with Adaptive Depth Range and Depth Interval
A multi-stage coarse-to-fine MVS framework that adaptively predicts the all-pixel depth range and refines depth intervals per stage, instead of using a fixed range with equal partitions.
ISPRS Deep Semantic-Visual Alignment for Zero-Shot Remote Sensing Image Scene Classification
A zero-shot scene classification model for remote sensing that automatically collects visually detectable attributes and aligns them with background context, rather than relying on hand-labelled attributes.
ECCV 2022 Distinctive Image Captioning via CLIP Guided Group Optimization
We use CLIP to quantify caption distinctiveness and propose a group-optimization training strategy that widens the embedding gap between a target image and its similar-image group.
IJCV Attribute Prototype Network for Any-Shot Learning
An extension of the attribute prototype network to the any-shot setting, showing that attribute-localized image representations help both zero-shot and few-shot classification of novel classes.
CVPR 2022 VGSE: Visually-Grounded Semantic Embeddings for Zero-Shot Learning
VGSE discovers semantic embeddings with discriminative visual properties without any human annotation, by clustering local image regions from seen classes and relating those clusters to unseen classes.
TPAMI On Distinctive Image Captioning via Comparing and Reweighting
A journal extension of CIDErBtw showing that human annotations differ in distinctiveness, and that reweighting training captions accordingly improves both distinctiveness and standard captioning accuracy.
ACM MM 2021 Group-based Distinctive Image Captioning with Memory Attention
GdisCap compares each image against a group of similar images and uses memory attention to highlight the unique objects and relations, producing captions that separate near-duplicate images.
NeurIPS 2020 Attribute Prototype Network for Zero-Shot Learning
A zero-shot representation learning framework that jointly learns global visual-semantic embeddings and local features through an attribute prototype network that regresses and decorrelates attributes.
ECCV 2020 Compare and Reweight: Distinctive Image Captioning Using Similar Images Sets
We introduce between-set CIDEr (CIDErBtw), a metric for caption distinctiveness, and weighted training strategies that push captioning models to describe what makes each image unique.
ECCV 2020 Generating Visual and Semantic Explanations with Multi-Task Network
A multi-task network that produces textual explanations for its predictions and grounds the mentioned attributes visually, so the two modalities of explanation reinforce each other.
JSTSP Where Is the Model Looking At? Concentrate and Explain the Network Attention
An Explainable Attribute-based Multi-task (EAT) framework that concentrates model attention on discriminative image regions and grounds attribute-based textual explanations back onto the image.
No publications match your search yet.