Multi-modal LLMs & Agents
Vision-language multi-modal LLMs, in-context learning, model compression and pruning, and agent systems for real-world tasks such as remote sensing.
Learn moreLed by Dr. Wenjia Xu, our group works at the intersection of computer vision, natural language processing and remote sensing, with a focus on multi-modal large language models, agents, and few-shot learning.
Our work spans vision, language and remote sensing, organized around perception, understanding and decision-making.
Vision-language multi-modal LLMs, in-context learning, model compression and pruning, and agent systems for real-world tasks such as remote sensing.
Learn more
Attribute prototype networks and visual-semantic embeddings that recognize novel categories from few or zero labeled samples.
Learn more
Image captioning, distinctive captioning and visual-semantic alignment - so models not only see, but describe.
Learn more
Super-resolution, UAV visual localization, height estimation and change detection for Earth observation and sustainability.
Learn moreWe are driven by creativity and curiosity to push back the frontiers of knowledge
Associate Professor · PhD Supervisor
School of Information and Communication Engineering, BUPT
Research across computer vision, natural language processing and remote sensing, focusing on agents, multimodal large language models and few-shot learning.
Explore critical questions in Earth observation with outstanding collaborators, turning remote sensing data into knowledge for a more sustainable future.
Fieldwork, conferences and the moments that connect our team across the world.
Follow our latest publications, awards, field activities, student achievements, and research updates from IntelliSensing Lab.
Our latest work in computer vision, language intelligence, multi-modal learning, and remote sensing.
ICCV, 2025 (Best Student Paper Honorable Mention, Best Paper Candidate, Oral Presentation)
View DetailsResearch notes, technical reflections, and ideas taking shape inside IntelliSensing Lab.