Please use this identifier to cite or link to this item: http://hdl.handle.net/10397/119951
PIRA download icon_1.1View/Download Full Text
Title: Adapting self-supervised vision foundation models for cross-domain surface water mapping
Authors: Lu, X 
Weng, Q 
Issue Date: 2026
Source: IEEE transactions on geoscience and remote sensing, 2026, v. 64, 5624114
Abstract: Accurate and scalable mapping of surface water is crucial for understanding environmental dynamics, ensuring sustainable water resource management, and facilitating informed decision-making. However, the extraction of surface water at the global scale remains challenging due to the scarcity of labeled samples, as well as substantial spatial heterogeneity across regions. Traditional deep learning methods frequently fail to generalize effectively under these conditions. With the advent of vision foundation models (VFMs), such as DINOv3, visual representation learning has demonstrated enhanced generalization and transferability across domains. However, their potential for downstream remote sensing applications, particularly in surface water mapping, remains underexplored. In this study, we propose SWM-DINOv3, a framework designed to efficiently adapt the self-supervised DINOv3 model for large-scale, cross-domain surface water mapping (SWM). The framework employs a pre-trained DINOv3 as the encoder and a LinkNet architecture as the decoder. The pre-trained DINOv3 parameters are kept frozen, while parameter-efficient fine-tuning is performed using a general-purpose low-rank adaptation (LoRA) module and two task-specific adapters. Among these, a pyramid pooling adapter (PPA) is integrated to address the large-scale variability of water bodies by extracting multi-scale contextual features, while an atrous depthwise convolution adapter (ADCA) introduces lightweight local inductive biases, improving boundary delineation for water bodies with low texture and fragmented shorelines. Extensive experiments conducted on two global-scale datasets—OpenEarthMap and GLH-Water—demonstrate that SWM-DINOv3 effectively leverages the generalization capabilities of vision foundation models, achieving superior accuracy and robustness in cross-domain surface water mapping compared to existing VFM-based approaches. Notably, SWM-DINOV3 outperforms Segment Anything Model (SAM)-based methods by over 10% in F1-score. This work establishes a pivotal and scalable solution for global water resource monitoring and management.
Keywords: Crossdomain
DINOv3
Parameter efficient fine-tuning
Satellite mapping
Surface Water Mapping (SWM)
Vision foundation models (VFMs)
Publisher: Institute of Electrical and Electronics Engineers
Journal: IEEE transactions on geoscience and remote sensing 
ISSN: 0196-2892
EISSN: 1558-0644
DOI: 10.1109/TGRS.2026.3695620
Rights: © 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.
The following publication X. Lu and Q. Weng, 'Adapting Self-Supervised Vision Foundation Models for Cross-Domain Surface Water Mapping,' in IEEE Transactions on Geoscience and Remote Sensing, vol. 64, Art no. 5624114, 2026 is available at https://doi.org/10.1109/TGRS.2026.3695620.
Appears in Collections:Journal/Magazine Article

Files in This Item:
File Description SizeFormat 
Lu_Adapting_Self-supervised_Vision.pdfPre-Published version8.82 MBAdobe PDFView/Open
Open Access Information
Status open access
File Version Final Accepted Manuscript
Access
View full-text via PolyU eLinks SFX Query
Show full item record

Google ScholarTM

Check

Altmetric


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.