Please use this identifier to cite or link to this item:
http://hdl.handle.net/10397/119951
| Title: | Adapting self-supervised vision foundation models for cross-domain surface water mapping | Authors: | Lu, X Weng, Q |
Issue Date: | 2026 | Source: | IEEE transactions on geoscience and remote sensing, 2026, v. 64, 5624114 | Abstract: | Accurate and scalable mapping of surface water is crucial for understanding environmental dynamics, ensuring sustainable water resource management, and facilitating informed decision-making. However, the extraction of surface water at the global scale remains challenging due to the scarcity of labeled samples, as well as substantial spatial heterogeneity across regions. Traditional deep learning methods frequently fail to generalize effectively under these conditions. With the advent of vision foundation models (VFMs), such as DINOv3, visual representation learning has demonstrated enhanced generalization and transferability across domains. However, their potential for downstream remote sensing applications, particularly in surface water mapping, remains underexplored. In this study, we propose SWM-DINOv3, a framework designed to efficiently adapt the self-supervised DINOv3 model for large-scale, cross-domain surface water mapping (SWM). The framework employs a pre-trained DINOv3 as the encoder and a LinkNet architecture as the decoder. The pre-trained DINOv3 parameters are kept frozen, while parameter-efficient fine-tuning is performed using a general-purpose low-rank adaptation (LoRA) module and two task-specific adapters. Among these, a pyramid pooling adapter (PPA) is integrated to address the large-scale variability of water bodies by extracting multi-scale contextual features, while an atrous depthwise convolution adapter (ADCA) introduces lightweight local inductive biases, improving boundary delineation for water bodies with low texture and fragmented shorelines. Extensive experiments conducted on two global-scale datasets—OpenEarthMap and GLH-Water—demonstrate that SWM-DINOv3 effectively leverages the generalization capabilities of vision foundation models, achieving superior accuracy and robustness in cross-domain surface water mapping compared to existing VFM-based approaches. Notably, SWM-DINOV3 outperforms Segment Anything Model (SAM)-based methods by over 10% in F1-score. This work establishes a pivotal and scalable solution for global water resource monitoring and management. | Keywords: | Crossdomain DINOv3 Parameter efficient fine-tuning Satellite mapping Surface Water Mapping (SWM) Vision foundation models (VFMs) |
Publisher: | Institute of Electrical and Electronics Engineers | Journal: | IEEE transactions on geoscience and remote sensing | ISSN: | 0196-2892 | EISSN: | 1558-0644 | DOI: | 10.1109/TGRS.2026.3695620 | Rights: | © 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works. The following publication X. Lu and Q. Weng, 'Adapting Self-Supervised Vision Foundation Models for Cross-Domain Surface Water Mapping,' in IEEE Transactions on Geoscience and Remote Sensing, vol. 64, Art no. 5624114, 2026 is available at https://doi.org/10.1109/TGRS.2026.3695620. |
| Appears in Collections: | Journal/Magazine Article |
Files in This Item:
| File | Description | Size | Format | |
|---|---|---|---|---|
| Lu_Adapting_Self-supervised_Vision.pdf | Pre-Published version | 8.82 MB | Adobe PDF | View/Open |
Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.



