Decomposed soft prompt guided fusion enhancing for compositional zero-shot learning

Lu, X; Guo, S; Liu, Z; Guo, J

doi:10.1109/CVPR52729.2023.02256

Please use this identifier to cite or link to this item: http://hdl.handle.net/10397/101453

Title:	Decomposed soft prompt guided fusion enhancing for compositional zero-shot learning
Authors:	Lu, X Guo, S Liu, Z Guo, J
Issue Date:	2023
Source:	2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 17-24 June 2023, p. 23560-23569
Abstract:	Compositional Zero-Shot Learning (CZSL) aims to recognize novel concepts formed by known states and objects during training. Existing methods either learn the combined state-object representation, challenging the generalization of unseen compositions, or design two classifiers to identify state and object separately from image features, ignoring the intrinsic relationship between them. To jointly eliminate the above issues and construct a more robust CZSL system, we propose a novel framework termed Decomposed Fusion with Soft Prompt (DFSP) 1 1 Code is available at: https://github.corn/Forest-art/DFSP.git, by involving vision-language models (VLMs)for unseen composition recognition. Specifically, DFSP constructs a vector combination of learnable soft prompts with state and object to establish the joint representation of them. In addition, a cross-modal decomposed fusion module is designed between the language and image branches, which decomposes state and object among language features instead of image features. Notably, being fused with the decomposed features, the image features can be more expressive for learning the relationship with states and objects, respectively, to improve the response of unseen compositions in the pair space, hence narrowing the domain gap between seen and unseen sets. Experimental results on three challenging benchmarks demonstrate that our approach significantly outperforms other state-of-the-art methods by large margins.
Keywords:	Low-level vision
Publisher:	IEEE
ISBN:	979-8-3503-0129-8 (Electronic) 979-8-3503-0130-4 (Print on Demand(PoD))
DOI:	10.1109/CVPR52729.2023.02256
Rights:	©2023 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works. The following publication X. Lu, S. Guo, Z. Liu and J. Guo, "Decomposed Soft Prompt Guided Fusion Enhancing for Compositional Zero-Shot Learning," 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 2023, pp. 23560-23569, is available at https://doi.org/10.1109/CVPR52729.2023.02256.
Appears in Collections:	Conference Paper

Files in This Item:

File	Description	Size	Format
Lu_Decomposed_Soft_Prompt.pdf	Pre-Published version	2.11 MB	Adobe PDF	View/Open

Open Access Information

Status	open access
File Version	Final Accepted Manuscript

Access

View full-text via PolyU eLinks

Show full item record

Page views

165

Citations as of Feb 9, 2026

Downloads

258

Citations as of Feb 9, 2026

SCOPUS^TM
Citations

3

Citations as of Jun 21, 2024

WEB OF SCIENCE^TM
Citations

3

Citations as of Oct 10, 2024

Google Scholar^TM

Check

Files in This Item:

Open Access Information

Access

Page views

Downloads

SCOPUSTM Citations

WEB OF SCIENCETM Citations

Google ScholarTM

Altmetric

SCOPUS^TM
Citations

WEB OF SCIENCE^TM
Citations

Google Scholar^TM