Please use this identifier to cite or link to this item: http://hdl.handle.net/10397/120074
DC FieldValueLanguage
dc.contributorDepartment of Computingen_US
dc.contributorDepartment of Land Surveying and Geospatial Scienceen_US
dc.creatorChen, Yen_US
dc.creatorLi, Men_US
dc.creatorRao, Zen_US
dc.creatorZeng, Den_US
dc.creatorGuo, Sen_US
dc.creatorGuo, Jen_US
dc.date.accessioned2026-07-22T03:54:20Z-
dc.date.available2026-07-22T03:54:20Z-
dc.identifier.urihttp://hdl.handle.net/10397/120074-
dc.language.isoenen_US
dc.titleLearning by neighbor-aware semantics, deciding by open-form flows : towards robust zero-shot skeleton action recognitionen_US
dc.typeConference Paperen_US
dc.identifier.spage3374en_US
dc.identifier.epage3383en_US
dcterms.abstractRecognizing unseen skeleton action categories remains highly challenging due to the absence of corresponding skeletal priors. Existing approaches generally follow an “align-then-classify” paradigm but face two fundamental issues: (i) fragile point-to-point alignment arising from imperfect semantics, and (ii) rigid classifiers restricted by static decision boundaries and coarse-grained anchors. To address these issues, we propose a novel method for zero-shot skeleton action recognition, termed Flora, which builds upon FlexibLe neighbOr-aware semantic attunement and open-form distRibution-aware flow clAssifier. Specifically, we flexibly attune textual semantics by incorporating neighboring inter-class contextual cues to form direction-aware regional semantics, coupled with a cross-modal geometric consistency objective that ensures stable and robust point-to-region alignment. Furthermore, we employ noise-free flow matching to bridge the modality distribution gap between semantic and skeleton latent embeddings, while a condition-free contrastive regularization enhances discriminability, leading to a distribution-aware classifier with fine-grained decision boundaries achieved through token-level velocity predictions. Extensive experiments on three benchmark datasets validate the effectiveness of our method, showing particularly impressive performance even when trained with only 10% of the seen data. Code is available at https://github.com/cseeyangchen/Flora.en_US
dcterms.accessRightsembargoed accessen_US
dcterms.bibliographicCitationThe IEEE/CVF Conference on Computer Vision and Pattern Recognition 2026, June 3 - June 7, 2026, Colorado Convention Center, p. 3374-3383en_US
dcterms.issued2026-
dc.description.validate202607 bcwcen_US
dc.description.oaNot applicableen_US
dc.identifier.FolderNumbera4666-
dc.identifier.SubFormID53538-
dc.description.fundingSourceRGCen_US
dc.description.fundingSourceOthersen_US
dc.description.fundingTextThis research was supported by the Hong Kong RGC General Research Fund (Grant Nos. 15221123, 15216424, and 15211525) and the Hong Kong PolyU Internal Research Fund (Grant Nos. P0058468 and P0056171).en_US
dc.description.pubStatusEarly releaseen_US
dc.date.embargo0000-00-00 (to be updated)en_US
dc.description.oaCategoryGreen (AAM)en_US
Appears in Collections:Conference Paper
Open Access Information
Status embargoed access
Embargo End Date 0000-00-00 (to be updated)
Show simple item record

Google ScholarTM

Check


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.