Please use this identifier to cite or link to this item: http://hdl.handle.net/10397/121714
DC FieldValueLanguage
dc.contributorDepartment of Data Science and Artificial Intelligenceen_US
dc.creatorJiang, Ken_US
dc.creatorHuang, Jen_US
dc.creatorZhang, Jen_US
dc.creatorXie, Wen_US
dc.creatorLi, Yen_US
dc.creatorWang, Yen_US
dc.creatorXiao, Aen_US
dc.creatorTao, Den_US
dc.date.accessioned2026-10-05T03:27:22Z-
dc.date.available2026-10-05T03:27:22Z-
dc.identifier.issn0302-9743en_US
dc.identifier.urihttp://hdl.handle.net/10397/121714-
dc.descriptionComputer Vision - ECCV 2026: 19th European Conference, Malmö, Sweden, September 8-12, 2026en_US
dc.language.isoenen_US
dc.publisherSpringeren_US
dc.subjectEmbodied AIen_US
dc.subjectKnowledge distillationen_US
dc.subjectLightweight modelen_US
dc.subjectSegment anything modelen_US
dc.subjectSpatial intelligenceen_US
dc.titleMobileSAM2 : lightweight segment anything for spatial intelligenceen_US
dc.typeConference Paperen_US
dc.identifier.spage494en_US
dc.identifier.epage514en_US
dc.identifier.doi10.1007/978-3-032-37032-7_27en_US
dcterms.abstractThe recent large video foundation model, SAM2, enables segment anything in both images and videos, serving as a powerful base model for various applications. However, many of such use cases require to operate on resource-constrained devices like mobile phones and laptops. In this work, we aim to make SAM2 more mobile-friendly by distilling the heavyweight SAM2 into a lightweight model, facilitating segment anything in both images and videos on mobile devices. To this end, we propose Hypergraphical Knowledge Distill (HyperKD), which introduces the idea of hypergraph into knowledge distillation, aiming to effectively model and transfer SAM2’s generalizable and comprehensive knowledge. HyperKD consists of Temporal HyperKD and Granularity HyperKD that construct hypergraphs to explicitly model and extract the generalizable temporal knowledge and the comprehensive multi-granularity knowledge from SAM2 respectively, which are then distilled into the lightweight student model by aligning it with the constructed hypergraphs. Besides, we present MobileSAM2, a new family of lightweight SAM2 that balances efficiency and effectiveness via searching the best model architectures with HyperKD during model size reduction. Extensive experiments validate MobileSAM2 across multiple benchmarks and show promising generalization performance on embodied AI tasks.en_US
dcterms.accessRightsembargoed accessen_US
dcterms.bibliographicCitationLecture notes in computer science (including subseries Lecture notes in artificial intelligence and lecture notes in bioinformatics), 2026, v. 17075, p. 494-514en_US
dcterms.isPartOfLecture notes in computer science (including subseries Lecture notes in artificial intelligence and lecture notes in bioinformatics)en_US
dcterms.issued2026-
dc.relation.conferenceEuropean Conference on Computer Vision [ECCV]en_US
dc.identifier.eissn1611-3349en_US
dc.identifier.artn17075en_US
dc.description.validate202609 bcchen_US
dc.description.oaNot applicableen_US
dc.identifier.FolderNumbera4711-
dc.identifier.SubFormID53733-
dc.description.fundingSourceOthersen_US
dc.description.fundingTextThis project is supported by the National Research Foundation, Singapore, under its NRF Professorship Award No. NRF-P2024-001. This work is also supported by PolyU Internal Fund.en_US
dc.description.pubStatusPublisheden_US
dc.date.embargo2027-09-18en_US
dc.description.oaCategoryGreen (AAM)en_US
Appears in Collections:Conference Paper
Open Access Information
Status embargoed access
Embargo End Date 2027-09-18
Access
View full-text via PolyU eLinks SFX Query
Show simple item record

Google ScholarTM

Check

Altmetric


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.