Please use this identifier to cite or link to this item: http://hdl.handle.net/10397/119984
PIRA download icon_1.1View/Download Full Text
DC FieldValueLanguage
dc.contributorDepartment of Language Science and Technology-
dc.creatorQiu, L-
dc.creatorChersoni, E-
dc.creatorVillavicencio, A-
dc.date.accessioned2026-07-17T09:20:05Z-
dc.date.available2026-07-17T09:20:05Z-
dc.identifier.isbn979-8-89176-340-1-
dc.identifier.urihttp://hdl.handle.net/10397/119984-
dc.descriptionThe 14th Joint Conference on Lexical and Computational Semantics (*SEM 2025), Suzhou, China, November 8-9, 2025en_US
dc.language.isoenen_US
dc.rights©2025 Association for Computational Linguisticsen_US
dc.rightsLicensed under the Creative Commons Attribution 4.0 International License (https://creativecommons.org/licenses/by/4.0/)en_US
dc.rightsThe following publication Le Qiu, Emmanuele Chersoni, and Aline Villavicencio. 2025. ChengyuSTS: An Intrinsic Perspective on Mandarin Idiom Representation. In Proceedings of the 14th Joint Conference on Lexical and Computational Semantics (*SEM 2025), pages 1–12, Suzhou, China. Association for Computational Linguistics is available at https://aclanthology.org/2025.starsem-1.1/.en_US
dc.titleChengyuSTS : an intrinsic perspective on Mandarin idiom representationen_US
dc.typeConference Paperen_US
dc.identifier.spage1-
dc.identifier.epage12-
dcterms.abstractChengyu, or four-character idioms, are ubiquitous in both spoken and written Chinese. Despite their importance, chengyu are often underexplored in NLP tasks, and existing evaluation frameworks are primarily based on extrinsic measures. In this paper, we introduce an intrinsic evaluation task for Chinese idiomatic understanding: idiomatic semantic textual similarity (iSTS), which evaluates how well models can capture the semantic similarity of sentences containing idioms. To this purpose, we present a curated dataset: ChengyuSTS. Our experiments show that current pre-trained sentence Transformer models generally fail to capture the idiomaticity of chengyu in a zero-shot setting. We then show results of fine-tuned models using the SimCSE contrastive learning framework, which demonstrate promising results for handling idiomatic expressions. We also presented the results of DeepSeek for reference-
dcterms.accessRightsopen accessen_US
dcterms.bibliographicCitationIn Proceedings of the 14th Joint Conference on Lexical and Computational Semantics (*SEM 2025), p.1-12. Kerrville : Association for Computational Linguistics, 2025-
dcterms.issued2025-
dc.relation.ispartofbookProceedings of the 14th Joint Conference on Lexical and Computational Semantics (*SEM 2025)-
dc.description.validate202607 bcwh-
dc.description.oaVersion of Recorden_US
dc.identifier.FolderNumbera4672en_US
dc.identifier.SubFormID53556en_US
dc.description.fundingSourceRGCen_US
dc.description.pubStatusPublisheden_US
dc.description.oaCategoryCCen_US
Appears in Collections:Conference Paper
Files in This Item:
File Description SizeFormat 
2025.starsem-1.1.pdf531.73 kBAdobe PDFView/Open
Open Access Information
Status open access
File Version Version of Record
Access
View full-text via PolyU eLinks SFX Query
Show simple item record

Google ScholarTM

Check


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.