Please use this identifier to cite or link to this item: http://hdl.handle.net/10397/119802
PIRA download icon_1.1View/Download Full Text
DC FieldValueLanguage
dc.contributorDepartment of Language Science and Technologyen_US
dc.creatorMa, Jen_US
dc.creatorFeng, Zen_US
dc.creatorSong, Hen_US
dc.creatorChersoni, Een_US
dc.creatorChen, Zen_US
dc.date.accessioned2026-07-10T04:13:36Z-
dc.date.available2026-07-10T04:13:36Z-
dc.identifier.urihttp://hdl.handle.net/10397/119802-
dc.descriptionThe 3rd Workshop on Towards Knowledgeable Foundation Models (KnowFM), Vienna, Austria, 1 Aug. 2025en_US
dc.language.isoenen_US
dc.publisherAssociation for Computational Linguisticsen_US
dc.rights©2025 Association for Computational Linguisticsen_US
dc.rightsThis publication is licensed on a Creative Commons Attribution 4.0 International License. (https://creativecommons.org/licenses/by/4.0/)en_US
dc.rightsThe following publication Jianfei Ma, Zhaoxin Feng, Huacheng Song, Emmanuele Chersoni, and Zheng Chen. 2025. Reasoning or Memorization? Investigating LLMs’ Capability in Restoring Chinese Internet Homophones. In Proceedings of the 3rd Workshop on Towards Knowledgeable Foundation Models (KnowFM), pages 120–139, Vienna, Austria. Association for Computational Linguistics is available at https://aclanthology.org/2025.knowllm-1.11/.en_US
dc.titleReasoning or memorization? Investigating LLMs’ capability in restoring Chinese internet homophonesen_US
dc.typeConference Paperen_US
dc.identifier.epage120en_US
dcterms.abstractChinese homophones, prevalent in Internet culture, bring rich linguistic twists that are challenging for language models. While native speakers disambiguate them through phonological reasoning and contextual understanding, it remains untested how well LLMs perform on this task and whether LLMs also achieve this via similar reasoning processes or merely through memorization of homophone-original word pairs during training.In this paper, we present HomoP-CN, the first Chinese Internet homophones dataset with systematic perturbations for evaluating LLMs’ homophone restoration capabilities. Using this benchmark, we investigated the influence of semantic, phonological, and graphemic features on LLMs’ restoration accuracy, measured the reliance levels of each model on memorization during restoration through consistency ratios under controlled perturbations, and assessed the effectiveness of various prompting strategies, including contextual cues, pinyin augmentation, few-shot learning, and thought-chain approaches.en_US
dcterms.accessRightsopen accessen_US
dcterms.bibliographicCitationIn Proceedings of the 3rd Workshop on Towards Knowledgeable Foundation Models (KnowFM), p. 120-139. Vienna, Austria: Association for Computational Linguistics, 2025en_US
dcterms.issued2025-
dc.relation.ispartofbookProceedings of the 3rd Workshop on Towards Knowledgeable Foundation Models (KnowFM)en_US
dc.relation.conferenceWorkshop on Towards Knowledgeable Foundation Models [KnowFM]en_US
dc.identifier.artn139en_US
dc.description.validate202607 bcwhen_US
dc.description.oaVersion of Recorden_US
dc.identifier.FolderNumbera4654-
dc.identifier.SubFormID53459-
dc.description.fundingSourceSelf-fundeden_US
dc.description.pubStatusPublisheden_US
dc.description.oaCategoryCCen_US
Appears in Collections:Conference Paper
Files in This Item:
File Description SizeFormat 
2025.knowllm-1.11v2.pdf1.65 MBAdobe PDFView/Open
Open Access Information
Status open access
File Version Version of Record
Access
View full-text via PolyU eLinks SFX Query
Show simple item record

Google ScholarTM

Check


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.