Please use this identifier to cite or link to this item:
http://hdl.handle.net/10397/119802
| DC Field | Value | Language |
|---|---|---|
| dc.contributor | Department of Language Science and Technology | en_US |
| dc.creator | Ma, J | en_US |
| dc.creator | Feng, Z | en_US |
| dc.creator | Song, H | en_US |
| dc.creator | Chersoni, E | en_US |
| dc.creator | Chen, Z | en_US |
| dc.date.accessioned | 2026-07-10T04:13:36Z | - |
| dc.date.available | 2026-07-10T04:13:36Z | - |
| dc.identifier.uri | http://hdl.handle.net/10397/119802 | - |
| dc.description | The 3rd Workshop on Towards Knowledgeable Foundation Models (KnowFM), Vienna, Austria, 1 Aug. 2025 | en_US |
| dc.language.iso | en | en_US |
| dc.publisher | Association for Computational Linguistics | en_US |
| dc.rights | ©2025 Association for Computational Linguistics | en_US |
| dc.rights | This publication is licensed on a Creative Commons Attribution 4.0 International License. (https://creativecommons.org/licenses/by/4.0/) | en_US |
| dc.rights | The following publication Jianfei Ma, Zhaoxin Feng, Huacheng Song, Emmanuele Chersoni, and Zheng Chen. 2025. Reasoning or Memorization? Investigating LLMs’ Capability in Restoring Chinese Internet Homophones. In Proceedings of the 3rd Workshop on Towards Knowledgeable Foundation Models (KnowFM), pages 120–139, Vienna, Austria. Association for Computational Linguistics is available at https://aclanthology.org/2025.knowllm-1.11/. | en_US |
| dc.title | Reasoning or memorization? Investigating LLMs’ capability in restoring Chinese internet homophones | en_US |
| dc.type | Conference Paper | en_US |
| dc.identifier.epage | 120 | en_US |
| dcterms.abstract | Chinese homophones, prevalent in Internet culture, bring rich linguistic twists that are challenging for language models. While native speakers disambiguate them through phonological reasoning and contextual understanding, it remains untested how well LLMs perform on this task and whether LLMs also achieve this via similar reasoning processes or merely through memorization of homophone-original word pairs during training.In this paper, we present HomoP-CN, the first Chinese Internet homophones dataset with systematic perturbations for evaluating LLMs’ homophone restoration capabilities. Using this benchmark, we investigated the influence of semantic, phonological, and graphemic features on LLMs’ restoration accuracy, measured the reliance levels of each model on memorization during restoration through consistency ratios under controlled perturbations, and assessed the effectiveness of various prompting strategies, including contextual cues, pinyin augmentation, few-shot learning, and thought-chain approaches. | en_US |
| dcterms.accessRights | open access | en_US |
| dcterms.bibliographicCitation | In Proceedings of the 3rd Workshop on Towards Knowledgeable Foundation Models (KnowFM), p. 120-139. Vienna, Austria: Association for Computational Linguistics, 2025 | en_US |
| dcterms.issued | 2025 | - |
| dc.relation.ispartofbook | Proceedings of the 3rd Workshop on Towards Knowledgeable Foundation Models (KnowFM) | en_US |
| dc.relation.conference | Workshop on Towards Knowledgeable Foundation Models [KnowFM] | en_US |
| dc.identifier.artn | 139 | en_US |
| dc.description.validate | 202607 bcwh | en_US |
| dc.description.oa | Version of Record | en_US |
| dc.identifier.FolderNumber | a4654 | - |
| dc.identifier.SubFormID | 53459 | - |
| dc.description.fundingSource | Self-funded | en_US |
| dc.description.pubStatus | Published | en_US |
| dc.description.oaCategory | CC | en_US |
| Appears in Collections: | Conference Paper | |
Files in This Item:
| File | Description | Size | Format | |
|---|---|---|---|---|
| 2025.knowllm-1.11v2.pdf | 1.65 MB | Adobe PDF | View/Open |
Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.



