Please use this identifier to cite or link to this item: http://hdl.handle.net/10397/119981
PIRA download icon_1.1View/Download Full Text
DC FieldValueLanguage
dc.contributorDepartment of Language Science and Technology-
dc.creatorFeng, Z-
dc.creatorMa, J-
dc.creatorChersoni, E-
dc.creatorZhao, X-
dc.creatorBao, X-
dc.date.accessioned2026-07-17T09:20:03Z-
dc.date.available2026-07-17T09:20:03Z-
dc.identifier.isbn979-8-89176-251-0-
dc.identifier.urihttp://hdl.handle.net/10397/119981-
dc.descriptionThe 63rd Annual Meeting of the Association for Computational Linguistics, Vienna, Austria, July 27- August 1, 2025en_US
dc.language.isoenen_US
dc.publisherAssociation for Computational Linguisticsen_US
dc.rights©2025 Association for Computational Linguisticsen_US
dc.rightsLicensed under the Creative Commons Attribution 4.0 International License (https://creativecommons.org/licenses/by/4.0/)en_US
dc.rightsThe following publication Zhaoxin Feng, Jianfei Ma, Emmanuele Chersoni, Xiaojing Zhao, and Xiaoyi Bao. 2025. Learning to Look at the Other Side: A Semantic Probing Study of Word Embeddings in LLMs with Enabled Bidirectional Attention. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 23226–23245, Vienna, Austria. Association for Computational Linguistics is available at https://aclanthology.org/2025.acl-long.1132/.en_US
dc.titleLearning to look at the other side : a semantic probing study of word embeddings in LLMs with enabled bidirectional attentionen_US
dc.typeConference Paperen_US
dc.identifier.spage23226-
dc.identifier.epage23245-
dc.identifier.volume1-
dc.identifier.doi10.18653/v1/2025.acl-long.1132-
dcterms.abstractAutoregressive Large Language Models (LLMs) demonstrate exceptional performance in language understanding and generation. However, their application in text embedding tasks has been relatively slow, along with the analysis of their semantic representation in probing tasks, due to the constraints of the unidirectional attention mechanism. This paper aims to explore whether such constraints can be overcome by enabling bidirectional attention in LLMs. We tested different variants of the Llama architecture through additional training steps, progressively enabling bidirectional attention and unsupervised/supervised contrastive learning. Our results show that bidirectional attention improves the LLMs’ ability to represent subsequent context but weakens their utilization of preceding context, while contrastive learning training can help to maintain both abilities.-
dcterms.accessRightsopen accessen_US
dcterms.bibliographicCitationIn Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (v. 1: Long Papers), p. 23226-23245. Kerrville : Association for Computational Linguistics, 2025-
dcterms.issued2025-
dc.relation.ispartofbookProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)-
dc.relation.conferenceAnnual Meeting of the Association for Computational Linguistics-
dc.identifier.artn23226-
dc.description.validate202607 bcwh-
dc.description.oaVersion of Recorden_US
dc.identifier.FolderNumbera4672en_US
dc.identifier.SubFormID53552en_US
dc.description.fundingSourceSelf-fundeden_US
dc.description.pubStatusPublisheden_US
dc.description.oaCategoryCCen_US
Appears in Collections:Conference Paper
Files in This Item:
File Description SizeFormat 
2025.acl-long.1132.pdf846.37 kBAdobe PDFView/Open
Open Access Information
Status open access
File Version Version of Record
Access
View full-text via PolyU eLinks SFX Query
Show simple item record

Google ScholarTM

Check

Altmetric


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.