Please use this identifier to cite or link to this item: http://hdl.handle.net/10397/119982
PIRA download icon_1.1View/Download Full Text
DC FieldValueLanguage
dc.contributorDepartment of Language Science and Technology-
dc.creatorMiliani, M-
dc.creatorAuriemma, S-
dc.creatorBondielli, A-
dc.creatorChersoni, E-
dc.creatorPassaro, LC-
dc.creatorSucameli, I-
dc.creatorLenci, A-
dc.date.accessioned2026-07-17T09:20:04Z-
dc.date.available2026-07-17T09:20:04Z-
dc.identifier.isbn979-8-89176-256-5-
dc.identifier.urihttp://hdl.handle.net/10397/119982-
dc.descriptionThe 63rd Annual Meeting of the Association for Computational Linguistics, Vienna, Austria, July 27- August 1, 2025en_US
dc.language.isoenen_US
dc.rights©2025 Association for Computational Linguisticsen_US
dc.rightsLicensed under the Creative Commons Attribution 4.0 International License (https://creativecommons.org/licenses/by/4.0/)en_US
dc.rightsThe following publication Martina Miliani, Serena Auriemma, Alessandro Bondielli, Emmanuele Chersoni, Lucia C. Passaro, Irene Sucameli, and Alessandro Lenci. 2025. ExpliCa: Evaluating Explicit Causal Reasoning in Large Language Models. In Findings of the Association for Computational Linguistics: ACL 2025, pages 17335–17355, Vienna, Austria. Association for Computational Linguistics is available at https://aclanthology.org/2025.findings-acl.891/.en_US
dc.titleExpliCa : evaluating explicit causal reasoning in large language modelsen_US
dc.typeConference Paperen_US
dc.identifier.spage17335-
dc.identifier.epage17355-
dc.identifier.doi10.18653/v1/2025.findings-acl.891-
dcterms.abstractLarge Language Models (LLMs) are increasingly used in tasks requiring interpretive and inferential accuracy. In this paper, we introduce ExpliCa, a new dataset for evaluating LLMs in explicit causal reasoning. ExpliCa uniquely integrates both causal and temporal relations presented in different linguistic orders and explicitly expressed by linguistic connectives. The dataset is enriched with crowdsourced human acceptability ratings. We tested LLMs on ExpliCa through prompting and perplexity-based metrics. We assessed seven commercial and open-source LLMs, revealing that even top models struggle to reach 0.80 accuracy. Interestingly, models tend to confound temporal relations with causal ones, and their performance is also strongly influenced by the linguistic order of the events. Finally, perplexity-based scores and prompting performance are differently affected by model size.-
dcterms.accessRightsopen accessen_US
dcterms.bibliographicCitationIn Findings of the Association for Computational Linguistics: ACL 2025, p. 17335-17355. Kerrville : Association for Computational Linguistics, 2025-
dcterms.issued2025-
dc.relation.ispartofbookFindings of the Association for Computational Linguistics: ACL 2025-
dc.relation.conferenceAnnual Meeting of the Association for Computational Linguistics-
dc.description.validate202607 bcwh-
dc.description.oaVersion of Recorden_US
dc.identifier.FolderNumbera4672en_US
dc.identifier.SubFormID53554en_US
dc.description.fundingSourceRGCen_US
dc.description.pubStatusPublisheden_US
dc.description.oaCategoryCCen_US
Appears in Collections:Conference Paper
Files in This Item:
File Description SizeFormat 
2025.acl-long.1132.pdf846.37 kBAdobe PDFView/Open
Open Access Information
Status open access
File Version Version of Record
Access
View full-text via PolyU eLinks SFX Query
Show simple item record

Google ScholarTM

Check

Altmetric


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.