Please use this identifier to cite or link to this item: http://hdl.handle.net/10397/121761
PIRA download icon_1.1View/Download Full Text
DC FieldValueLanguage
dc.contributorDepartment of Management and Marketing-
dc.contributorIndustrial Centre-
dc.creatorYeung, Een_US
dc.creatorXiao, JJen_US
dc.creatorLee, Jen_US
dc.creatorBettoni, Men_US
dc.creatorWong, Ren_US
dc.date.accessioned2026-10-09T08:03:40Z-
dc.date.available2026-10-09T08:03:40Z-
dc.identifier.urihttp://hdl.handle.net/10397/121761-
dc.descriptionInternational Conference on GenAI and Pedagogical Innovations (GaPI), Hong Kong, 20-22 May 2026en_US
dc.language.isoenen_US
dc.rightsCopyright © 2026en_US
dc.rightsCopyright of the papers is retained by the authors. No part of this collection may be reproduced by any process without prior written permission of the copyright holders.en_US
dc.rightsPosted with permission of the publisher.en_US
dc.subjectAI gradingen_US
dc.subjectAutomated Essay Scoring (AES)en_US
dc.subjectHigher education assessmenten_US
dc.subjectOpen-source LLMen_US
dc.subjectReflective essay scoringen_US
dc.titleA multiple LLM approach to automated essay scoring : comparing seven open-source models with human gradersen_US
dc.typeConference Paperen_US
dc.identifier.spage328en_US
dc.identifier.epage334en_US
dcterms.abstractThis study evaluates whether open-source large language models (LLMs) can support automated essay scoring (AES) of reflective writing in higher education. Forty-nine reflective essays from a foundation entrepreneurship and innovation course were scored by lecturers and independently graded by seven open-source LLMs (Llama 3 8B, DeepSeek R1 32B, Ministral 3 8B, Qwen 3 8B, GPT-OSS 20B, Gemma 3 27B, and Phi 3 14B) using the same rubric. Model performance was assessed via exact match, mean absolute error (MAE), Pearson and Spearman correlations, and paired t-tests/Wilcoxon tests against human scores. Ministral produced complete outputs and demonstrated the most balanced alignment with human grading, showing no significant difference in mean or median scores. Other models exhibited larger deviations, conservative bias, or missing outputs, raising scalability concerns. Results suggest that LLMs may serve as decision-support tools but do not yet replace human judgement for reflective assessment.-
dcterms.accessRightsopen accessen_US
dcterms.bibliographicCitationIn Chen, J., Leung, A., Tsang, E., Ng, A., Chau, J., Kam, R., Patel, M., Lo, D., Tam, B., Chon, L., Cheung, K., Tang, E., & Ho, K. (Eds.). Collection of Selected Papers from the International Conference on GenAI and Pedagogical Innovations 2026, p. 328-334. Hong Kong : Educational Development Centre, Hong Kong Polytechnic University, 2026en_US
dcterms.issued2026-
dc.relation.conferenceInternational Conference on GenAI and Pedagogical Innovations [GaPI]-
dc.description.validate202610 bcch-
dc.description.oaVersion of Recorden_US
dc.identifier.FolderNumbera4809-n18-
dc.description.fundingSourceSelf-fundeden_US
dc.description.pubStatusPublisheden_US
dc.description.oaCategoryPublisher permissionen_US
Appears in Collections:Conference Paper
Files in This Item:
File Description SizeFormat 
Yeung_Multiple_LLM_Approach.pdf137.13 kBAdobe PDFView/Open
Open Access Information
Status open access
File Version Version of Record
Show simple item record

Google ScholarTM

Check


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.