Please use this identifier to cite or link to this item:
http://hdl.handle.net/10397/121761
| DC Field | Value | Language |
|---|---|---|
| dc.contributor | Department of Management and Marketing | - |
| dc.contributor | Industrial Centre | - |
| dc.creator | Yeung, E | en_US |
| dc.creator | Xiao, JJ | en_US |
| dc.creator | Lee, J | en_US |
| dc.creator | Bettoni, M | en_US |
| dc.creator | Wong, R | en_US |
| dc.date.accessioned | 2026-10-09T08:03:40Z | - |
| dc.date.available | 2026-10-09T08:03:40Z | - |
| dc.identifier.uri | http://hdl.handle.net/10397/121761 | - |
| dc.description | International Conference on GenAI and Pedagogical Innovations (GaPI), Hong Kong, 20-22 May 2026 | en_US |
| dc.language.iso | en | en_US |
| dc.rights | Copyright © 2026 | en_US |
| dc.rights | Copyright of the papers is retained by the authors. No part of this collection may be reproduced by any process without prior written permission of the copyright holders. | en_US |
| dc.rights | Posted with permission of the publisher. | en_US |
| dc.subject | AI grading | en_US |
| dc.subject | Automated Essay Scoring (AES) | en_US |
| dc.subject | Higher education assessment | en_US |
| dc.subject | Open-source LLM | en_US |
| dc.subject | Reflective essay scoring | en_US |
| dc.title | A multiple LLM approach to automated essay scoring : comparing seven open-source models with human graders | en_US |
| dc.type | Conference Paper | en_US |
| dc.identifier.spage | 328 | en_US |
| dc.identifier.epage | 334 | en_US |
| dcterms.abstract | This study evaluates whether open-source large language models (LLMs) can support automated essay scoring (AES) of reflective writing in higher education. Forty-nine reflective essays from a foundation entrepreneurship and innovation course were scored by lecturers and independently graded by seven open-source LLMs (Llama 3 8B, DeepSeek R1 32B, Ministral 3 8B, Qwen 3 8B, GPT-OSS 20B, Gemma 3 27B, and Phi 3 14B) using the same rubric. Model performance was assessed via exact match, mean absolute error (MAE), Pearson and Spearman correlations, and paired t-tests/Wilcoxon tests against human scores. Ministral produced complete outputs and demonstrated the most balanced alignment with human grading, showing no significant difference in mean or median scores. Other models exhibited larger deviations, conservative bias, or missing outputs, raising scalability concerns. Results suggest that LLMs may serve as decision-support tools but do not yet replace human judgement for reflective assessment. | - |
| dcterms.accessRights | open access | en_US |
| dcterms.bibliographicCitation | In Chen, J., Leung, A., Tsang, E., Ng, A., Chau, J., Kam, R., Patel, M., Lo, D., Tam, B., Chon, L., Cheung, K., Tang, E., & Ho, K. (Eds.). Collection of Selected Papers from the International Conference on GenAI and Pedagogical Innovations 2026, p. 328-334. Hong Kong : Educational Development Centre, Hong Kong Polytechnic University, 2026 | en_US |
| dcterms.issued | 2026 | - |
| dc.relation.conference | International Conference on GenAI and Pedagogical Innovations [GaPI] | - |
| dc.description.validate | 202610 bcch | - |
| dc.description.oa | Version of Record | en_US |
| dc.identifier.FolderNumber | a4809-n18 | - |
| dc.description.fundingSource | Self-funded | en_US |
| dc.description.pubStatus | Published | en_US |
| dc.description.oaCategory | Publisher permission | en_US |
| Appears in Collections: | Conference Paper | |
Files in This Item:
| File | Description | Size | Format | |
|---|---|---|---|---|
| Yeung_Multiple_LLM_Approach.pdf | 137.13 kB | Adobe PDF | View/Open |
Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.


