Please use this identifier to cite or link to this item: http://hdl.handle.net/10397/120226
DC FieldValueLanguage
dc.contributorDepartment of English and Communicationen_US
dc.creatorWang, BXen_US
dc.creatorZhang, Cen_US
dc.creatorChan, RKWen_US
dc.date.accessioned2026-07-27T09:14:18Z-
dc.date.available2026-07-27T09:14:18Z-
dc.identifier.issn0379-0738en_US
dc.identifier.urihttp://hdl.handle.net/10397/120226-
dc.language.isoenen_US
dc.publisherElsevier BVen_US
dc.subjectChinese languagesen_US
dc.subjectForensic automatic speaker recognitionen_US
dc.subjectForensic voice comparisonen_US
dc.subjectLikelihood ratioen_US
dc.titleValidation of an embedding-based forensic voice comparison system using short speech samples in Chinese languagesen_US
dc.typeJournal/Magazine Articleen_US
dc.identifier.volume388en_US
dc.identifier.doi10.1016/j.forsciint.2026.113068en_US
dcterms.abstractThe current study evaluates the performance of an ECAPA-TDNN x-vector system in forensic voice comparison as a function of speech duration and calibration size. Forensically realistic speech from 90 Hong Kong Cantonese (HKC) and 90 Northeast Mandarin (NEM) speakers were used. Speaker embeddings were extracted from a pre-trained ECAPA-TDNN embedding model, followed by PLDA scoring and logistic regression calibration. System performance was evaluated with speech duration increasing from 2 to 20 s with a one-second increase and calibration group size from 20 to 60 speakers with a 10-speaker increase. Across both datasets, performance improved markedly as duration increased up to 10 s; gains beyond 10 s were modest. For HKC, the best result was achieved with 19 s and 40 calibration speakers (Cllr = 0.112), while the worst was at 2 s and 20 speakers (Cllr = 0.702). For NEM, the best was 20 s and 60 speakers (Cllr = 0.172), with the worst at 2 s and 20 speakers (Cllr = 0.884). The number of calibration speakers had a minor effect relative to duration; fluctuation in Cllr cal was more evident with fewer than 30 speakers, especially for short samples. HKC consistently outperformed NEM, likely due to greater content similarity across participants. Results were discussed in relation to previous studies using other languages.en_US
dcterms.accessRightsembargoed accessen_US
dcterms.bibliographicCitationForensic science international, Nov. 2026, v. 388, 113068en_US
dcterms.isPartOfForensic science internationalen_US
dcterms.issued2026-11-
dc.identifier.eissn1872-6283en_US
dc.identifier.artn113068en_US
dc.description.validate202607 bcchen_US
dc.description.oaNot applicableen_US
dc.identifier.FolderNumbera4755-
dc.identifier.SubFormID53858-
dc.description.fundingSourceRGCen_US
dc.description.fundingSourceOthersen_US
dc.description.fundingTextThis work was primarily supported by the Hong Kong Polytechnic University Start-up Fund for Research Assistant Professors under the Strategic Hiring Scheme (P0049447; 1-BDUM) and was partially supported by the Research Grants Council of the Hong Kong Special Administrative Region, China ([PolyU], [P0056858], [HKU], [17605124]), the Natural Science Foundation of Chongqing, China (CSTB2022NSCQ-LZX0007), and the Key Project of the Science and Technology Research Program of the Chongqing Municipal Education Commission (KJZD-K202600302).en_US
dc.description.pubStatusPublisheden_US
dc.date.embargo2027-11-30en_US
dc.description.oaCategoryGreen (AAM)en_US
Appears in Collections:Journal/Magazine Article
Open Access Information
Status embargoed access
Embargo End Date 2027-11-30
Access
View full-text via PolyU eLinks SFX Query
Show simple item record

Google ScholarTM

Check

Altmetric


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.