Please use this identifier to cite or link to this item:
http://hdl.handle.net/10397/120226
| DC Field | Value | Language |
|---|---|---|
| dc.contributor | Department of English and Communication | en_US |
| dc.creator | Wang, BX | en_US |
| dc.creator | Zhang, C | en_US |
| dc.creator | Chan, RKW | en_US |
| dc.date.accessioned | 2026-07-27T09:14:18Z | - |
| dc.date.available | 2026-07-27T09:14:18Z | - |
| dc.identifier.issn | 0379-0738 | en_US |
| dc.identifier.uri | http://hdl.handle.net/10397/120226 | - |
| dc.language.iso | en | en_US |
| dc.publisher | Elsevier BV | en_US |
| dc.subject | Chinese languages | en_US |
| dc.subject | Forensic automatic speaker recognition | en_US |
| dc.subject | Forensic voice comparison | en_US |
| dc.subject | Likelihood ratio | en_US |
| dc.title | Validation of an embedding-based forensic voice comparison system using short speech samples in Chinese languages | en_US |
| dc.type | Journal/Magazine Article | en_US |
| dc.identifier.volume | 388 | en_US |
| dc.identifier.doi | 10.1016/j.forsciint.2026.113068 | en_US |
| dcterms.abstract | The current study evaluates the performance of an ECAPA-TDNN x-vector system in forensic voice comparison as a function of speech duration and calibration size. Forensically realistic speech from 90 Hong Kong Cantonese (HKC) and 90 Northeast Mandarin (NEM) speakers were used. Speaker embeddings were extracted from a pre-trained ECAPA-TDNN embedding model, followed by PLDA scoring and logistic regression calibration. System performance was evaluated with speech duration increasing from 2 to 20 s with a one-second increase and calibration group size from 20 to 60 speakers with a 10-speaker increase. Across both datasets, performance improved markedly as duration increased up to 10 s; gains beyond 10 s were modest. For HKC, the best result was achieved with 19 s and 40 calibration speakers (Cllr = 0.112), while the worst was at 2 s and 20 speakers (Cllr = 0.702). For NEM, the best was 20 s and 60 speakers (Cllr = 0.172), with the worst at 2 s and 20 speakers (Cllr = 0.884). The number of calibration speakers had a minor effect relative to duration; fluctuation in Cllr cal was more evident with fewer than 30 speakers, especially for short samples. HKC consistently outperformed NEM, likely due to greater content similarity across participants. Results were discussed in relation to previous studies using other languages. | en_US |
| dcterms.accessRights | embargoed access | en_US |
| dcterms.bibliographicCitation | Forensic science international, Nov. 2026, v. 388, 113068 | en_US |
| dcterms.isPartOf | Forensic science international | en_US |
| dcterms.issued | 2026-11 | - |
| dc.identifier.eissn | 1872-6283 | en_US |
| dc.identifier.artn | 113068 | en_US |
| dc.description.validate | 202607 bcch | en_US |
| dc.description.oa | Not applicable | en_US |
| dc.identifier.FolderNumber | a4755 | - |
| dc.identifier.SubFormID | 53858 | - |
| dc.description.fundingSource | RGC | en_US |
| dc.description.fundingSource | Others | en_US |
| dc.description.fundingText | This work was primarily supported by the Hong Kong Polytechnic University Start-up Fund for Research Assistant Professors under the Strategic Hiring Scheme (P0049447; 1-BDUM) and was partially supported by the Research Grants Council of the Hong Kong Special Administrative Region, China ([PolyU], [P0056858], [HKU], [17605124]), the Natural Science Foundation of Chongqing, China (CSTB2022NSCQ-LZX0007), and the Key Project of the Science and Technology Research Program of the Chongqing Municipal Education Commission (KJZD-K202600302). | en_US |
| dc.description.pubStatus | Published | en_US |
| dc.date.embargo | 2027-11-30 | en_US |
| dc.description.oaCategory | Green (AAM) | en_US |
| Appears in Collections: | Journal/Magazine Article | |
Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.



