Please use this identifier to cite or link to this item: http://hdl.handle.net/10397/120073
PIRA download icon_1.1View/Download Full Text
DC FieldValueLanguage
dc.contributorDepartment of Computingen_US
dc.contributorDepartment of Land Surveying and Geospatial Scienceen_US
dc.creatorSun, Yen_US
dc.creatorZhang, Ren_US
dc.creatorSun, Aen_US
dc.creatorLi, Xen_US
dc.creatorLiu, Zen_US
dc.creatorGuo, Jen_US
dc.date.accessioned2026-07-22T03:54:19Z-
dc.date.available2026-07-22T03:54:19Z-
dc.identifier.urihttp://hdl.handle.net/10397/120073-
dc.language.isoenen_US
dc.publisherOpenReview.neten_US
dc.rightsCC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)en_US
dc.rightsThe following publication Sun, Y., Zhang, R., Sun, A., Li, X., Liu, Z., & Guo, J. (2026). D&R : recovery-based AI-generated text detection via a single black-box LLM call. In The Fourteenth International Conference on Learning Representations is available at https://openreview.net/forum?id=FiMZSxo4DO.en_US
dc.subjectAI-generated text detectionen_US
dc.subjectBlack-box detectionen_US
dc.subjectLarge language modelsen_US
dc.subjectRecovery-based detectionen_US
dc.subjectRobustnessen_US
dc.subjectTraining-free methodsen_US
dc.titleD&R : recovery-based AI-generated text detection via a single black-box LLM callen_US
dc.typeConference Paperen_US
dcterms.abstractLarge language models (LLMs) generate increasingly human-like text, raising concerns about misinformation and authenticity. Detecting AI-generated text remains challenging: existing methods often underperform, especially on short texts, require probability access unavailable in real-world black-box settings, incur high costs from multiple calls, or fail to generalize across models. We propose Disrupt-and-Recover (D&R), a recovery-based detection framework grounded in posterior concentration. D&R disrupts text via model-free Within-Chunk Shuffling, performs a single black-box LLM recovery, and measures semantic–structural recovery similarity as a proxy for concentration. This design ensures efficiency, black-box practicality, and is theoretically supported under the concentration assumption. Extensive experiments across four datasets and six source models show that D&R achieves state-of-the-art performance, with AUROC 0.96 on long texts and 0.87 on short texts, surpassing the strongest baseline by +0.08 and +0.14. D&R further remains robust under source–recovery mismatch and model variation. Our code and data are available at https://github.com/Yuxia-Sun/D-R.en_US
dcterms.accessRightsopen accessen_US
dcterms.bibliographicCitationThe Fourteenth International Conference on Learning Representations, ICLR 2026, Rio de Janeiro, Brazil, Apr 23-27 2026, https://openreview.net/forum?id=FiMZSxo4DOen_US
dcterms.issued2026-
dc.description.validate202607 bcwcen_US
dc.description.oaVersion of Recorden_US
dc.identifier.FolderNumbera4666-
dc.identifier.SubFormID53537-
dc.description.fundingSourceRGCen_US
dc.description.pubStatusPublisheden_US
dc.description.oaCategoryCCen_US
Appears in Collections:Conference Paper
Files in This Item:
File Description SizeFormat 
13074_D_R_Recovery_based_AI_Ge.pdf533.98 kBAdobe PDFView/Open
Open Access Information
Status open access
File Version Version of Record
Show simple item record

Google ScholarTM

Check


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.