Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews

LLM HDCE AI Safety / AI Ethics DWAT
2024年03月11日
我们提出了一种方法,用于估计大语言模型(LLM)可能会大量修改或生成的大型语料库中文本的比例。我们的最大似然模型利用专家编写和AI生成的参考文本,在语料库级别上准确高效地检查现实世界中LLM使用情况。我们将此方法应用于一个科学同行评审的案例研究,该研究发生在ChatGPT发布之后的人工智能会议上:ICLR 2024、NeurIPS 2023、CoRL 2023和EMNLP 2023。我们的结果表明,提交给这些会议的同行评审文本中,有6.5%至16.9%的文本可能会被LLM大量修改,即超出了拼写检查或小的写作更新。生成文本出现的情况可以提供关于用户行为的见解:在报告较低置信度、接近截止日期提交的评论以及不太可能回应作者反驳的审稿人的评论中,估计的LLM生成文本比例更高。我们还观察到生成文本的语料库级别趋势,这可能在个体级别上太微妙而无法检测,并讨论了这些趋势对同行评审的影响。我们呼吁未来跨学科研究来研究LLM使用如何改变我们的信息和知识实践。
We present an approach for estimating the fraction of text in a large corpus which is likely to be substantially modified or produced by a large language model (LLM). Our maximum likelihood model leverages expert-written and AI-generated reference texts to accurately and efficiently examine real-world LLM-use at the corpus level. We apply this approach to a case study of scientific peer review in AI conferences that took place after the release of ChatGPT: ICLR 2024, NeurIPS 2023, CoRL 2023 and EMNLP 2023. Our results suggest that between 6.5% and 16.9% of text submitted as peer reviews to these conferences could have been substantially modified by LLMs, i.e. beyond spell-checking or minor writing updates. The circumstances in which generated text occurs offer insight into user behavior: the estimated fraction of LLM-generated text is higher in reviews which report lower confidence, were submitted close to the deadline, and from reviewers who are less likely to respond to author rebuttals. We also observe corpus-level trends in generated text which may be too subtle to detect at the individual level, and discuss the implications of such trends on peer review. We call for future interdisciplinary work to examine how LLM use is changing our information and knowledge practices.
许愿