SlopShape: Identifying AI-Generated Commercial Web Content

LLM Other LLM NLP Other NLP AI Safety / AI Ethics DWAT
词级检测器几乎能完美识别未经修改的AI生成文本,但现有文献已指出,这类检测器在文本被改写后极易失效;此外,词级得分既无法刻画文本的整体特征,也无法确定其出自哪一种AI模型。因此,我们提出一个问题:能否在更深层次上识别AI生成文本——即从其结构特征入手,考察信息如何呈现、以何种顺序展开、依托哪些证据支撑,以及采用何种叙述口吻?我们复现了StoryScope(Russell等,2026)的研究方法——该研究曾揭示AI生成小说所具有的特定结构模式,并将这一方法拓展至商业内容领域:以268个企业网站发布的2,250篇ChatGPT问世前的人类撰写的博客文章为基准,与五种前沿AI模型生成的11,250篇对应“镜像”文本进行对比分析。我们构建了一套包含214个特征的结构化测量工具,由大语言模型(LLM)执行评估,并通过人工黄金标注环节完成验证(人工标注者之间的一致性kappa系数达0.928,人工标注与模型标注之间的一致性kappa系数达0.946)。仅凭其中187个结构特征,该工具即可在未参与训练的公司样本上实现98.0的宏平均F1值;即使对每一篇AI生成的博文均由其原始生成模型自行重写(即彻底改写),检测性能也几乎不受影响(宏平均F1值仍达98.1)。这一结构信号不仅具有表征能力,还具备归因能力:AI生成的博文普遍呈现出一种规整、自我指涉的结构形态;在源模型归因任务中,79.3%的AI博文被准确识别出其真实生成模型,远高于随机猜测的16.7%基线水平;而人类撰写的博文则多分布于极为罕见的结构配置之中。上述所有效应均成功复现了StoryScope的研究发现,且效应方向完全一致、强度更为显著。我们已开源全部流程、测量工具、提示词(prompts)、代码及聚合后的分析成果。
Word-level detectors identify unedited AI-generated text almost perfectly, but the literature documents their brittleness under rewording, and a word-level score neither characterizes a text nor identifies which AI model wrote it. We ask whether AI-generated text can be identified one level deeper, from structural signatures: how information is presented, in what order, with what evidence, and in what voice. We replicate StoryScope (Russell et al., 2026), which showed such patterns for AI-generated fiction, on commercial content: 2,250 pre-ChatGPT human blog posts from 268 company domains against 11,250 AI mirrors from five frontier models. A 214-feature instrument, applied by an LLM and validated in a human gold-annotation session (human-human kappa 0.928, human-model 0.946), detects AI posts from its 187 structural features alone at 98.0 macro-F1 on held-out companies, unchanged (98.1) when every AI post is reworded by its own model. The signal characterizes and attributes: AI posts share a tidy, self-announcing shape, 79.3% are attributed to the correct source against a 16.7% chance rate, and human posts occupy rare structural configurations. All effects replicate StoryScope's, consistent in direction and larger in magnitude. We release pipeline, instrument, prompts, code, and aggregate artifacts.
许愿