Data-Driven Persona-Conditioned Agents for A/B Test Simulation

LLM Other LLM Agent Multi-agent Collaboration Agent Eval Benchmarks
A/B测试是评估产品变更的黄金标准,但每次实验均需消耗真实的用户流量、投入工程资源,并耗费数周时间进行效果测量。我们提出了一种仿真框架,利用基于大语言模型(LLM)的智能体来预测A/B测试结果;这些智能体以数据驱动的用户画像为条件,而该画像则扎根于真实用户的实际行为信号。与以往依赖合成数据或基于规则构建用户画像的研究不同,我们的智能体直接从匿名化的行为数据中构建而成——包括用户活动模式、参与度信号以及推断得出的人口统计学特征——从而实现对目标用户群体更忠实、更精准的建模。我们将A/B测试仿真任务建模为一项结构化问答任务,并系统性地探究了以下四个关键问题:(i)问题设计的不同形式;(ii)用户画像数据来源及其与目标业务领域的匹配程度所带来的影响;(iii)单个画像所刻画的行为深度与整体用户群体多样性之间的权衡关系;(iv)面向大规模群体的高效子采样策略。在涵盖两类核心指标、共计40组真实A/B测试构成的基准测试集上,我们最优配置的仿真方法在不同测试指标下实现了75%–90%的方向性准确率,充分表明:基于真实数据构建的用户画像,是一条切实可行的技术路径,可支撑快速、低成本的实验预筛与可行性评估。
A/B testing is the gold standard for evaluating product changes, but each experiment requires real user traffic, engineering effort, and weeks of measurement. We propose a simulation framework that predicts A/B test outcomes using LLM-powered agents conditioned on data-driven personas grounded in real user behavioral signals. Unlike prior work that relies on synthetic or rule-based personas, our agents are constructed from anonymized behavioral data-activity patterns, engagement signals, and inferred demographics-enabling more faithful population modeling. We frame A/B test simulation as a structured question task and systematically study (i) question design formats, (ii) the impact of persona data source and domain alignment, (iii) the trade-off between per-persona behavioral depth and population diversity, and (iv) efficient population subsampling. On a benchmark of 40 A/B tests spanning two metric types, our best configuration achieves 0.75-0.90 directional accuracy depending on the test metric, demonstrating that data-driven personas are a viable path toward fast, low-cost experiment pre-screening.
许愿