Clarify: Improving Model Robustness With Natural Language Corrections

在监督学习中,模型被训练从静态数据集中提取相关性。这经常导致模型依赖于高级误解。为了防止这种误解,我们必须提供超出训练数据的附加信息。现有方法包括形式的附加实例级监督,例如针对虚假特征的标签或来自平衡分布的附加标记数据。由于这些策略需要近乎原始训练数据规模的附加注释,因此对于大规模数据集来说,这些策略可能成本过高。我们假设有关模型误解的有针对性的自然语言反馈是一种更有效的附加监督形式。我们引入了Clarify,一种新颖的界面和方法,用于交互式纠正模型的误解。通过Clarify,用户只需要提供一个简短的文本描述来描述模型的一致性失败模式。然后,我们完全自动化地使用这些描述来通过重新加权训练数据或收集附加的有针对性数据来改善训练过程。我们的用户研究表明,非专业用户可以通过Clarify成功地描述模型的误解,在两个数据集中平均提高了最差组的准确性17.1%。此外,我们使用Clarify在ImageNet数据集中找到并纠正了31个新的困难子群体,将少数派分割的准确性从21.1%提高到28.7%。
In supervised learning, models are trained to extract correlations from a static dataset. This often leads to models that rely on high-level misconceptions. To prevent such misconceptions, we must necessarily provide additional information beyond the training data. Existing methods incorporate forms of additional instance-level supervision, such as labels for spurious features or additional labeled data from a balanced distribution. Such strategies can become prohibitively costly for large-scale datasets since they require additional annotation at a scale close to the original training data. We hypothesize that targeted natural language feedback about a model's misconceptions is a more efficient form of additional supervision. We introduce Clarify, a novel interface and method for interactively correcting model misconceptions. Through Clarify, users need only provide a short text description to describe a model's consistent failure patterns. Then, in an entirely automated way, we use such descriptions to improve the training process by reweighting the training data or gathering additional targeted data. Our user studies show that non-expert users can successfully describe model misconceptions via Clarify, improving worst-group accuracy by an average of 17.1% in two datasets. Additionally, we use Clarify to find and rectify 31 novel hard subpopulations in the ImageNet dataset, improving minority-split accuracy from 21.1% to 28.7%.
许愿