Understanding and Mitigating Language Confusion in LLMs

LLM SFT NLP MLRCC CLKTA Benchmarks
2024年06月28日
我们调查了LLMs的一个令人惊讶的限制:它们无法在用户期望的语言中始终生成文本。我们创建了语言混淆基准(LCB)来评估这种失败,涵盖了15种类型上不同的语言,包括现有的和新创建的英语和多语种提示。我们评估了一系列LLMs在单语和跨语言生成上的表现,反映了实际使用情况,发现Llama Instruct和Mistral模型表现出高度的语言混淆,即使是最强大的模型也无法始终以正确的语言回应。我们观察到基础和以英语为中心的指导模型更容易出现语言混淆,这在复杂提示和高采样温度下更为严重。我们发现,通过少量提示、多语言SFT和偏好调整,可以部分缓解语言混淆。我们发布了我们的语言混淆基准,它作为一种高效、可扩展的多语言评估的第一层,网址为https://github.com/for-ai/language-confusion。
We investigate a surprising limitation of LLMs: their inability to consistently generate text in a user's desired language. We create the Language Confusion Benchmark (LCB) to evaluate such failures, covering 15 typologically diverse languages with existing and newly-created English and multilingual prompts. We evaluate a range of LLMs on monolingual and cross-lingual generation reflecting practical use cases, finding that Llama Instruct and Mistral models exhibit high degrees of language confusion and even the strongest models fail to consistently respond in the correct language. We observe that base and English-centric instruct models are more prone to language confusion, which is aggravated by complex prompts and high sampling temperatures. We find that language confusion can be partially mitigated via few-shot prompting, multilingual SFT and preference tuning. We release our language confusion benchmark, which serves as a first layer of efficient, scalable multilingual evaluation at https://github.com/for-ai/language-confusion.
许愿