Habibi: Laying the Open-Source Foundation of Unified-Dialectal Arabic Speech Synthesis

GenAI TTS
目前,阿拉伯语方言的语音合成研究与开发仍存在显著空白,尤其在统一建模视角下尤为突出。尽管该方向具有极高的实际应用价值,但阿拉伯语方言固有的语言复杂性,加之缺乏标准化的数据集、基准测试体系及评估规范,致使研究人员往往倾向于选择更稳妥的研究路径。为弥合这一鸿沟,我们推出了“哈比比”(Habibi)——一套专门设计且具备统一架构的文本到语音(TTS)模型。该模型充分利用现有的开源自动语音识别(ASR)语料库,借助语言学驱动的课程学习策略,有效支持从高资源到低资源的多种阿拉伯语方言。实验结果表明,我们的方法在语音生成质量上超越了当前领先的商业语音合成服务;同时,该模型无需对输入文本进行变音符号(diacritization)标注,即可通过高效的上下文内学习(in-context learning)保持良好的可扩展性。我们承诺将模型完全开源,并构建首个面向多方言阿拉伯语语音合成的系统性基准测试集。此外,我们还深入剖析了该领域面临的核心挑战,确立了科学、可复现的评估标准,旨在为后续研究奠定坚实基础。相关资源详见:https://SWivid.github.io/Habibi/。
A notable gap persists in speech synthesis research and development for Arabic dialects, particularly from a unified modeling perspective. Despite its high practical value, the inherent linguistic complexity of Arabic dialects, further compounded by a lack of standardized data, benchmarks, and evaluation guidelines, steers researchers toward safer ground. To bridge this divide, we present Habibi, a suite of specialized and unified text-to-speech models that harnesses existing open-source ASR corpora to support a wide range of high- to low-resource Arabic dialects through linguistically-informed curriculum learning. Our approach outperforms the leading commercial service in generation quality, while maintaining extensibility through effective in-context learning, without requiring text diacritization. We are committed to open-sourcing the model, along with creating the first systematic benchmark for multi-dialect Arabic speech synthesis. Furthermore, by identifying the key challenges in and establishing evaluation standards for the process, we aim to provide a solid groundwork for subsequent research. Resources at https://SWivid.github.io/Habibi/ .
许愿