HILCodec: High Fidelity and Lightweight Neural Audio Codec

GenAI TTS AI Systems and Hardware MCPKD
2024年05月08日
最近,端到端神经音频编解码器的进展使得可以在非常低的比特率下压缩音频,同时以高保真度重构输出音频。然而,这种改进往往以增加模型复杂度为代价。在本文中,我们确定并解决了现有神经音频编解码器存在的问题。我们发现,Wave-U-Net的性能并不随着网络深度的增加而一致提高。我们分析了这种现象的根本原因,并提出了一种方差约束设计。此外,我们揭示了以前波形域鉴别器中的各种失真,并提出了一种新的无失真鉴别器。由此产生的模型HILCodec是一种实时流媒体音频编解码器,展示了在各种比特率和音频类型下的最先进的质量。
The recent advancement of end-to-end neural audio codecs enables compressing audio at very low bitrates while reconstructing the output audio with high fidelity. Nonetheless, such improvements often come at the cost of increased model complexity. In this paper, we identify and address the problems of existing neural audio codecs. We show that the performance of Wave-U-Net does not increase consistently as the network depth increases. We analyze the root cause of such a phenomenon and suggest a variance-constrained design. Also, we reveal various distortions in previous waveform domain discriminators and propose a novel distortion-free discriminator. The resulting model, \textit{HILCodec}, is a real-time streaming audio codec that demonstrates state-of-the-art quality across various bitrates and audio types.
许愿