Ensuring Consistency for In-Image Translation

NLP NMT Multimodal Intelligence Vision-Language Pre-training TGIE
2024年12月24日
图像内机器翻译任务涉及将嵌入图像中的文本进行翻译,并以图像格式呈现翻译结果。虽然这一任务在电影海报翻译和日常场景图像翻译等各种场景中有许多应用,但现有的方法经常忽视了整个过程的一致性。我们认为在这个任务中需要保持两种一致性:翻译一致性和图像生成一致性。前者是指在翻译过程中融入图像信息,而后者是指保持文本图像与原始图像的风格一致性,确保背景完整性。为了解决这些一致性要求,我们提出了一种新的两阶段框架,称为HCIIT(高一致性图像内翻译),该框架首先使用多模态多语言大模型进行文本图像翻译,然后在第二阶段使用扩散模型进行图像填充。在第一阶段,采用链式思维学习来增强模型在翻译过程中利用图像信息的能力。随后,一个经过训练以生成风格一致的文本图像的扩散模型确保了图像中文本风格的统一,并保留了背景细节。我们整理了一个包含40万个风格一致的伪文本图像对的数据集用于模型训练。在整理的测试集和真实图像测试集上获得的结果验证了我们框架在确保一致性和生成高质量翻译图像方面的有效性。
The in-image machine translation task involves translating text embedded within images, with the translated results presented in image format. While this task has numerous applications in various scenarios such as film poster translation and everyday scene image translation, existing methods frequently neglect the aspect of consistency throughout this process. We propose the need to uphold two types of consistency in this task: translation consistency and image generation consistency. The former entails incorporating image information during translation, while the latter involves maintaining consistency between the style of the text-image and the original image, ensuring background integrity. To address these consistency requirements, we introduce a novel two-stage framework named HCIIT (High-Consistency In-Image Translation) which involves text-image translation using a multimodal multilingual large language model in the first stage and image backfilling with a diffusion model in the second stage. Chain of thought learning is utilized in the first stage to enhance the model's ability to leverage image information during translation. Subsequently, a diffusion model trained for style-consistent text-image generation ensures uniformity in text style within images and preserves background details. A dataset comprising 400,000 style-consistent pseudo text-image pairs is curated for model training. Results obtained on both curated test sets and authentic image test sets validate the effectiveness of our framework in ensuring consistency and producing high-quality translated images.
许愿