See-through: Single-image Layer Decomposition for Anime Characters

GenAI Diffusion Inpainting
我们提出了一种全新框架,可自动将静态动漫插画转化为可操控的2.5D模型。当前专业工作流需依赖繁琐的手动图像分割,并凭借艺术家主观“想象”来补全被遮挡区域,方能实现动画运动;而本方法通过将单张图像分解为若干语义明确、绘制顺序经推理确定且已完全修复(即无缺失区域)的图层,从根本上克服了上述瓶颈。为应对训练数据严重匮乏的挑战,我们设计了一套可扩展的引擎,能够从商用Live2D模型中自举生成高质量监督信号,精准捕获像素级语义信息及隐含的几何结构。本方法将基于扩散模型的“身体部位一致性模块”与像素级伪深度推断机制有机结合:前者保障整体几何结构的全局一致性,后者则实现精细的空间层次解析——例如准确区分相互穿插的发丝。二者协同作用,实现了对动漫角色复杂图层结构的稳健解耦与动态重建。实验表明,本方法生成的模型不仅保真度高,且具备优异的可操控性,完全满足专业级实时动画制作的实际需求。
We introduce a framework that automates the transformation of static anime illustrations into manipulatable 2.5D models. Current professional workflows require tedious manual segmentation and the artistic ``hallucination'' of occluded regions to enable motion. Our approach overcomes this by decomposing a single image into fully inpainted, semantically distinct layers with inferred drawing orders. To address the scarcity of training data, we introduce a scalable engine that bootstraps high-quality supervision from commercial Live2D models, capturing pixel-perfect semantics and hidden geometry. Our methodology couples a diffusion-based Body Part Consistency Module, which enforces global geometric coherence, with a pixel-level pseudo-depth inference mechanism. This combination resolves the intricate stratification of anime characters, e.g., interleaving hair strands, allowing for dynamic layer reconstruction. We demonstrate that our approach yields high-fidelity, manipulatable models suitable for professional, real-time animation applications.
许愿