Music-to-Dance Generation via Atomic Movements

Multimodal Intelligence SDFA Other MMI GenAI Other GenAI CV HPMR
2026年07月15日
音乐驱动的舞蹈生成旨在生成既在节奏上与音乐同步、又在语义上与音乐保持一致的人体运动。尽管近期基于神经网络的方法已在视觉真实感方面取得显著成果,但它们通常将运动建模为连续信号,忽视了运动本身固有的结构性与可组合性,导致所生成的舞蹈在结构上缺乏连贯性,且难以进行有效控制。本研究提出一种结构感知型框架,将编舞建模为一系列原子动作——即具有明确语义可解释性的基本运动事件,这些事件构成舞蹈的底层构建单元。为构建这一原子动作词表,我们首先对大规模舞蹈数据进行切分,并将其聚类为若干原子动作组;随后,利用大语言模型对各聚类结果进行语义重标注与精细化处理,最终获得一组兼具可解释性与可复用性的原子动作。基于上述原子动作标注,我们设计了一种两阶段生成框架,以模拟人类编舞的实际过程:在第一阶段(原子动作规划阶段),模型依据输入音乐预测每个原子动作的类型、持续时长及发生时机,从而生成一种符号化的舞蹈编排方案;在第二阶段(动作补全阶段),一个具备过渡感知能力的生成器则依据该结构化规划,合成出平滑自然、风格统一的连续运动序列。大量实验表明,相较于现有各类基线方法,本方法生成的舞蹈在结构连贯性、节奏对齐度以及感知自然度等方面均取得显著提升;同时,得益于显式的结构化表征,本方法还具备更强的可解释性与可控编辑能力。
Music-driven dance generation aims to produce human motion that is both rhythmically synchronized and semantically consistent with music. While recent neural approaches have achieved impressive visual realism, they typically model motion as a continuous signal and neglect its compositional nature, making generated dances structurally incoherent and difficult to control. In this work, we introduce a structure-aware framework that models choreography as a sequence of atomic movements-semantically interpretable motion events that serve as the building blocks of dance. To construct this atomic movement vocabulary, we first segment large-scale dance data and cluster them into atomic movement groups. We then employ a large language model to semantically relabel and refine the clusters, yielding a set of interpretable and reusable atomic movements. Based on these atomic movement annotations, we design a two-stage generation framework that mirrors the human choreography process. In the atomic movement planning stage, the model predicts the type, duration, and timing of atomic movements conditioned on the input music, forming a symbolic dance allocation. In the completion stage, a transition-aware generator synthesizes smooth and stylistically coherent motion conditioned on the planned structure. Extensive experiments demonstrate that our method produces dances with significantly improved structural coherence, rhythmic alignment, and perceptual naturalness compared to existing baselines, while providing enhanced interpretability and controllable editing through explicit structural representation.
许愿