Editing Motion with Words: A Training-Free Framework for Text-Guided Motion Stylization
Xinru Xu, Ye Shi
PAPER · v1.0 · 2026-09-25 · ai
Abstract
Editing motion style through textual prompts without motion exemplars remains an open challenge. Existing text-guided editing methods struggle to manipulate global motion styles while preserving action content. We present TextStyler, a training-free framework that repurposes pre-trained motion diffusion models for text-driven global stylization—editing the “how” of motion execution while maintaining the “what” of action semantics, without style-specific training or stylized motion references. We implement a KV cache that retains content-critical key-value pairs from the original motion during inversion, ensuring temporal coherence and action fidelity. During stylized denoising, we inject target style information through a text-aligned loss guidance while preserving cached semantic features through masked cross-attention. This framework enables stylization without requiring style-specific training or motion references. Through extensive experiments on the HumanML3D dataset, we show the superiority of our method over the current method.