A Training-Free Framework for Human Motion Generation and Editing in 3D Scenes with Large Language Models

Xinru Xu, Ye Shi

PAPER · v1.0 · 2026-09-25 · human

Abstract

We propose a training-free framework for text-guided human motion generation and editing in 3D indoor scenes. Unlike prior scene-agnostic methods, our approach explicitly models spatial reasoning and physical constraints for human-scene interaction. Given a textual instruction and 3D scene, a large language model decomposes the input into structured motion prompts and localizes target objects. A stochastic A*-based planner then generates a collision-free and diverse root trajectory guided by scene geometry. Conditioned on the prompts and trajectory, a pretrained diffusion model generates human motion, which is further refined via loss-guided inference with differentiable constraints on trajectory alignment, object interaction, and collision avoidance. Experiments on SceneFun3D demonstrate improved scene consistency, motion realism, and action alignment, highlighting the effectiveness of integrating LLM reasoning and diffusion-based generation for controllable 3D human motion generation and editing.

Keywords

Diffusion Human motion generation LLM reasoning

Download PDF