{"slug": "4dhumandiff-direct-text-to-4dgs-generation-for-consistent-360-degree-dynamic", "title": "4DHumanDiff: Direct Text-to-4DGS Generation for Consistent 360-Degree Dynamic Humans", "summary": "Researchers introduced 4DHumanDiff, a diffusion framework that directly generates dynamic human assets represented by 4D Gaussian Splatting (4DGS) from text prompts, avoiding video pre-generation and per-scene reconstruction. The model, built on a 3D U-Net backbone with temporal attention, generates consistent 360-degree dynamic humans within one minute, achieving better temporal and multi-view consistency and reducing inference time by more than 10x. The team also constructed a large-scale text-to-4DGS dataset with 60,000 high-quality pairs and introduced 2D regularization and training-free 4D interpolation to improve rendering quality and motion smoothness.", "body_md": "arXiv:2607.27634v1 Announce Type: new\nAbstract: Generating high-quality 360-degree dynamic human assets from text prompts is challenging. Existing methods usually synthesize monocular or multi-view videos first and then fit a 4D representation, which is expensive and often causes incomplete geometry or view-inconsistent renderings. We present 4DHumanDiff, a diffusion framework that directly generates dynamic humans represented by 4D Gaussian Splatting (4DGS) from text prompts. By modeling the structured 4D representation space end-to-end, 4DHumanDiff avoids video pre-generation and per-scene reconstruction, making it better suited for view-consistent and temporally coherent asset generation. The model uses a 3D U-Net backbone with temporal attention for motion-aware generation. We further construct a large-scale text-to-4DGS dataset with 60,000 high-quality pairs, and introduce 2D regularization and training-free 4D interpolation to improve rendering quality and motion smoothness. Experiments show that 4DHumanDiff generates consistent 360-degree dynamic humans within one minute, achieves better temporal and multi-view consistency, and reduces inference time by more than 10x.", "url": "https://wpnews.pro/news/4dhumandiff-direct-text-to-4dgs-generation-for-consistent-360-degree-dynamic", "canonical_source": "https://arxiv.org/abs/2607.27634", "published_at": "2026-07-31 04:00:00+00:00", "updated_at": "2026-07-31 04:39:52.563945+00:00", "lang": "en", "topics": ["generative-ai", "computer-vision", "artificial-intelligence"], "entities": ["4DHumanDiff", "4D Gaussian Splatting", "3D U-Net"], "alternates": {"html": "https://wpnews.pro/news/4dhumandiff-direct-text-to-4dgs-generation-for-consistent-360-degree-dynamic", "markdown": "https://wpnews.pro/news/4dhumandiff-direct-text-to-4dgs-generation-for-consistent-360-degree-dynamic.md", "text": "https://wpnews.pro/news/4dhumandiff-direct-text-to-4dgs-generation-for-consistent-360-degree-dynamic.txt", "jsonld": "https://wpnews.pro/news/4dhumandiff-direct-text-to-4dgs-generation-for-consistent-360-degree-dynamic.jsonld"}}