{"slug": "surgical-video-generation-from-diffusion-to-world-models-a-survey", "title": "Surgical Video Generation From Diffusion to World Models: A Survey", "summary": "A new survey from arXiv (2608.26214v1) organizes the 2024-2026 literature on surgical video generation into three categories—unconditional generation, conditional generation, and world modeling generation—highlighting a shift from synthesizing visually plausible frames to modeling causal dynamics of surgical scenes. The authors identify generalization, physical realism, controllability, and interpretability as key bottlenecks, and summarize experimental results on public datasets to provide a quantitative reference for researchers in intelligent perception, multi-modal fusion, generative AI, and surgical data science.", "body_md": "arXiv:2608.26214v1 Announce Type: new\nAbstract: Surgical video data provides the primary training resource for models of intraoperative perception, surgical workflow understanding, and robotic decision-making. However, clinical data acquisition remains constrained by privacy, cost, and class imbalance. Surgical video generation has emerged as a transformative approach to addressing data scarcity and as a foundation for surgical simulation, training, and robotic policy learning. The field has developed rapidly without a clear conceptual framework. This survey organizes the 2024-2026 literature into three categories: unconditional generation, conditional generation, and world modeling generation, revealing a fundamental shift in how the task is defined from synthesizing visually plausible frames to modeling the causal dynamics of surgical scenes. We examine the persistent gap between pixel-level fidelity and clinical plausibility, and identify generalization, physical realism, controllability, and interpretability as bottlenecks. We further summarize experimental results of representative methods on public datasets to provide a quantitative reference for the field. This survey provides a structured overview of the current state and open challenges, offering a reference for researchers working at the intersection of intelligent perception, multi-modal fusion, generative AI, and surgical data science.", "url": "https://wpnews.pro/news/surgical-video-generation-from-diffusion-to-world-models-a-survey", "canonical_source": "https://arxiv.org/abs/2608.26214", "published_at": "2026-08-28 04:00:00+00:00", "updated_at": "2026-08-28 04:21:51.648635+00:00", "lang": "en", "topics": ["generative-ai", "artificial-intelligence", "computer-vision"], "entities": ["arXiv"], "alternates": {"html": "https://wpnews.pro/news/surgical-video-generation-from-diffusion-to-world-models-a-survey", "markdown": "https://wpnews.pro/news/surgical-video-generation-from-diffusion-to-world-models-a-survey.md", "text": "https://wpnews.pro/news/surgical-video-generation-from-diffusion-to-world-models-a-survey.txt", "jsonld": "https://wpnews.pro/news/surgical-video-generation-from-diffusion-to-world-models-a-survey.jsonld"}}