Why China Is Winning the AI Video Race ByteDance's Seedance 2.5 and MiniMax's H3, released on the same day two weeks ago, have propelled China to dominate the AI video race, with nine of the top 10 text-to-video systems on Artificial Analysis now made in China. OpenAI's retreat from Sora as a standalone video business, partly due to estimated operating costs of $15 million per day, highlights that China's success stems not just from model performance but from an integrated economy where creators on platforms like Douyin pay for credits and API calls, sustaining the models. Why China Is Winning the AI Video Race Seedance and MiniMax are succeeding for a reason that has little to do with benchmarks: China built an economy around generating video. Two weeks ago, ByteDance released Seedance 2.5 https://seed.bytedance.com/en/seedance2 5 on the same day MiniMax unveiled H3 https://r.search.yahoo.com/ ylt=Awr1SSs1uX5qagIA3q fzAt.; ylu=Y29sbwNzZzMEcG9zAzMEdnRpZAMEc2VjA3Ny/RV=2/RE=1787899445/RO=10/RU=https%3a%2f%2fwww.minimax.io%2fblog%2fminimax-h3/RK=2/RS=4qLz3pnLSRbPzGauXp0iJEOh1W8- , setting off another round of excitement https://www.bloomberg.com/news/articles/2026-07-31/china-s-minimax-and-bytedance-release-dueling-ai-video-models around AI video. H3 emphasized visual effects and post-production workflows. Seedance 2.5 pushed further into editing, letting creators make precise changes to existing clips instead of generating everything from scratch. In the U.S., the category feels much quieter. OpenAI has retreated from Sora as a standalone video business https://www.cnbc.com/2026/03/24/openai-shutters-short-form-video-app-sora-as-company-reels-in-costs.html , while the biggest American AI companies have not produced anything that has matched the momentum around Seedance. On Artificial Analysis https://artificialanalysis.ai/video/leaderboard/text-to-video ’ text-to-video leaderboard, nine of the top 10 systems are now made in China. When OpenAI began stepping back from Sora, some Chinese technology publications concluded that generative video was simply a bad business. Instead, American retrenchment created more room for Chinese companies. AI-generated shorts are increasingly accepted on Chinese video platforms. Creators use them for memes https://www.asiaone.com/digital/did-you-save-fox-snowy-mountain-viral-chinese-ai-video , fictional stories https://x.com/PJaccetturo/status/2053475536538845522?s=20 , short dramas https://www.sixthtone.com/news/1018827 and advertising, then pay for more credits or API calls to make the next one. Those payments keep the video models alive. That is the part of China’s success that is easy to miss. This is not just a story about model performance. It is a story about operations, distribution and, ultimately, the business environment around the model. Sora Had a Model. It Never Had an Economy. This helps explain what went wrong with Sora. OpenAI is excellent at building AI products, but it has never been a content-platform company. ChatGPT is fundamentally private: you ask something, it answers you. Your output is not automatically placed in front of millions of other users. Sora inherited that weakness. A creator could generate a video and post it to YouTube, TikTok or X. But OpenAI did not own the audience on the other side. There was no built-in economic loop that paid creators for producing more Sora videos, and therefore little reason for an independent studio to keep buying enormous volumes of Sora inference. OpenAI experimented https://www.axios.com/2025/09/30/openai-sora-app-social-ai with a more social Sora experience, but building a new video destination is very different from adding video generation to an existing AI product. People already have TikTok, YouTube and Instagram. They do not necessarily need another app whose defining feature is that everything inside it is synthetic. That is particularly painful because video inference is brutally expensive. SemiAnalysis at one point estimated Sora’s operating cost https://r.search.yahoo.com/ ylt=Awrx wQIvH5qQgIAsmHfzAt.; ylu=Y29sbwNzZzMEcG9zAzIEdnRpZAMEc2VjA3Ny/RV=2/RE=1787900169/RO=10/RU=https%3a%2f%2fwww.forbes.com.au%2fnews%2finnovation%2fopenai-could-be-blowing-as-much-as-15-million-per-day-on-silly-sora-videos%2f/RK=2/RS=kLaAaIfYaVzODc7sEX.ROW0Y2gc- at roughly $15 million per day, equivalent to an annualized burn rate of around $5.4 billion. Even if that estimate is only directionally correct, the economics illustrate the problem: a video model needs a very large amount of paid usage to justify its compute bill. ByteDance can subsidize that loop because the model strengthens businesses it already owns. It can build the model, help creators make the video, distribute the video, monetize the traffic and then sell the creator more tokens. OpenAI cannot reproduce that chain. The U.S. has another obstacle: copyright. Video models inevitably collide with recognizable actors, characters and visual styles. Soon after Sora’s rollout, copyrighted characters from major Hollywood studios began appearing in generated clips, prompting objections from studios https://r.search.yahoo.com/ ylt=AwrKD3ZgvH5qPgIAW8PfzAt.; ylu=Y29sbwNzZzMEcG9zAzMEdnRpZAMEc2VjA3Ny/RV=2/RE=1787900256/RO=10/RU=https%3a%2f%2fdeadline.com%2f2025%2f10%2fsora-2-hollywood-ai-sam-altman-1236572662%2f/RK=2/RS=FwZ3YJ.ZcV3U5vpoIFdN32FySG0- , talent representatives and rights holders. The legal objections are understandable. The commercial consequence is still awkward. Stronger guardrails mean more prompts get rejected or constrained, making the tool less predictable for creators who want to work with familiar cultural references. Seedance faces copyright pressure too. Rights holders including Disney have challenged https://www.bbc.com/news/articles/c93wq6xqgy1o the use of protected material in generative video, and ByteDance has responded by restricting some capabilities involving real people, faces and copyrighted characters. But ByteDance has one advantage OpenAI does not: Seedance does not need the American market to survive. China alone can provide creators, audiences and paying customers. Google theoretically has the closest American equivalent. It has video models and it owns YouTube. But YouTube has little reason to aggressively subsidize a flood of cheap AI video when the platform is already trying to manage audience frustration with low-quality synthetic content. The technology may be similar. The incentives are not. Seedance Has Something Sora Never Had: Douyin When Seedance 2.0 broke out earlier this year, Western coverage mostly treated it as another sign that China had caught up in AI video. The more important question was where all that AI slop was going. The answer was Douyin https://www.nytimes.com/2026/05/03/world/asia/china-microdrama-ai-backlash.html , ByteDance’s Chinese version of TikTok. AI-generated content has faced less cultural resistance in China than it has in parts of the U.S. Before Seedance, Chinese creators were already using Sora https://www.bilibili.com/video/BV1y6x7zuEGY/?spm id from=333.337.search-card.all.click , Google’s video models and MiniMax’s Hailuo AI to make memes. The output was often obviously synthetic: soft depth of field, strange faces and physics that did not quite make sense. But it created an early generation of users who learned how to make AI videos entertaining. Then Seedance became much more usable. Its early viral examples did not simply show beautiful random scenes generated from text. Creators demonstrated familiar cinematic styles— old Hong Kong martial-arts films https://v.douyin.com/elDe7 uxLP0/ , Japanese tokusatsu https://v.douyin.com/Y3mXF--fekU/ and tightly choreographed fight scenes—while keeping characters, costumes and environments unusually consistent. That matters more than prettier demos. For creators, a video model becomes useful when the same character can survive from one shot to the next. A tool that produces one spectacular clip and then loses the character’s face is a lottery ticket. A model that can preserve the visual asset is production software. ByteDance also has an enormous structural advantage through Douyin, one of the world’s largest short-video platforms. It gives the company access to an immense ecosystem of short-form visual content, creator behavior and immediate distribution. Exactly how user content feeds model training is not publicly disclosed, but the product advantage is obvious: ByteDance understands what people make, what they watch and what they remix. That created a loop. Seedance users made videos. They uploaded them to Douyin. Viral examples attracted more people to Seedance, who created their own variations and pushed the model into new meme formats. Then ByteDance added another layer: short dramas. The company owns TomatoFiction (番茄小说) an enormous source of serialized web fiction, and Hongguo (红果短剧,another name is Melolo), its short-drama platform https://www.nytimes.com/2026/05/03/world/asia/china-microdrama-ai-backlash.html . The novels provide ready-made plots filled with romance, betrayal, revenge and endless cliffhangers. Hongguo provides somewhere to monetize the resulting videos. ByteDance’s familiar strategy is to subsidize creators first, then make traffic the performance metric. In AI video, that strategy worked almost too well. By early 2026, AI-generated animated dramas had become one of China’s hottest content businesses. According to 36Kr https://finance.biggo.com/news/98d36dcf-1051-463b-a962-430c503d70ca , some production companies began buying the highest-tier Seedance 2.0 API packages through ByteDance’s Volcano Engine. One reported top-up reached $7.4 million . Seedance 2.0’s 720p generation reportedly cost about $6.82 , per million tokens, translating to roughly 15 cents , per second of video in one common configuration. Kling AI from Kuaishou Technology could be cheaper, at roughly 9 cents per second, while Vidu from ShengShu Technology could come in near 10 cents . But reliability changes the calculation. Saving a few cents does not matter if a generation has to be thrown away. Bilibili is moving in the same direction with its AI creation tools https://www.updream.cn/ , allowing users to build a story, supply images and other assets, and then generate the video through models including Seedance. Seedance therefore has something more valuable than a benchmark score: inputs, creators, distribution and a reason to keep generating. China Is Turning Video Generation Into a Production Line MiniMax H3 pushes the argument one step further. Released alongside Seedance 2.5, H3 received less attention. But its significance is different: MiniMax is trying to prove that AI video can be industrialized. Its H3-Context-IR https://r.search.yahoo.com/ ylt=Awrx Ogown5qXAIAaerfzAt.; ylu=Y29sbwNzZzMEcG9zAzEEdnRpZAMEc2VjA3Ny/RV=2/RE=1787901736/RO=10/RU=https%3a%2f%2fplatform.minimaxi.com%2fdocs%2fapi-reference%2fvideo-generation-v2-h3-context-ir/RK=2/RS=TlNtjlfeQQbbne78ybIIlSM2DaI- system uses an “intermediate representation,” borrowing a term familiar from software compilers. Text, images, video and audio can be broken into elements such as characters, scenes, actions and sound, then reused across generations. The practical goal is consistency. That is what AI video needs most. Once assets can move reliably between shots without characters randomly changing faces or objects disappearing, generation stops looking like a demo and starts looking like a workflow. Seedance 2.5 is moving toward the same destination from another direction. Its editing tools increasingly resemble a video editor rather than a prompt box. Creators can modify specific time ranges https://www.atlascloud.ai/zh/blog/ai-updates/seedance-2.5-whatsnew , extract camera movement and visual styles from existing footage, and reuse those elements elsewhere https://www.atlascloud.ai/zh/blog/ai-updates/seedance-2.5-whatsnew . Third-party platforms such as LibTV https://r.search.yahoo.com/ ylt=AwrKBA0Ww35qSQIAiT7fzAt.; ylu=Y29sbwNzZzMEcG9zAzEEdnRpZAMEc2VjA3Ny/RV=2/RE=1787901975/RO=10/RU=https%3a%2f%2fwww.liblib.tv%2f%3fsourceid%3d005902%26utm%3dcg%26cgv%3d9omkl4jn4d/RK=2/RS=npalGC.LyHGKxMR8IjT42emEr2g- are already building workflows around these capabilities. H3 wants to become part of post-production. Seedance wants generation to feel more like editing software. Both approaches matter because video is one of the most aggressive ways to consume tokens. Industry estimates https://stock.10jqka.com.cn/20260604/c677226962.shtml cited by COL Group suggest that short dramas and video generation already account for 55% share of AI token usage in China. Whether the exact percentage holds across the entire market is difficult to verify, but the direction is clear: video consumes far more inference than a chatbot response. That makes every improvement in video workflow economically important. Better consistency means fewer wasted generations. Easier editing means more generations. More creators mean more tokens. Seedance is being integrated into ByteDance products such as Dreamina and CapCut. Kling AI and MiniMax sell relatively inexpensive subscription credits to individual creators. A freelancer can now insert generative video into a workflow for an advertisement, an e-commerce clip, a cheap short drama or a visual effect that would previously have required a small production team. Most of the output still looks like junk food. But junk food can support an enormous industry. More usage produces more revenue. Revenue funds better models. Better models make creators more willing to use them for real work. That feedback loop is what the U.S. video-model market has struggled to establish. MiniMax, which does not own a Douyin-sized platform, has to work harder. Its opportunity is to sell directly into advertising, games, e-commerce and post-production—industries where generative video penetration is still low. Someone will always need another commercial. H3 is trying to become the cheapest person in the editing room. The Bigger Prize Is Not Video AI-generated short dramas are not why I care about video models. The more interesting destination is the world model . To generate convincing video, a model eventually has to learn more than pixels. It needs some representation of motion, causality, object permanence and physics. A glass falls because gravity exists. A person walking behind a wall should still exist when they emerge from the other side. A robot pushing a box needs to understand that the box pushes back. There is an interesting historical parallel with Nvidia https://r.search.yahoo.com/ ylt=Awrx OhKw35qbQIAcr fzAt.; ylu=Y29sbwNzZzMEcG9zAzIEdnRpZAMEc2VjA3Ny/RV=2/RE=1787902026/RO=10/RU=https%3a%2f%2fwww.nvidia.com%2fen-us%2fabout-nvidia%2fcorporate-timeline%2f/RK=2/RS=MENJhHquVdx0yBynR .ywlnFkWE- . GPUs were originally pushed forward by gamers demanding better graphics. But rendering better graphics required increasingly sophisticated simulation of light, surfaces and physical movement. The architecture built for that visual workload eventually became the foundation for accelerated computing—and then modern AI. Video models may follow a similarly indirect path. A video model is not automatically a world model. But a system that becomes increasingly good at predicting how visual environments change over time is learning something that language alone cannot provide. According to LatePost, ByteDance is considering training a model with more than 5 trillion parameters. That would make it larger than Alibaba’s Qwen 3.8-Max, with 2.4 trillion parameters, and Moonshot AI’s K3, with 2.8 trillion—making it the largest model publicly known to be under development in China. The model would almost certainly be multimodal, meaning it could support both text and visual reasoning from launch. That would move it another step closer to the capabilities expected of a world model. That matters for robotics and autonomous driving. China currently has hundreds of companies trying to build humanoid robots, yet the industry still struggles with the same fundamental problem: bodies are arriving faster than brains. Unitree says cumulative humanoid shipments have already reached 18000 https://x.com/UnitreeRobotics/status/2087475885658210719?s=20 , while many Western competitors remain at much smaller production volumes. Physical training data is expensive. Robots have to exist, move, fail and be reset. A sufficiently capable video or world model could eventually provide a simulated environment in which some of that learning happens before the robot ever leaves the factory. I have always been cautious about the idea that language models alone can lead to general intelligence. Language cannot fully describe smell, light or the complicated physical relationship between a hand and an object. Humans learn by interacting with the world; language is, in many ways, the compressed record of those interactions. AI may need something similar. Video offers an imperfect simulated world in which a machine can begin learning what happens when things move, collide and change. As Bloomberg columnist Catherine Thorbecke https://www.bloomberg.com/opinion/articles/2026-08-09/chinese-ai-video-is-coming-for-more-than-hollywood has argued, large language models still absorb most of the attention and capital in AI. But the next major breakthrough may come from systems that can navigate the real world rather than merely describe it. If video becomes the training ground for those systems, China’s early advantage will matter for far more than annoying Hollywood. It could help determine who leads the next phase of AI.