Friday 28 August 2026
Sarabangla English
বাংলা.
Home » Tech

China’s AI video boom could give it edge in world models

SB Desk
19 August 2026 15:13 Updated: 19 August 2026 15:13

China’s growing dominance in artificial intelligence video generation could give the country an advantage in developing so-called “world models”, potentially extending its AI influence far beyond entertainment and Hollywood.

While large language models such as Moonshot’s Kimi K3 have attracted much of the attention surrounding China’s AI ambitions, video generation may offer a clearer indication of the country’s growing capabilities. Outside Google, nearly all of the leading video-generation models are Chinese, accounting for nine of the top 10 text-to-video systems on Artificial Analysis’ benchmark leaderboard.

ByteDance and MiniMax Group recently released major updates to their competing video-generation models, following the surprise emergence of “Happy Horse”, a previously secretive video system that drew industry attention before being identified as an Alibaba Group project.

Other major Chinese players include Kuaishou Technology, with its Kling model, and Beijing-based ShengShu Technology, which developed Vidu. Together with a growing number of startups, they have created an unusually competitive Chinese ecosystem in a field that major US AI companies, including OpenAI and Anthropic, have given less priority.

The technology is already being adopted by advertising agencies and entertainment studios and has helped fuel the rapid growth of the microdrama industry in China and overseas. Some industry observers see AI-generated video as a potential source of revenue for an entertainment sector struggling with price competition and uncertain business models.

Advertisement

But the strategic significance of the technology could be much greater.

From video generation to ‘world models’

Researchers and business leaders increasingly believe video-generation systems could become the foundation for AI systems capable of operating in the physical world.

Such “world models” would need to understand more than language. They would have to learn how objects move, how physical forces work and how different objects interact in unpredictable real-world environments.

The technology could eventually be used to power humanoid robots, autonomous vehicles and other physical AI systems. Some AI researchers believe this approach could offer a more promising path toward artificial general intelligence than simply building increasingly powerful chatbots.

Producing convincing video requires more than learning visual aesthetics. AI models need at least a basic understanding of motion, causality and physics to avoid producing unrealistic or absurd scenes.

As video models become more sophisticated, researchers have found that basic physical reasoning can emerge without being explicitly programmed. Although the systems do not understand the world in the same way humans do, they are becoming increasingly capable of predicting what the world should look like and what is likely to happen next.

Researchers are also beginning to use video models to train robots, linking advances in visual AI with developments in physical AI.

China moves toward broader AI systems

Several leading Chinese AI video companies, including ShengShu, are already working on world models and broader multimodal systems.

The trend is not limited to China. US-based Runway AI and Germany’s Black Forest Labs, two of the West’s leading visual AI companies, are pursuing similar directions.

Runway CEO Cris Valenzuela has described the transition as a natural progression. A capable video model must accurately simulate how events unfold in the real world. For example, a ball should roll across a field rather than appear to glide above it.

The same principle applies to robots. A robot cannot fold laundry or stock supermarket shelves simply by recognising objects. It must understand how those objects will respond when it moves or interacts with them.

The transition from video generation to world models remains at an early stage. However, after OpenAI abandoned its Sora video project, Chinese developers have increasingly taken the lead in the field.

Data and copyright challenges

China’s advantage in AI video is not without significant challenges. Video-generation models require enormous amounts of data and computing power, and developing world models could demand even greater resources.

Copyright is another major concern. While the issue affects the AI industry as a whole, it can be particularly difficult to obscure the sources of training data used by video-generation systems because their outputs can reveal characteristics of the material on which they were trained.

Copyright concerns have already emerged as a major obstacle to the overseas expansion of Chinese AI video platforms.

Despite these challenges, China’s rapid progress in generative video is positioning the country at the forefront of a technology that could eventually serve as a bridge between AI systems that understand digital content and machines capable of interacting with the physical world.

Advertisement

More

Related