How did the Chinese labs catch up so fast if most of the early training work was in the US? Chinese AI labs including DeepSeek closed much of the training gap with US labs such as OpenAI and Google despite the US holding a years-long head start, according to an analysis of the question. The analysis attributes the catch-up largely to open-source research papers and compute availability, while noting the data component of the equation remains unclear. OpenAI, Google, etc has a years long head start on training. Then Deepseek and the other Chinese labs closed a lot of that gap in a much shorter window. How did they do this? I know open source papers and compute factor greatly. What I don't get is the data part of the equation. Besides US data labs