China's artificial intelligence ambitions are hitting a wall that no amount of chip stockpiling can solve: the country is running out of high-quality Chinese-language training data. While US export controls on advanced semiconductors have dominated debate over China's AI capabilities, Chinese experts increasingly warn that data scarcity could be just as limiting. Chinese accounts for only 1.3% of global web content, according to internet tracker W3Techs, compared with nearly half for English, 6% for Spanish, and 5% for Japanese.
The problem is global in scope but hits China particularly hard. Large language models depend on vast amounts of human-written text to learn grammar, reasoning, cultural context, and domain knowledge. When that text is scarce, models must be trained with less material, often yielding weaker performance in native-language tasks. OpenAI co-founder Andrej Karpathy has warned of a "data wall" by the end of the decade, and Epoch AI estimates that the worldwide supply of high-quality, publicly available text could be fully exhausted within six years. For Chinese developers, that wall is already visible.
The data gap behind the chip debate
Much of the Western policy discussion about China's AI sector has focused on hardware restrictions. The US government has imposed successive rounds of export controls on advanced chips, chipmaking equipment, and related technology, aiming to slow the development of Chinese supercomputing and AI models. In response, Chinese firms have stockpiled GPUs, accelerated domestic chip development, and sought alternative sources of supply. But data is different. Unlike semiconductors, high-quality training text cannot be fabricated in a factory or smuggled through third countries. It must be produced, collected, cleaned, and made accessible to researchers and companies.
Chinese-language digital content is scarce in absolute terms. English enjoys the advantages of being the lingua franca of science, business, and the internet, generating an enormous share of global web traffic and scholarly output. Spanish and Japanese, with smaller speaker populations than Mandarin, still each have more online content than Chinese. The reasons are historical and structural: Chinese internet culture grew rapidly but largely within domestic platforms that do not always encourage open data sharing; much of the country's offline knowledge has not yet been digitized; and the language barrier limits the global creation of Chinese-language content by non-native speakers.
That scarcity has economic consequences. Chinese developers reportedly pay more per useful token than their Western counterparts because they must work harder with less native-language material. Models trained on insufficient data may require more fine-tuning, more manual annotation, and more iterative training cycles to reach acceptable levels of accuracy. For startups and research labs with limited budgets, these costs can be prohibitive, widening the gap between Chinese AI systems and the best English-centric models.
Walled gardens and closed data
China's digital ecosystem makes the shortage worse. The country's largest online platforms, including WeChat and Douyin, operate as walled gardens, collecting enormous amounts of user-generated text but sharing almost none of it with third-party developers. AI labs cannot easily access conversations, posts, comments, and articles that are locked inside these platforms. As a result, they are left to train on lower-quality sources such as web scrapes, government documents, and academic papers—useful, but not enough to replicate the rich diversity of human expression found on open social media and forums.
This is a sharp contrast with the English-language AI ecosystem, where researchers can draw on huge public datasets derived from Reddit, Wikipedia, books, academic articles, and open web crawls. Even when English-language data is messy, its sheer volume allows models to learn robust patterns. Chinese AI developers do not enjoy the same abundance, and the fragmented nature of China's internet further complicates efforts to build centralized, high-quality corpora.
Platforms have little incentive to open up. Data is a competitive asset, and sharing it could erode commercial advantages. Even where institutions or regulators have pushed for data sharing, progress has been slow. The result is a paradox: one of the world's largest internet populations generates vast amounts of data, but very little of it is available for AI training outside the platforms that collect it.
Beijing's strategic response
Beijing is treating data as strategic infrastructure. In June, the National Data Administration unveiled a nationwide plan to build validated AI training datasets by 2028, covering manufacturing, energy, healthcare, finance, agriculture, autonomous driving, and embodied AI. The plan aims to create reliable, standardized data resources that can be used by researchers and companies across the country. “Competition in the AI era is not only about models and computing power, but also about high-quality data supply systems,” said Yu Xiaohui, president of the state-affiliated China Academy of Information and Communications Technology.
Tsinghua University computer scientist Sun Maosong has urged authorities to digitize historical archives, ancient manuscripts, scientific literature, and regional dialects. Such materials are not only culturally valuable; they also contain linguistic patterns and knowledge that cannot be found in modern web text. Dialects, in particular, are an underrepresented aspect of Chinese-language data. Mandarin is the standard, but hundreds of regional varieties exist, and many have no widely used written form. For AI systems to truly serve China's population, they may need to understand or at least process these varieties, adding another layer of complexity to the data challenge.
Government-led dataset projects could help in the long term, but they are not quick fixes. Building validated datasets requires time, coordination, and significant investment. It also requires content owners to participate voluntarily or be compelled by regulation. The publishing industry, however, is already starting to lock the door.
Publishers push back
Not everyone wants to be digitized for AI consumption. Huaxia Publishing House recently added a warning to a new translation: “It is prohibited to use the content of this book for artificial intelligence training. Violators will be held legally responsible.” The statement reflects a growing unease across the publishing industry about AI companies using copyrighted texts without permission or compensation. Similar debates have erupted in the United States and Europe, where authors and publishers have sued AI developers over the use of their works in training datasets.
China's publishing industry is less globally visible than its Anglo-American counterparts, but it is large and increasingly aware of AI's commercial implications. If major Chinese publishers follow Huaxia's lead, the pool of available Chinese-language text could shrink further. Books, academic monographs, and literary works have long been considered high-quality training data because they are edited, structured, and fact-checked. Losing access to such sources would be a significant blow for Chinese AI labs already struggling with data shortages.
The contrast with the United States is striking. Anthropic's Project Panama reportedly bought and destroyed millions of physical books to scan them, creating a vast high-quality corpus for training its AI models. The project illustrates how far some American companies are willing to go to secure data. Chinese firms, by contrast, face a more fragmented landscape where public libraries, archives, and publishers are not yet centrally organized to support AI research on a similar scale.
Why language data matters for AI performance
The importance of native-language training data cannot be overstated. Advanced AI models are built on pattern recognition, and patterns differ from language to language. English-language texts encode certain ways of reasoning, storytelling, and knowledge organization; Chinese-language texts encode different ones. A model trained mostly on English may produce fluent and accurate responses in English, but its Chinese output can feel stilted, culturally tone-deaf, or factually unreliable when dealing with topics anchored in Chinese society, history, and politics.
There are technical approaches to bridge the gap. Multilingual models can transfer knowledge from high-resource languages like English to low-resource languages through cross-lingual learning. Chinese AI developers can and do use English datasets to supplement their native-language material. But this approach has limits. Subtle meanings, idioms, proverbs, and cultural references are often lost in translation. Moreover, some knowledge is uniquely expressed in Chinese: government policies, regional laws, historical scholarship, and local news often have no equivalent in English. A model that cannot read these sources cannot answer questions about them accurately.
Synthetic data is another possible solution. AI systems can generate artificial training examples that mimic real text, and some researchers argue that synthetic data can reduce dependency on scarce human-generated content. However, synthetic data carries risks of amplifying biases and errors. If the underlying real data is limited, the synthetic data will only reproduce those limitations. Many experts view synthetic data as a supplement, not a fix for the underlying shortage.
The long road to a Chinese data infrastructure
China's data problem is not new, but it is becoming more urgent as the world moves toward increasingly large and powerful AI models. The first generation of Chinese AI systems relied heavily on open source English-language datasets, adapting them for Chinese tasks. That approach worked while models were smaller and benchmarks were simpler. Today, cutting-edge models are trained on trillions of tokens, and the competition has shifted to who can assemble the best and largest corpus. China's 1.3% share of web content is a severe disadvantage in this race.
There are some promising signs. China has a rich textual tradition, with thousands of years of historical records, literary classics, philosophical treatises, and scientific writings. Digitizing these materials could provide unique high-quality data that no other country can offer. Regional archives, local gazetteers, and traditional medical texts are untapped resources that could give Chinese models a distinctive edge. The challenge is that digitization is expensive and slow, and the most valuable documents are often the hardest to scan and transcribe.
The Chinese government's plan to create national datasets by 2028 reflects an understanding that data is not merely a technical asset but a matter of strategic competitiveness. Whether the plan can be implemented in time is another question. AI development is moving quickly, and models training today need data today, not in four years. The gap between the demand for Chinese-language data and the supply of high-quality sources is likely to widen before it narrows.
What is clear is that the AI race is no longer just about computing power. It is also about who controls the words, images, and knowledge shaped by human culture. For China, the next bottleneck may not be a chip fab or a semiconductor lithography machine, but the difficulty of digitizing, preserving, and opening up its own linguistic heritage. With platforms sealing their data and publishers adding AI training bans, the window for building comprehensive Chinese-language datasets is closing. Whether these interventions can close the language-data gap remains an open question, but they highlight a fundamental reality: in the AI era, data is strategic terrain.