None
EN
China faces new AI bottleneck as it runs out of Chinese-language training data
['Ben Jiang']
News - South China Morning Post
China’s high-stakes race to build next-generation artificial intelligence models is entering a critical new phase, where a less visible yet far more existential threat is coming into view: a severe shortage of high-quality training data.
It is a challenge confronting AI giants on both sides of the Pacific – and some US companies are already resorting to aggressive measures to stay ahead.
The global supply of high-quality, publicly available human-generated text could be fully exhausted within the next six years, according to US-based research institute Epoch AI.
OpenAI co-founder Andrej Karpathy has also warned of a looming “data wall” by the end of this decade, beyond which model capabilities could hit a plateau unless they were fed fresh, reliable information.
Top American labs are spending lavishly to mine offline human knowledge, igniting a fierce ethical debate in the process.