China Aims to Dominate AI Training Data by 2028 to Shape Global Narratives
China is undertaking a state-sponsored, well-funded campaign to become the leading global supplier of training data for artificial intelligence models by 2028. The initiative involves creating high-quality datasets across more than a dozen strategic fields, recruiting experts to label data, launching academic courses, and distributing open-source data aligned with Communist Party values. This effort aims to counter Western narratives, particularly those led by the United States, which dominate AI models trained primarily on English-language data and reflect Western worldviews on issues like human rights and Taiwan.
Experts warn that AI chatbots, such as ChatGPT, can be influenced by the volume and prominence of online content supporting specific narratives. Cybersecurity specialist Ohad Cohen explains that while AI systems use ranking and filtering mechanisms, the lack of transparency about source selection means it is difficult to know who influences the information AI models rely on. Research published in Nature highlights how governments indirectly shape large language models by controlling local information environments, a phenomenon termed "institutional influence."
In a related development, Israel has launched a $100,000 influence campaign in the U.S. to feed positive information about the IDF and the Gaza conflict into AI language models. This campaign involved creating a website with anonymous Q&A-style articles designed to be absorbed by AI systems, with content cited by ChatGPT and Perplexity. OpenAI has also identified and blocked covert influence operations from Russia, China, Iran, and Israel using its AI tools to sway global public opinion.
Studies show that political questions posed in Chinese to AI models receive responses more favorable to the Chinese regime compared to those asked in English. Additionally, AI models tend to avoid criticizing authoritarian leaders due to safety concerns, as revealed by Meta's oversight council research. With Israel's upcoming elections, experts express concern that AI chatbots could become powerful tools for influence, potentially affecting voter turnout and political opinions in subtle, personalized ways that are difficult to detect in real time.
Dr. Tehila Schwartz Altshuler of the Israeli Democracy Institute notes that Israel's high chatbot penetration combined with moderate digital literacy creates a fertile ground for AI-driven influence campaigns. Unlike social media, chatbot responses are individualized, complicating efforts to identify coordinated manipulation during elections. This underscores the growing challenge of ensuring transparency and integrity in AI-generated information amid geopolitical and domestic political pressures.