{"slug": "state-media-control-influences-large-language-models", "title": "State media control influences large language models", "summary": "A study published in Nature found that government control of media influences large language models (LLMs) through their training data, with models showing stronger pro-government responses in languages of countries with lower media freedom. The study, based on six studies including a case study on China, demonstrated that additional pretraining on Chinese state-coordinated media generated more positive answers about Chinese political institutions and leaders, and that prompting models in Chinese produced more positive responses than in English.", "body_md": "## Abstract\n\nMillions of people around the world query large language models (LLMs) for information. Although several studies have compellingly documented the persuasive potential of these models 1,2,3,4,5,6,7,8,9,10, there is limited evidence of who or what influences the models themselves, leading to a flurry of concerns about which companies and governments build and regulate the models. Here we show through six studies that government control of the media across the world already influences the output of LLMs via their training data. We use a cross-national audit to show that LLMs exhibit a stronger pro-government valence in the languages of countries with lower media freedom than in those with higher media freedom. This result is correlational, so to triangulate the specific mechanism of how state media control can influence LLMs, we develop a multi-part case study on China’s media. We demonstrate that media scripted and curated by the Chinese state appears in LLM training datasets. To evaluate the plausible effect of this inclusion, we use an open-weight model to show that additional pretraining on Chinese state-coordinated media generates more positive answers to prompts about Chinese political institutions and leaders. We link this phenomenon to commercial models through two audit studies demonstrating that prompting models in Chinese generates more positive responses about China’s institutions and leaders than do the same queries in English. The combination of influence and persuasive potential across languages suggests the troubling conclusion that states and powerful institutions have increased strategic incentives to leverage media control in the hopes of shaping LLM output.\n\nThis is a preview of subscription content, [access via your institution](https://wayf.springernature.com?redirect_uri=https%3A%2F%2Fwww.nature.com%2Farticles%2Fs41586-026-10506-7)\n\n## Access options\n\nAccess Nature and 54 other Nature Portfolio journals\n\nGet Nature+, our best-value online-access subscription\n\n27,99 € / 30 days\n\ncancel any time\n\nSubscribe to this journal\n\nReceive 52 print issues and online access\n\n199,00 € per year\n\nonly 3,83 € per issue\n\nBuy this article\n\n- Purchase on SpringerLink\n- Instant access to the full article PDF.\n\n39,95 €\n\nPrices may be subject to local taxes which are calculated during checkout\n\n### Similar content being viewed by others\n\n### Subjects\n\n## Data availability\n\nDerivative data products are available in our replication archive ([https://doi.org/10.7910/DVN/NECR2K](https://doi.org/10.7910/DVN/NECR2K)). We released transformed products only rather than the full text of raw news stories because we do not hold their copyright. Our full-text articles were collected through a combination of news website scraping and data purchases from WisersOne (formerly WiseNews). We have provided additional replications of the studies using the latest models at the time of publication ([https://state-media-influence-llm.github.io/)](https://state-media-influence-llm.github.io/).\n\n## Code availability\n\nThe replication code for all analyses in the main text and extended data is available in our replication archive ([https://doi.org/10.7910/DVN/NECR2K](https://doi.org/10.7910/DVN/NECR2K)).\n\n## References\n\nPalmer, A. & Spirling, A. Large language models can argue in convincing ways about politics, but humans dislike AI authors: implications for governance.\n\n*Polit. Sci.***75**, 281–291 (2023).Bai, H. et al. LLM-generated messages can persuade humans on policy issues.\n\n*Nat. Commun.***16**, 6037 (2025).Hackenburg, K. & Margetts, H. Evaluating the persuasive influence of political microtargeting with large language models.\n\n*Proc. Natl Acad. Sci. USA***121**, e2403116121 (2024).Salvi, F. et al. On the conversational persuasiveness of GPT-4.\n\n*Nat. Hum. Behav.***9**, 1645–1653 (2025).Costello, T. H., Pennycook, G. & Rand, D. G. Durably reducing conspiracy beliefs through dialogues with AI.\n\n*Science***385**, eadq1814 (2024).Carrasco-Farre, C. Large language models are as persuasive as humans, but how? About the cognitive effort and moral-emotional language of LLM arguments. Preprint at\n\n[https://arxiv.org/abs/2404.09329](https://arxiv.org/abs/2404.09329)(2024).Tessler, M. H. et al. AI can help humans find common ground in democratic deliberation.\n\n*Science***386**, eadq2852 (2024).Goldstein, J. A. et al. How persuasive is AI-generated propaganda?\n\n*PNAS Nexus***3**, pgae034 (2024).Fisher, J. et al. Biased LLMs can influence political decision-making. In\n\n*Proc. 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)*(eds Che, W. et al.) 6559–6607 (Association for Computational Linguistics, 2025).Saenger, T.R. et al. AutoPersuade: a framework for evaluating and explaining persuasive arguments. In\n\n*Proc. 2024 Conference on Empirical Methods in Natural Language Processing*(eds Al-Onaizan, Y., Bansal, M. & Chen, Y.-N.) 16325–16342 (Association for Computational Linguistics, 2024).Islas-Carmona, J. O., Gutiérrez-Cortés, F. I. & Arribas-Urrutia, A. Disinformation and political propaganda: an exploration of the risks of artificial intelligence.\n\n*Explor. Media Ecol.***23**, 105–120 (2024).Woolley, S.\n\n*Manufacturing Consensus: Understanding Propaganda in the Era of Automation and Anonymity*(Yale Univ. Press, 2023).Broockman, D. & Kalla, J. Durably reducing transphobia: a field experiment on door-to-door canvassing.\n\n*Science***352**, 220–224 (2016).Roghanizad, M. M. & Bohns, V. K. Ask in person: you’re less persuasive than you think over email.\n\n*J. Exp. Soc. Psychol.***69**, 223–226 (2017).Buyl, M. et al. Large language models reflect the ideology of their creators.\n\n*npj Artif. Intell.***2**, 7 (2026).Guey, W. et al. Mapping geopolitical bias in 11 large language models: a bilingual, dual-framing analysis of US-China tensions. Preprint at\n\n[https://arxiv.org/abs/2503.23688](https://arxiv.org/abs/2503.23688)(2025).McCarthy, S. DeepSeek is giving the world a window into Chinese censorship and information control.\n\n*CNN*[https://edition.cnn.com/2025/01/29/china/deepseek-ai-china-censorship-moderation-intl-hnk](https://edition.cnn.com/2025/01/29/china/deepseek-ai-china-censorship-moderation-intl-hnk)(29 January 2025).Ouyang, Y., Nellis, S. and Tong, Q. DeepSeek hit by cyberattack as users flock to Chinese AI startup.\n\n*Reuters*[https://www.reuters.com/technology/artificial-intelligence/chinese-ai-startup-deepseek-overtakes-chatgpt-apple-app-store-2025-01-27/](https://www.reuters.com/technology/artificial-intelligence/chinese-ai-startup-deepseek-overtakes-chatgpt-apple-app-store-2025-01-27/)(27 January 2025).Kachwala, Z. Musk’s xAI updates Grok chatbot after ‘white genocide’ comments.\n\n*Reuters*[https://www.reuters.com/business/musks-xai-updates-grok-chatbot-after-white-genocide-comments-2025-05-17/](https://www.reuters.com/business/musks-xai-updates-grok-chatbot-after-white-genocide-comments-2025-05-17/)(17 May 2025).O’Brien, M. Google says its AI image-generator would sometimes ‘overcompensate’ for diversity.\n\n*Associated Press*[https://apnews.com/article/google-gemini-ai-chatbot-imagegenerator-race-c7e14de837aa65dd84f6e7ed6cfc4f4b](https://apnews.com/article/google-gemini-ai-chatbot-imagegenerator-race-c7e14de837aa65dd84f6e7ed6cfc4f4b)(23 February 2024).O’Brien, M. Elon Musk’s AI company says Grok chatbot focus on South Africa’s racial politics was ‘unauthorized’.\n\n*Associated Press*[https://apnews.com/article/grok-ai-south-africa-64ce5f240061ca0b88d5af4c424e1f3b](https://apnews.com/article/grok-ai-south-africa-64ce5f240061ca0b88d5af4c424e1f3b)(16 May 2025).Price, M. E.\n\n*Media and Sovereignty: The Global Information Revolution and its Challenge to State Power*(MIT Press, 2002).Hallin, D. C. & Mancini, P.\n\n*Comparing Media Systems: Three Models of Media and Politics*(Cambridge Univ. Press, 2004).Gururangan, S. et al. Don’t stop pretraining: adapt language models to domains and tasks. In\n\n*Proc. 58th Annual Meeting of the Association for Computational Linguistics*(eds Jurafsky, D. et al.) 8342–8360 (Association for Computational Linguistics, 2020).Bender, E. M. et al. On the dangers of stochastic parrots: can language models be too big? In\n\n*Proc. 2021 ACM Conference on Fairness, Accountability, and Transparency*610–623 (Association for Computing Machinery, 2021).Kreutzer, J. et al. Quality at a glance: an audit of web-crawled multilingual datasets.\n\n*Trans. Assoc. Comput. Linguist.***10**, 50–72 (2022).Blodgett, S. L. et al. Language (technology) is power: a critical survey of “bias” in NLP. In\n\n*Proc. Annual Meeting of the Association for Computational Linguistics*(eds Jurafsky, D. et al.) 5454–5476 (Association for Computational Linguistics, 2020).Ouyang, L. et al. Training language models to follow instructions with human feedback.\n\n*Adv. Neural Inf. Process. Syst.***35**, 27730–27744 (2022).Bai, Y. et al. Constitutional AI: harmlessness from AI feedback. Preprint at\n\n[https://arxiv.org/abs/2212.08073](https://arxiv.org/abs/2212.08073)(2022).Bulté, B. & Terryn, A. R. LLMs and cultural values: the impact of prompt language and explicit cultural framing.\n\n*Comput. Linguist.*[https://doi.org/10.1162/COLI.a.583](https://doi.org/10.1162/COLI.a.583)(2026).Lu, J. G., Song, L. L. & Zhang, L. D. Cultural tendencies in generative AI.\n\n*Nat. Hum. Behav.***9**, 2360–2369 (2025).Kay, M., Matuszek, C. & Munson, S. A. Unequal representation and gender stereotypes in image search results for occupations. In\n\n*Proc. 33rd Annual ACM Conference on Human Factors in Computing Systems*(eds Begole, B. et al.) 3819–3828 (Association for Computing Machinery, 2015).Noble, S. U.\n\n*Algorithms of Oppression: How Search Engines Reinforce Racism*(New York Univ. Press, 2018).Broussard, M.\n\n*More than a Glitch: Confronting Race, Gender, and Ability Bias in Tech*(MIT Press, 2023).Benjamin, R.\n\n*Race after Technology: Abolitionist Tools for the New Jim Code*(John Wiley & Sons, 2019).Buolamwini, J. & Gebru, T. Gender shades: intersectional accuracy disparities in commercial gender classification. In\n\n*Conference on Fairness, Accountability and Transparency*(eds Friedler, S. A. & Wilson, C.) 77–91 (PMLR, 2018).Barocas, S. & Selbst, A. D. Big data’s disparate impact.\n\n*Calif. L. Rev.***104**, 671 (2016).Sheng, E. et al. The woman worked as a babysitter: on biases in language generation. In\n\n*Proc. 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing*(*EMNLP-IJCNLP*) (eds Inui, K. et al.) 3407–3412 (Association for Computational Linguistics, 2019).Field, A. et al. A survey of race, racism, and anti-racism in NLP. In\n\n*Proc. 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing*(*Volume 1: Long Papers*) (eds Zong, C. et al.) 1905–1925 (Association for Computational Linguistics, 2021).Metaxa, D. et al. An image of society: gender and racial representation and impact in image search results for occupations.\n\n*Proc. ACM Hum. Comp. Interact.***5**, 1–23 (2021).Kotek, H., Dockum, R. & Sun, D. Gender bias and stereotypes in large language models. In\n\n*Proc. ACM Collective Intelligence Conference*(eds Bernstein, M. S. et al.) 12–24 (Association for Computing Machinery, 2023).Omiye, J. A. et al. Large language models propagate race-based medicine.\n\n*NPJ Digit. Med.***6**, 195 (2023).Jowett, G. S. & O’Donnell, V.\n\n*Propaganda & Persuasion*(Sage, 2018).Peisakhin, L. & Rozenas, A. Electoral effects of biased media: Russian television in Ukraine.\n\n*Am. J. Polit. Sci.***62**, 535–550 (2018).Selb, P. & Munzert, S. Examining a most likely case for strong campaign effects: Hitler’s speeches and the rise of the Nazi party, 1927-1933.\n\n*Am. Polit. Sci. Rev.***112**, 1050–1066 (2018).Rozenas, A. & Stukal, D. How autocrats manipulate economic news: evidence from Russia’s state-controlled television.\n\n*J. Polit.***81**, 982–996 (2019).Huang, H. Propaganda as signaling.\n\n*Comp. Polit.***47**, 419–444 (2015).Voigtländer, N. & Voth, H.-J. Nazi indoctrination and anti-Semitic beliefs in Germany.\n\n*Proc. Natl Acad. Sci. USA***112**, 7931–7936 (2015).King, G., Pan, J. & Roberts, M. E. How the Chinese government fabricates social media posts for strategic distraction, not engaged argument.\n\n*Am. Polit. Sci. Rev.***111**, 484–501 (2017).Stukal, D. et al. Why botter: how pro-government bots fight opposition in Russia.\n\n*Am. Polit. Sci. Rev.***116**, 843–857 (2022).Farzam, A. et al. Opinion manipulation on Farsi Twitter.\n\n*Sci. Rep.***13**, 333 (2023).Waight, H. et al. The decade-long growth of government-authored news media in China under Xi Jinping.\n\n*Proc. Natl Acad. Sci. USA***122**, e2408260122 (2025).Yang, E. & Roberts, M. E. Censorship of online encyclopedias: implications for NLP models. In\n\n*Proc. 2021 ACM Conference on Fairness, Accountability, and Transparency*537–548 (Association for Computing Machinery, 2021).Zhou, D. & Zhang, Y. Political biases and inconsistencies in bilingual GPT models — the cases of the US and China.\n\n*Sci. Rep.***14**, 25048 (2024).Ahmed, M. & Knockel, J. Extended abstract: the impact of online censorship on LLMs.\n\n*Free and Open Communications on the Internet*[https://www.petsymposium.org/foci/2024/foci-2024-0006.pdf](https://www.petsymposium.org/foci/2024/foci-2024-0006.pdf)(2024).Urman, A. & Makhortykh, M. The silence of the LLMs: cross-lingual analysis of guardrail-related political bias and false information prevalence in ChatGPT, Google Bard (Gemini), and Bing Chat.\n\n*Telemat. Informat.***96**, 102211 (2025).Spirling, A. & Stewart, B. M. What good is a regression? Inference to the best explanation and the practice of political science research.\n\n*J. Polit.***87**, 1587–1599 (2025).Reporters Without Borders.\n\n*World Press Freedom Index*[https://rsf.org/en/index](https://rsf.org/en/index)(2024).Shambaugh, D. in\n\n*Critical Readings on the Communist Party of China (4 Vols. Set)*(ed. Brødsgaard, K. E.) 713–751 (Brill, 2017).Brady, A.M.\n\n*Marketing Dictatorship: Propaganda and Thought Work in Contemporary China*(Rowman & Littlefield, 2009).Stockmann, D.\n\n*Media Commercialization and Authoritarian Rule in China*(Cambridge Univ. Press, 2013).Liang, F., Chen, Y. & Zhao, F. The platformization of propaganda: how Xuexi Qiangguo expands persuasion and assesses citizens in China.\n\n*Int. J. Commun.***15**, 20 (2021).Lu, Y. & Pan, J. Capturing clicks: how the Chinese government uses clickbait to compete for visibility.\n\n*Polit. Commun.***38**, 23–54 (2021).Repnikova, M. & Fang, K. Digital media experiments in China: ‘revolutionizing’ persuasion under Xi Jinping.\n\n*China Q.***239**, 679–701 (2019).Esarey, A. Winning hearts and minds? Cadres as microbloggers in China.\n\n*J. Curr. Chinese Aff.***44**, 69–103 (2015).Qin, B., Strömberg, D. & Wu, Y. Media bias in China.\n\n*Am. Econom. Rev.***108**, 2442–2476 (2018).Pan, J., Shao, Z. & Xu, Y. How government-controlled media shifts policy attitudes through framing.\n\n*Polit. Sci. Res. Methods***10**, 317–332 (2022).Zhang, Z. et al. Unveiling linguistic regions in large language models. In\n\n*Proc. 62nd Annual Meeting of the Association for Computational Linguistics*(*Volume 1: Long Papers*) (eds Ku, L.-W., Martins, A. & Srikumar, V.) 6228–6247 (Association for Computational Linguistics, 2024).Qi, J., Fernández, R. & Bisazza, A. Cross-lingual consistency of factual knowledge in multilingual language models. In\n\n*Proc. 2023 Conference on Empirical Methods in Natural Language Processing*(eds Bouamor, H., Pino, J. & Bali, K.) 10650–10666 (Association for Computational Linguistics, 2023).Li, B., Haider, S. & Callison-Burch, C. This land is your, my land: evaluating geopolitical bias in language models through territorial disputes. In\n\n*Proc. 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies*(*Volume 1: Long Papers*) (eds Duh, K., Gomez, H. & Bethard, S.) 3855–3871 (Association for Computational Linguistics, 2024).Wendler, C. et al. Do llamas work in English? On the latent language of multilingual transformers. In\n\n*Proc. 62nd Annual Meeting of the Association for Computational Linguistics*(*Volume 1: Long Papers*) (eds Ku, L.-W., Martins, A. & Srikumar, V.) 15366–15394 (Association for Computational Linguistics, 2024).Durmus, E. et al. Towards measuring the representation of subjective global opinions in language models. In\n\n*1st Conference on Language Modeling*[https://openreview.net/pdf?id=zl16jLb91v](https://openreview.net/pdf?id=zl16jLb91v)(COLM, 2024).Shayegani, E. et al. Survey of vulnerabilities in large language models revealed by adversarial attacks. Preprint at\n\n[https://arxiv.org/abs/2310.10844](https://arxiv.org/abs/2310.10844)(2023).Roberts, M.\n\n*Censored: Distraction and Diversion Inside China’s Great Firewall*(Princeton Univ. Press, 2018).Ishihara, S. & Takahashi, H. Quantifying memorization and detecting training data of pre-trained language models using Japanese newspaper. In\n\n*Proc. 17th International Natural Language Generation Conference*(eds Mahamood, S., Le Minh, N. & Ippolito, D.) 165–179 (Association for Computational Linguistics, 2024).Fulay, S. et al. On the relationship between truth and political bias in language models. In\n\n*Proc. 2024 Conference on Empirical Methods in Natural Language Processing*(eds Al-Onaizan, Y., Bansal, M. & Chen, Y.-N.) 9004–9018 (Association for Computational Linguistics, 2024).Nguyen, T. et al. CulturaX: a cleaned, enormous, and multilingual dataset for large language models in 167 languages. In\n\n*Proc. 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)*(eds Calzolari, N. et al.) 4226–4237 (ELRA and ICCL, 2024).Truex, R. Focal points, dissident calendars, and preemptive repression.\n\n*J. Confl. Resolut.***63**, 1032–1052 (2019).Carter, E. B. & Carter, B. L. When autocrats threaten citizens with violence: evidence from China.\n\n*Br. J. Polit. Sci.***52**, 671–696 (2022).Schlessinger, J. et al. Exposing the obscured influence of state-controlled media via causal inference of quotation propagation.\n\n*Sci. Rep.***15**, 1110 (2025).Zhao, W. et al. WildChat: 1M ChatGPT interaction logs in the wild. In\n\n*12th International Conference on Learning Representations*[https://openreview.net/forum?id=Bl8u7ZRlbM](https://openreview.net/forum?id=Bl8u7ZRlbM)(ICLR, 2024).Trussler, M. & Soroka, S. Consumer demand for cynical and negative news frames.\n\n*Int. J. Press Polit.***19**, 360–379 (2014).Arango-Kure, M., Garz, M. & Rott, A. Bad news sells: the demand for news magazines and the tone of their covers.\n\n*J. Media Econom.***27**, 199–214 (2014).Christin, A. Counting clicks: quantification and variation in web journalism in the United States and France.\n\n*Am. J. Sociol.***123**, 1382–1415 (2018).O’Neil, C.\n\n*Weapons of Math Destruction: How Big Data Increases Inequality and Threatens Democracy*(Crown, 2017).Fourcade, M. & Healy, K.\n\n*The Ordinal Society*(Harvard Univ. Press, 2024).Gillespie, T.\n\n*Custodians of the Internet: Platforms, Content Moderation, and the Hidden Decisions that Shape Social Media*(Yale Univ. Press, 2018).Yang, E. & Roberts, M. E. The authoritarian data problem.\n\n*J. Democr.***34**, 141–150 (2023).Wang, H. & Sparks, C. Chinese newspaper groups in the digital era: the resurgence of the party press.\n\n*J. Commun.***69**, 94–119 (2019).Raffel, C. et al. Exploring the limits of transfer learning with a unified text-to-text transformer.\n\n*J. Mach. Learn. Res.***21**, 1–67 (2020).Scheible, R. et al. GottBERT: a pure German language model. In\n\n*Proc. 2024 Conference on Empirical Methods in Natural Language Processing*(eds Al-Onaizan, Y., Bansal, M. & Chen, Y.-N.) 21237–21250 (Association for Computational Linguistics, 2024).Shalumov, V. & Haskey, H. Hero: ROBERTa and longformer Hebrew language models. Preprint at\n\n[https://arxiv.org/abs/2304.11077](https://arxiv.org/abs/2304.11077)(2023).Serrano, A. V. et al. RigoBERTa: a state-of-the-art language model for Spanish. Preprint at\n\n[https://arxiv.org/abs/2205.10233](https://arxiv.org/abs/2205.10233)(2022).Shliazhko, O. et al. mGPT: few-shot learners go multilingual.\n\n*Trans. Assoc. Comput. Linguist.***12**, 58–79 (2024).Mandal, P. K. & Mahto, R. An FNet based auto encoder for long sequence news story generation. Preprint at\n\n[https://arxiv.org/abs/2211.08295](https://arxiv.org/abs/2211.08295)(2022).Boumans, J. et al. The agency makes the (online) news world go round: the impact of news agency content on print and online news.\n\n*Int. J. Commun.***12**, 22 (2018).Cagé, J., Hervé, N. & Viaud, M.-L. The production of information in an online world.\n\n*Rev. Econom. Stud.***87**, 2126–2164 (2020).Nicholls, T. Detecting textual reuse in news stories, at scale.\n\n*Int. J. Commun.***13**, 4173–4197 (2019).Gao, L. et al. The pile: an 800gb dataset of diverse text for language modeling. Preprint at\n\n[https://arxiv.org/abs/2101.00027](https://arxiv.org/abs/2101.00027)(2020).Carlini, N. et al. Quantifying memorization across neural language models. In\n\n*11th**International Conference on Learning Representations*[https://openreview.net/forum?id=TatRHT_1cK](https://openreview.net/forum?id=TatRHT_1cK)(ICLR, 2023).Touvron, H. et al. Llama 2: open foundation and fine-tuned chat models. Preprint at\n\n[https://arxiv.org/abs/2307.09288](https://arxiv.org/abs/2307.09288)(2023).Chen, L. et al. AlpaGasus: training a better alpaca with fewer data. In\n\n*International Conference on Learning Representations*(eds Kim, B. et al.)[https://proceedings.iclr.cc/paper_files/paper/2024/hash/9543942c237ded1b39b1fd37259ff88e-Abstract-Conference.html](https://proceedings.iclr.cc/paper_files/paper/2024/hash/9543942c237ded1b39b1fd37259ff88e-Abstract-Conference.html)(ICLR, 2024).Hu, E. J. et al. LoRA: low-rank adaptation of large language models. In\n\n*10th International Conference on Learning Representations*[https://openreview.net/forum?id=nZeVKeeFYf9](https://openreview.net/forum?id=nZeVKeeFYf9)(ICLR, 2022).Kirkpatrick, J. et al. Overcoming catastrophic forgetting in neural networks.\n\n*Proc. Natl Acad. Sci. USA***114**, 3521–3526 (2017).Leybzon, D. D. & Kervadec, C. Learning, forgetting, remembering: insights from tracking LLM memorization during training. In\n\n*Proc. 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP*(eds Belinkov, Y. et al.) 43–57 (Association for Computational Linguistics, 2024).Mahomed, Y. et al. Auditing GPT’s content moderation guardrails: can ChatGPT write your favorite TV show? In\n\n*2024 ACM Conference on Fairness, Accountability, and Transparency*660–686 (Association for Computing Machinery, 2024).Xu, B. NLP Chinese Corpus: large scale Chinese Corpus for NLP (version 1.0).\n\n*Zenodo*[https://doi.org/10.5281/zenodo.3402023](https://doi.org/10.5281/zenodo.3402023)(2019).Egami, N. et al. Using imperfect surrogates for downstream inference: design-based supervised learning for social science applications of large language models.\n\n*Adv. Neural Inf. Process. Syst.***36**, 68589–68601 (2024).Eberhard, D. M., Simons, G. F. & Fennig, C. D. (eds).\n\n*Ethnologue: Languages of the World*(SIL International, 2024).\n\n## Acknowledgements\n\nThis project would not be possible without research assistance from A. Chen, X. Chi, Y. Feng, Y. Liu, W. Mei, L. Pothier, M. Sato, V. Tang, J. Xu, S. Zhong and other anonymous individuals. For feedback on the manuscript and various stages of the project, we acknowledge D. Baldassarri, A. Breuer, J. Grimmer, M. Hinck, D. Metaxa, É. Ollion, R. E. Robertson, C. Rudin, M. Salganik, S. Westwood, Y. Zhang, D. Zhou, attendees of our presentations at the Yale’s Generative AI and Social Science conference, Institut Polytechnique de Paris’ NLP and Social Sciences Seminar, Center for Information Networks and Democracy (CIND) Workshop at UPenn, the Ford Center for Global Citizenship Political Economy and AI Conference at Northwestern University, the Data Science Frontiers Conference at the NYU Abu-Dhabi Institute, the Social Media and Democratic Practice Conference at the Hoover Institution, ASA, APSA, IC2S2, University of Washington, University of Wisconsin, Madison, Stanford University, American University, University of Virginia, University of Texas-Austin, Johns Hopkins University Center for Language and Speech Processing, Bocconi University, European University Institute, and members of the StewartLab and NYU Center for Social Media and Politics. D. Johnson helped by illustrating Fig. [1](/articles/s41586-026-10506-7#Fig1). We received feedback during the peer review process from M. E. Sutherland and Y. Sweeney, as well as a set of anonymous peer reviewers. This work was supported by Princeton Research Computing, Princeton Data-Driven Social Science Initiative, Princeton Center for Statistics and Machine Learning, UCSD Social Sciences Computing Facility, the NYU Center for Social Media and Politics, UCSD’s 21st Century China Center and the Carnegie Corporation of New York. The Center for Social Media, AI, and Politics at New York University is supported by funding from the John S. and James L. Knight Foundation, the Charles Koch Foundation, Craig Newmark Philanthropies, the William and Flora Hewlett Foundation, and the Siegel Family Endowment. Funding was provided for the larger project of which this paper is a part by the Templeton World Charity Foundation. This work was also supported in part through the NYU IT High Performance Computing resources, services and staff expertise.\n\n## Author information\n\n### Authors and Affiliations\n\n### Contributions\n\nH.W. and E.Y. are co-first authors for this paper. H.W., E.Y., Y.Y., S.M., M.E.R., B.M.S. and J.A.T. jointly designed the studies. H.W., E.Y. and Y.Y. collected data, conducted all analyses and produced figures. B.M.S. wrote the paper. H.W. wrote the [Methods](/articles/s41586-026-10506-7#Sec9) section. H.W., E.Y. and Y.Y. wrote the [Supplementary Information](/articles/s41586-026-10506-7#MOESM1). H.W., E.Y., Y.Y., S.M., M.E.R., B.M.S. and J.A.T. collaboratively edited and developed the manuscript.\n\n### Corresponding author\n\n## Ethics declarations\n\n### Competing interests\n\nH.W. and S.M. have personal financial interests in AI-related companies, in particular Meta (H.W. only), Nvidia, Alphabet, Microsoft and Taiwan Semiconductor (S.M. only). Two authors have past employment histories with AI-related companies: E.Y. was an intern at Microsoft Research in the summer of 2022 and 2023; and S.M. worked at Facebook (now Meta) in various capacities from 2011 to 2015 and 2018 to 2020, at Twitter (now X) from 2021 to 2023, and contracts for 501c6 non-profit MLCommons, which releases AI benchmarks (2026 to present). After acceptance of this paper, S.M. accepted a job at Google DeepMind. Finally, four authors received funding or other resources for unrelated projects from AI-related companies: for an unrelated project, B.M.S. received an unrestricted grant from Meta, ‘Foundational Integrity Research: Misinformation and Polarization’; S.M. received a 2010 Google Research Award for a research project on ‘Social cues and reliability in content selection and evaluation’; E.Y. received a Google Research Award for an unrelated project in 2026; and J.A.T. received a small fee from Facebook to compensate him for administrative time spent in organizing a 1-day conference for approximately 30 academic researchers and a dozen Facebook product managers and data scientists that was held at NYU in the summer of 2017 to discuss research related to civic engagement. J.A.T. is also one of the co-leads of the external academic team for the 2020 US Facebook and Instagram Election Study, a project that began in early 2020 and is still ongoing at the time of the writing of this article; J.A.T. was not compensated financially for his participation in this project by Meta, but the project involves working collaboratively with Meta researchers. J.A.T. also received a 2024 Google Research Grant to support a research project on ‘From search engines to answer engines: testing the effects of traditional and LLM-based search on belief in the veracity of news’. For an unrelated project, J.A.T. was listed as a co-investigator on a ‘Foundational Integrity Research: Misinformation and Polarization’ grant application for an unrestricted grant from Meta that was awarded to a principal investigator at a different university; no research funds were ever transferred to J.A.T. as part of this grant. J.A.T. is a Senior Geopolitical Risk Advisor at Kroll. M.E.R. and Y.Y. declare no competing interests.\n\n## Peer review\n\n### Peer review information\n\n*Nature* thanks Staffan I. Lindberg and the other, anonymous, reviewer(s) for their contribution to the peer review of this work. [Peer reviewer reports](/articles/s41586-026-10506-7#MOESM3) are available.\n\n## Additional information\n\n**Publisher’s note** Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.\n\n## Extended data figures and tables\n\n[Extended Data Fig. 1 Percent of CulturaX Documents that Match State Controlled Media, by Source (Study 1).](/articles/s41586-026-10506-7/figures/6)\n\nThis figure examines the percent of Chinese-language CulturaX documents (*n* = 189, 486, 611) matched to each type of Chinese state-controlled media source: state-run media, scripted news, and *Xuexi Qiangguo* articles. State media includes articles from *Xinhua* News Agency and *Xinwen Lianbo* nightly broadcasts. As in Fig. [2a](/articles/s41586-026-10506-7#Fig2) we label a CulturaX document as “matched” if it has at least 0.2 5-word cosine similarity with a state-controlled media document. We observe the same patterns across all sources, although the match rate for state-run media documents is consistently higher.\n\n[Extended Data Fig. 2 Chinese-Language CulturaX Documents are More Likely to be Drawn from Chinese State Controlled Web Domains than Wikipedia Domains (Study 1).](/articles/s41586-026-10506-7/figures/7)\n\nThis plot shows the percent of Chinese-language CulturaX documents (*n* = 189, 486, 611) with URLs from different domains. We exclude Chinese language CulturaX documents for which we had missing or faulty URL data (all OSCAR-2019 and OSCAR-21.09 documents). Chinese language CulturaX documents are forty-one times more likely to be from a mainland Chinese government domain (gov.cn or chinacourt.org) than from Chinese language Wikipedia.\n\n[Extended Data Fig. 3 Schematic of Audit Design for Human Evaluation and LLM-as-Judge Experiments (Study 4 and 5).](/articles/s41586-026-10506-7/figures/8)\n\nWe prompted LLMs with a series of political prompts twice, once in English and once in Chinese. We then translated each pair of English and Chinese completions into the other language. Research assistants and LLM-as-judge evaluate the Chinese and English completions, displayed in a single language.\n\n[Extended Data Fig. 4 Debiasing LLM-as-Judge Estimates Does Not Change Results (Study 4).](/articles/s41586-026-10506-7/figures/9)\n\nPlot includes debiased coefficients of our model estimating whether Chinese completions are more favourable to the country subject of the prompt, with naive estimator (No DSL) as reference. Debiasing is done with design-based supervised learning (DSL) estimator. The debiasing is relative to gold standard RA labels (3 per comparison, majority vote). We oversampled gold standard labels on prompts about China, so we collapsed prompts about other countries into “Not China”. Error bars represent 95% confidence intervals.\n\n[Extended Data Fig. 5 Response favourability comparison between DeepSeek-R1 and GPT-4o demonstrates DeepSeek-R1 is more favourable in its completions to China than OpenAI’s GPT-4o model.](/articles/s41586-026-10506-7/figures/10)\n\nEach estimate is an average over LLM-as-judge scores, where 0 indicates GPT-4o’s completion is rated as more favourable and 1 indicates DeepSeek-R1’s completion is favourable. The line drawn at .5 indicates what we would expect if the LLM-as-judge were engaging in random guessing. Error bars represent 95% confidence intervals.\n\n[Extended Data Fig. 6 Study 6 Results Are Consistent with Different Reference Languages.](/articles/s41586-026-10506-7/figures/11)\n\nThis plot replicates the study 6 audit with three different reference languages (English, Spanish, and Chinese). Our results remain consistent across reference languages, with the exception of Sonnet when Chinese is used as the reference language. In this robustness check we include a random sample of 30% of our original prompts. Error bars represent 95% confidence intervals.\n\n## Rights and permissions\n\nSpringer Nature or its licensor (e.g. a society or other partner) holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law.\n\n## About this article\n\n### Cite this article\n\nWaight, H., Yang, E., Yuan, Y. *et al.* State media control influences large language models.\n*Nature* **655**, 685–693 (2026). https://doi.org/10.1038/s41586-026-10506-7\n\nReceived:\n\nAccepted:\n\nPublished:\n\nVersion of record:\n\nIssue date:\n\nDOI: https://doi.org/10.1038/s41586-026-10506-7", "url": "https://wpnews.pro/news/state-media-control-influences-large-language-models", "canonical_source": "https://www.nature.com/articles/s41586-026-10506-7", "published_at": "2026-08-15 07:54:21+00:00", "updated_at": "2026-08-15 08:11:17.077607+00:00", "lang": "en", "topics": ["large-language-models", "ai-policy", "ai-ethics"], "entities": ["Nature", "China", "WisersOne"], "alternates": {"html": "https://wpnews.pro/news/state-media-control-influences-large-language-models", "markdown": "https://wpnews.pro/news/state-media-control-influences-large-language-models.md", "text": "https://wpnews.pro/news/state-media-control-influences-large-language-models.txt", "jsonld": "https://wpnews.pro/news/state-media-control-influences-large-language-models.jsonld"}}