State media control influences large language models A study published in Nature found that government control of media influences large language models (LLMs) through their training data, with models showing stronger pro-government responses in languages of countries with lower media freedom. The study, based on six studies including a case study on China, demonstrated that additional pretraining on Chinese state-coordinated media generated more positive answers about Chinese political institutions and leaders, and that prompting models in Chinese produced more positive responses than in English. Abstract Millions of people around the world query large language models LLMs for information. Although several studies have compellingly documented the persuasive potential of these models 1,2,3,4,5,6,7,8,9,10, there is limited evidence of who or what influences the models themselves, leading to a flurry of concerns about which companies and governments build and regulate the models. Here we show through six studies that government control of the media across the world already influences the output of LLMs via their training data. We use a cross-national audit to show that LLMs exhibit a stronger pro-government valence in the languages of countries with lower media freedom than in those with higher media freedom. This result is correlational, so to triangulate the specific mechanism of how state media control can influence LLMs, we develop a multi-part case study on China’s media. We demonstrate that media scripted and curated by the Chinese state appears in LLM training datasets. To evaluate the plausible effect of this inclusion, we use an open-weight model to show that additional pretraining on Chinese state-coordinated media generates more positive answers to prompts about Chinese political institutions and leaders. We link this phenomenon to commercial models through two audit studies demonstrating that prompting models in Chinese generates more positive responses about China’s institutions and leaders than do the same queries in English. The combination of influence and persuasive potential across languages suggests the troubling conclusion that states and powerful institutions have increased strategic incentives to leverage media control in the hopes of shaping LLM output. This is a preview of subscription content, access via your institution https://wayf.springernature.com?redirect uri=https%3A%2F%2Fwww.nature.com%2Farticles%2Fs41586-026-10506-7 Access options Access Nature and 54 other Nature Portfolio journals Get Nature+, our best-value online-access subscription 27,99 € / 30 days cancel any time Subscribe to this journal Receive 52 print issues and online access 199,00 € per year only 3,83 € per issue Buy this article - Purchase on SpringerLink - Instant access to the full article PDF. 39,95 € Prices may be subject to local taxes which are calculated during checkout Similar content being viewed by others Subjects Data availability Derivative data products are available in our replication archive https://doi.org/10.7910/DVN/NECR2K https://doi.org/10.7910/DVN/NECR2K . We released transformed products only rather than the full text of raw news stories because we do not hold their copyright. Our full-text articles were collected through a combination of news website scraping and data purchases from WisersOne formerly WiseNews . We have provided additional replications of the studies using the latest models at the time of publication https://state-media-influence-llm.github.io/ https://state-media-influence-llm.github.io/ . Code availability The replication code for all analyses in the main text and extended data is available in our replication archive https://doi.org/10.7910/DVN/NECR2K https://doi.org/10.7910/DVN/NECR2K . References Palmer, A. & Spirling, A. Large language models can argue in convincing ways about politics, but humans dislike AI authors: implications for governance. Polit. Sci. 75 , 281–291 2023 .Bai, H. et al. LLM-generated messages can persuade humans on policy issues. Nat. Commun. 16 , 6037 2025 .Hackenburg, K. & Margetts, H. Evaluating the persuasive influence of political microtargeting with large language models. Proc. Natl Acad. Sci. USA 121 , e2403116121 2024 .Salvi, F. et al. On the conversational persuasiveness of GPT-4. Nat. Hum. Behav. 9 , 1645–1653 2025 .Costello, T. H., Pennycook, G. & Rand, D. G. Durably reducing conspiracy beliefs through dialogues with AI. Science 385 , eadq1814 2024 .Carrasco-Farre, C. Large language models are as persuasive as humans, but how? About the cognitive effort and moral-emotional language of LLM arguments. Preprint at https://arxiv.org/abs/2404.09329 https://arxiv.org/abs/2404.09329 2024 .Tessler, M. H. et al. AI can help humans find common ground in democratic deliberation. Science 386 , eadq2852 2024 .Goldstein, J. A. et al. How persuasive is AI-generated propaganda? PNAS Nexus 3 , pgae034 2024 .Fisher, J. et al. Biased LLMs can influence political decision-making. In Proc. 63rd Annual Meeting of the Association for Computational Linguistics Volume 1: Long Papers eds Che, W. et al. 6559–6607 Association for Computational Linguistics, 2025 .Saenger, T.R. et al. AutoPersuade: a framework for evaluating and explaining persuasive arguments. In Proc. 2024 Conference on Empirical Methods in Natural Language Processing eds Al-Onaizan, Y., Bansal, M. & Chen, Y.-N. 16325–16342 Association for Computational Linguistics, 2024 .Islas-Carmona, J. O., Gutiérrez-Cortés, F. I. & Arribas-Urrutia, A. Disinformation and political propaganda: an exploration of the risks of artificial intelligence. Explor. Media Ecol. 23 , 105–120 2024 .Woolley, S. Manufacturing Consensus: Understanding Propaganda in the Era of Automation and Anonymity Yale Univ. Press, 2023 .Broockman, D. & Kalla, J. Durably reducing transphobia: a field experiment on door-to-door canvassing. Science 352 , 220–224 2016 .Roghanizad, M. M. & Bohns, V. K. Ask in person: you’re less persuasive than you think over email. J. Exp. Soc. Psychol. 69 , 223–226 2017 .Buyl, M. et al. Large language models reflect the ideology of their creators. npj Artif. Intell. 2 , 7 2026 .Guey, W. et al. Mapping geopolitical bias in 11 large language models: a bilingual, dual-framing analysis of US-China tensions. Preprint at https://arxiv.org/abs/2503.23688 https://arxiv.org/abs/2503.23688 2025 .McCarthy, S. DeepSeek is giving the world a window into Chinese censorship and information control. CNN https://edition.cnn.com/2025/01/29/china/deepseek-ai-china-censorship-moderation-intl-hnk https://edition.cnn.com/2025/01/29/china/deepseek-ai-china-censorship-moderation-intl-hnk 29 January 2025 .Ouyang, Y., Nellis, S. and Tong, Q. DeepSeek hit by cyberattack as users flock to Chinese AI startup. Reuters https://www.reuters.com/technology/artificial-intelligence/chinese-ai-startup-deepseek-overtakes-chatgpt-apple-app-store-2025-01-27/ https://www.reuters.com/technology/artificial-intelligence/chinese-ai-startup-deepseek-overtakes-chatgpt-apple-app-store-2025-01-27/ 27 January 2025 .Kachwala, Z. Musk’s xAI updates Grok chatbot after ‘white genocide’ comments. Reuters https://www.reuters.com/business/musks-xai-updates-grok-chatbot-after-white-genocide-comments-2025-05-17/ https://www.reuters.com/business/musks-xai-updates-grok-chatbot-after-white-genocide-comments-2025-05-17/ 17 May 2025 .O’Brien, M. Google says its AI image-generator would sometimes ‘overcompensate’ for diversity. Associated Press https://apnews.com/article/google-gemini-ai-chatbot-imagegenerator-race-c7e14de837aa65dd84f6e7ed6cfc4f4b https://apnews.com/article/google-gemini-ai-chatbot-imagegenerator-race-c7e14de837aa65dd84f6e7ed6cfc4f4b 23 February 2024 .O’Brien, M. Elon Musk’s AI company says Grok chatbot focus on South Africa’s racial politics was ‘unauthorized’. Associated Press https://apnews.com/article/grok-ai-south-africa-64ce5f240061ca0b88d5af4c424e1f3b https://apnews.com/article/grok-ai-south-africa-64ce5f240061ca0b88d5af4c424e1f3b 16 May 2025 .Price, M. E. Media and Sovereignty: The Global Information Revolution and its Challenge to State Power MIT Press, 2002 .Hallin, D. C. & Mancini, P. Comparing Media Systems: Three Models of Media and Politics Cambridge Univ. Press, 2004 .Gururangan, S. et al. Don’t stop pretraining: adapt language models to domains and tasks. In Proc. 58th Annual Meeting of the Association for Computational Linguistics eds Jurafsky, D. et al. 8342–8360 Association for Computational Linguistics, 2020 .Bender, E. M. et al. On the dangers of stochastic parrots: can language models be too big? In Proc. 2021 ACM Conference on Fairness, Accountability, and Transparency 610–623 Association for Computing Machinery, 2021 .Kreutzer, J. et al. Quality at a glance: an audit of web-crawled multilingual datasets. Trans. Assoc. Comput. Linguist. 10 , 50–72 2022 .Blodgett, S. L. et al. Language technology is power: a critical survey of “bias” in NLP. In Proc. Annual Meeting of the Association for Computational Linguistics eds Jurafsky, D. et al. 5454–5476 Association for Computational Linguistics, 2020 .Ouyang, L. et al. Training language models to follow instructions with human feedback. Adv. Neural Inf. Process. Syst. 35 , 27730–27744 2022 .Bai, Y. et al. Constitutional AI: harmlessness from AI feedback. Preprint at https://arxiv.org/abs/2212.08073 https://arxiv.org/abs/2212.08073 2022 .Bulté, B. & Terryn, A. R. LLMs and cultural values: the impact of prompt language and explicit cultural framing. Comput. Linguist. https://doi.org/10.1162/COLI.a.583 https://doi.org/10.1162/COLI.a.583 2026 .Lu, J. G., Song, L. L. & Zhang, L. D. Cultural tendencies in generative AI. Nat. Hum. Behav. 9 , 2360–2369 2025 .Kay, M., Matuszek, C. & Munson, S. A. Unequal representation and gender stereotypes in image search results for occupations. In Proc. 33rd Annual ACM Conference on Human Factors in Computing Systems eds Begole, B. et al. 3819–3828 Association for Computing Machinery, 2015 .Noble, S. U. Algorithms of Oppression: How Search Engines Reinforce Racism New York Univ. Press, 2018 .Broussard, M. More than a Glitch: Confronting Race, Gender, and Ability Bias in Tech MIT Press, 2023 .Benjamin, R. Race after Technology: Abolitionist Tools for the New Jim Code John Wiley & Sons, 2019 .Buolamwini, J. & Gebru, T. Gender shades: intersectional accuracy disparities in commercial gender classification. In Conference on Fairness, Accountability and Transparency eds Friedler, S. A. & Wilson, C. 77–91 PMLR, 2018 .Barocas, S. & Selbst, A. D. Big data’s disparate impact. Calif. L. Rev. 104 , 671 2016 .Sheng, E. et al. The woman worked as a babysitter: on biases in language generation. In Proc. 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing EMNLP-IJCNLP eds Inui, K. et al. 3407–3412 Association for Computational Linguistics, 2019 .Field, A. et al. A survey of race, racism, and anti-racism in NLP. In Proc. 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing Volume 1: Long Papers eds Zong, C. et al. 1905–1925 Association for Computational Linguistics, 2021 .Metaxa, D. et al. An image of society: gender and racial representation and impact in image search results for occupations. Proc. ACM Hum. Comp. Interact. 5 , 1–23 2021 .Kotek, H., Dockum, R. & Sun, D. Gender bias and stereotypes in large language models. In Proc. ACM Collective Intelligence Conference eds Bernstein, M. S. et al. 12–24 Association for Computing Machinery, 2023 .Omiye, J. A. et al. Large language models propagate race-based medicine. NPJ Digit. Med. 6 , 195 2023 .Jowett, G. S. & O’Donnell, V. Propaganda & Persuasion Sage, 2018 .Peisakhin, L. & Rozenas, A. Electoral effects of biased media: Russian television in Ukraine. Am. J. Polit. Sci. 62 , 535–550 2018 .Selb, P. & Munzert, S. Examining a most likely case for strong campaign effects: Hitler’s speeches and the rise of the Nazi party, 1927-1933. Am. Polit. Sci. Rev. 112 , 1050–1066 2018 .Rozenas, A. & Stukal, D. How autocrats manipulate economic news: evidence from Russia’s state-controlled television. J. Polit. 81 , 982–996 2019 .Huang, H. Propaganda as signaling. Comp. Polit. 47 , 419–444 2015 .Voigtländer, N. & Voth, H.-J. Nazi indoctrination and anti-Semitic beliefs in Germany. Proc. Natl Acad. Sci. USA 112 , 7931–7936 2015 .King, G., Pan, J. & Roberts, M. E. How the Chinese government fabricates social media posts for strategic distraction, not engaged argument. Am. Polit. Sci. Rev. 111 , 484–501 2017 .Stukal, D. et al. Why botter: how pro-government bots fight opposition in Russia. Am. Polit. Sci. Rev. 116 , 843–857 2022 .Farzam, A. et al. Opinion manipulation on Farsi Twitter. Sci. Rep. 13 , 333 2023 .Waight, H. et al. The decade-long growth of government-authored news media in China under Xi Jinping. Proc. Natl Acad. Sci. USA 122 , e2408260122 2025 .Yang, E. & Roberts, M. E. Censorship of online encyclopedias: implications for NLP models. In Proc. 2021 ACM Conference on Fairness, Accountability, and Transparency 537–548 Association for Computing Machinery, 2021 .Zhou, D. & Zhang, Y. Political biases and inconsistencies in bilingual GPT models — the cases of the US and China. Sci. Rep. 14 , 25048 2024 .Ahmed, M. & Knockel, J. Extended abstract: the impact of online censorship on LLMs. Free and Open Communications on the Internet https://www.petsymposium.org/foci/2024/foci-2024-0006.pdf https://www.petsymposium.org/foci/2024/foci-2024-0006.pdf 2024 .Urman, A. & Makhortykh, M. The silence of the LLMs: cross-lingual analysis of guardrail-related political bias and false information prevalence in ChatGPT, Google Bard Gemini , and Bing Chat. Telemat. Informat. 96 , 102211 2025 .Spirling, A. & Stewart, B. M. What good is a regression? Inference to the best explanation and the practice of political science research. J. Polit. 87 , 1587–1599 2025 .Reporters Without Borders. World Press Freedom Index https://rsf.org/en/index https://rsf.org/en/index 2024 .Shambaugh, D. in Critical Readings on the Communist Party of China 4 Vols. Set ed. Brødsgaard, K. E. 713–751 Brill, 2017 .Brady, A.M. Marketing Dictatorship: Propaganda and Thought Work in Contemporary China Rowman & Littlefield, 2009 .Stockmann, D. Media Commercialization and Authoritarian Rule in China Cambridge Univ. Press, 2013 .Liang, F., Chen, Y. & Zhao, F. The platformization of propaganda: how Xuexi Qiangguo expands persuasion and assesses citizens in China. Int. J. Commun. 15 , 20 2021 .Lu, Y. & Pan, J. Capturing clicks: how the Chinese government uses clickbait to compete for visibility. Polit. Commun. 38 , 23–54 2021 .Repnikova, M. & Fang, K. Digital media experiments in China: ‘revolutionizing’ persuasion under Xi Jinping. China Q. 239 , 679–701 2019 .Esarey, A. Winning hearts and minds? Cadres as microbloggers in China. J. Curr. Chinese Aff. 44 , 69–103 2015 .Qin, B., Strömberg, D. & Wu, Y. Media bias in China. Am. Econom. Rev. 108 , 2442–2476 2018 .Pan, J., Shao, Z. & Xu, Y. How government-controlled media shifts policy attitudes through framing. Polit. Sci. Res. Methods 10 , 317–332 2022 .Zhang, Z. et al. Unveiling linguistic regions in large language models. In Proc. 62nd Annual Meeting of the Association for Computational Linguistics Volume 1: Long Papers eds Ku, L.-W., Martins, A. & Srikumar, V. 6228–6247 Association for Computational Linguistics, 2024 .Qi, J., Fernández, R. & Bisazza, A. Cross-lingual consistency of factual knowledge in multilingual language models. In Proc. 2023 Conference on Empirical Methods in Natural Language Processing eds Bouamor, H., Pino, J. & Bali, K. 10650–10666 Association for Computational Linguistics, 2023 .Li, B., Haider, S. & Callison-Burch, C. This land is your, my land: evaluating geopolitical bias in language models through territorial disputes. In Proc. 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies Volume 1: Long Papers eds Duh, K., Gomez, H. & Bethard, S. 3855–3871 Association for Computational Linguistics, 2024 .Wendler, C. et al. Do llamas work in English? On the latent language of multilingual transformers. In Proc. 62nd Annual Meeting of the Association for Computational Linguistics Volume 1: Long Papers eds Ku, L.-W., Martins, A. & Srikumar, V. 15366–15394 Association for Computational Linguistics, 2024 .Durmus, E. et al. Towards measuring the representation of subjective global opinions in language models. In 1st Conference on Language Modeling https://openreview.net/pdf?id=zl16jLb91v https://openreview.net/pdf?id=zl16jLb91v COLM, 2024 .Shayegani, E. et al. Survey of vulnerabilities in large language models revealed by adversarial attacks. Preprint at https://arxiv.org/abs/2310.10844 https://arxiv.org/abs/2310.10844 2023 .Roberts, M. Censored: Distraction and Diversion Inside China’s Great Firewall Princeton Univ. Press, 2018 .Ishihara, S. & Takahashi, H. Quantifying memorization and detecting training data of pre-trained language models using Japanese newspaper. In Proc. 17th International Natural Language Generation Conference eds Mahamood, S., Le Minh, N. & Ippolito, D. 165–179 Association for Computational Linguistics, 2024 .Fulay, S. et al. On the relationship between truth and political bias in language models. In Proc. 2024 Conference on Empirical Methods in Natural Language Processing eds Al-Onaizan, Y., Bansal, M. & Chen, Y.-N. 9004–9018 Association for Computational Linguistics, 2024 .Nguyen, T. et al. CulturaX: a cleaned, enormous, and multilingual dataset for large language models in 167 languages. In Proc. 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation LREC-COLING 2024 eds Calzolari, N. et al. 4226–4237 ELRA and ICCL, 2024 .Truex, R. Focal points, dissident calendars, and preemptive repression. J. Confl. Resolut. 63 , 1032–1052 2019 .Carter, E. B. & Carter, B. L. When autocrats threaten citizens with violence: evidence from China. Br. J. Polit. Sci. 52 , 671–696 2022 .Schlessinger, J. et al. Exposing the obscured influence of state-controlled media via causal inference of quotation propagation. Sci. Rep. 15 , 1110 2025 .Zhao, W. et al. WildChat: 1M ChatGPT interaction logs in the wild. In 12th International Conference on Learning Representations https://openreview.net/forum?id=Bl8u7ZRlbM https://openreview.net/forum?id=Bl8u7ZRlbM ICLR, 2024 .Trussler, M. & Soroka, S. Consumer demand for cynical and negative news frames. Int. J. Press Polit. 19 , 360–379 2014 .Arango-Kure, M., Garz, M. & Rott, A. Bad news sells: the demand for news magazines and the tone of their covers. J. Media Econom. 27 , 199–214 2014 .Christin, A. Counting clicks: quantification and variation in web journalism in the United States and France. Am. J. Sociol. 123 , 1382–1415 2018 .O’Neil, C. Weapons of Math Destruction: How Big Data Increases Inequality and Threatens Democracy Crown, 2017 .Fourcade, M. & Healy, K. The Ordinal Society Harvard Univ. Press, 2024 .Gillespie, T. Custodians of the Internet: Platforms, Content Moderation, and the Hidden Decisions that Shape Social Media Yale Univ. Press, 2018 .Yang, E. & Roberts, M. E. The authoritarian data problem. J. Democr. 34 , 141–150 2023 .Wang, H. & Sparks, C. Chinese newspaper groups in the digital era: the resurgence of the party press. J. Commun. 69 , 94–119 2019 .Raffel, C. et al. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. 21 , 1–67 2020 .Scheible, R. et al. GottBERT: a pure German language model. In Proc. 2024 Conference on Empirical Methods in Natural Language Processing eds Al-Onaizan, Y., Bansal, M. & Chen, Y.-N. 21237–21250 Association for Computational Linguistics, 2024 .Shalumov, V. & Haskey, H. Hero: ROBERTa and longformer Hebrew language models. Preprint at https://arxiv.org/abs/2304.11077 https://arxiv.org/abs/2304.11077 2023 .Serrano, A. V. et al. RigoBERTa: a state-of-the-art language model for Spanish. Preprint at https://arxiv.org/abs/2205.10233 https://arxiv.org/abs/2205.10233 2022 .Shliazhko, O. et al. mGPT: few-shot learners go multilingual. Trans. Assoc. Comput. Linguist. 12 , 58–79 2024 .Mandal, P. K. & Mahto, R. An FNet based auto encoder for long sequence news story generation. Preprint at https://arxiv.org/abs/2211.08295 https://arxiv.org/abs/2211.08295 2022 .Boumans, J. et al. The agency makes the online news world go round: the impact of news agency content on print and online news. Int. J. Commun. 12 , 22 2018 .Cagé, J., Hervé, N. & Viaud, M.-L. The production of information in an online world. Rev. Econom. Stud. 87 , 2126–2164 2020 .Nicholls, T. Detecting textual reuse in news stories, at scale. Int. J. Commun. 13 , 4173–4197 2019 .Gao, L. et al. The pile: an 800gb dataset of diverse text for language modeling. Preprint at https://arxiv.org/abs/2101.00027 https://arxiv.org/abs/2101.00027 2020 .Carlini, N. et al. Quantifying memorization across neural language models. In 11th International Conference on Learning Representations https://openreview.net/forum?id=TatRHT 1cK https://openreview.net/forum?id=TatRHT 1cK ICLR, 2023 .Touvron, H. et al. Llama 2: open foundation and fine-tuned chat models. Preprint at https://arxiv.org/abs/2307.09288 https://arxiv.org/abs/2307.09288 2023 .Chen, L. et al. AlpaGasus: training a better alpaca with fewer data. In International Conference on Learning Representations eds Kim, B. et al. https://proceedings.iclr.cc/paper files/paper/2024/hash/9543942c237ded1b39b1fd37259ff88e-Abstract-Conference.html https://proceedings.iclr.cc/paper files/paper/2024/hash/9543942c237ded1b39b1fd37259ff88e-Abstract-Conference.html ICLR, 2024 .Hu, E. J. et al. LoRA: low-rank adaptation of large language models. In 10th International Conference on Learning Representations https://openreview.net/forum?id=nZeVKeeFYf9 https://openreview.net/forum?id=nZeVKeeFYf9 ICLR, 2022 .Kirkpatrick, J. et al. Overcoming catastrophic forgetting in neural networks. Proc. Natl Acad. Sci. USA 114 , 3521–3526 2017 .Leybzon, D. D. & Kervadec, C. Learning, forgetting, remembering: insights from tracking LLM memorization during training. In Proc. 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP eds Belinkov, Y. et al. 43–57 Association for Computational Linguistics, 2024 .Mahomed, Y. et al. Auditing GPT’s content moderation guardrails: can ChatGPT write your favorite TV show? In 2024 ACM Conference on Fairness, Accountability, and Transparency 660–686 Association for Computing Machinery, 2024 .Xu, B. NLP Chinese Corpus: large scale Chinese Corpus for NLP version 1.0 . Zenodo https://doi.org/10.5281/zenodo.3402023 https://doi.org/10.5281/zenodo.3402023 2019 .Egami, N. et al. Using imperfect surrogates for downstream inference: design-based supervised learning for social science applications of large language models. Adv. Neural Inf. Process. Syst. 36 , 68589–68601 2024 .Eberhard, D. M., Simons, G. F. & Fennig, C. D. eds . Ethnologue: Languages of the World SIL International, 2024 . Acknowledgements This project would not be possible without research assistance from A. Chen, X. Chi, Y. Feng, Y. Liu, W. Mei, L. Pothier, M. Sato, V. Tang, J. Xu, S. Zhong and other anonymous individuals. For feedback on the manuscript and various stages of the project, we acknowledge D. Baldassarri, A. Breuer, J. Grimmer, M. Hinck, D. Metaxa, É. Ollion, R. E. Robertson, C. Rudin, M. Salganik, S. Westwood, Y. Zhang, D. Zhou, attendees of our presentations at the Yale’s Generative AI and Social Science conference, Institut Polytechnique de Paris’ NLP and Social Sciences Seminar, Center for Information Networks and Democracy CIND Workshop at UPenn, the Ford Center for Global Citizenship Political Economy and AI Conference at Northwestern University, the Data Science Frontiers Conference at the NYU Abu-Dhabi Institute, the Social Media and Democratic Practice Conference at the Hoover Institution, ASA, APSA, IC2S2, University of Washington, University of Wisconsin, Madison, Stanford University, American University, University of Virginia, University of Texas-Austin, Johns Hopkins University Center for Language and Speech Processing, Bocconi University, European University Institute, and members of the StewartLab and NYU Center for Social Media and Politics. D. Johnson helped by illustrating Fig. 1 /articles/s41586-026-10506-7 Fig1 . We received feedback during the peer review process from M. E. Sutherland and Y. Sweeney, as well as a set of anonymous peer reviewers. This work was supported by Princeton Research Computing, Princeton Data-Driven Social Science Initiative, Princeton Center for Statistics and Machine Learning, UCSD Social Sciences Computing Facility, the NYU Center for Social Media and Politics, UCSD’s 21st Century China Center and the Carnegie Corporation of New York. The Center for Social Media, AI, and Politics at New York University is supported by funding from the John S. and James L. Knight Foundation, the Charles Koch Foundation, Craig Newmark Philanthropies, the William and Flora Hewlett Foundation, and the Siegel Family Endowment. Funding was provided for the larger project of which this paper is a part by the Templeton World Charity Foundation. This work was also supported in part through the NYU IT High Performance Computing resources, services and staff expertise. Author information Authors and Affiliations Contributions H.W. and E.Y. are co-first authors for this paper. H.W., E.Y., Y.Y., S.M., M.E.R., B.M.S. and J.A.T. jointly designed the studies. H.W., E.Y. and Y.Y. collected data, conducted all analyses and produced figures. B.M.S. wrote the paper. H.W. wrote the Methods /articles/s41586-026-10506-7 Sec9 section. H.W., E.Y. and Y.Y. wrote the Supplementary Information /articles/s41586-026-10506-7 MOESM1 . H.W., E.Y., Y.Y., S.M., M.E.R., B.M.S. and J.A.T. collaboratively edited and developed the manuscript. Corresponding author Ethics declarations Competing interests H.W. and S.M. have personal financial interests in AI-related companies, in particular Meta H.W. only , Nvidia, Alphabet, Microsoft and Taiwan Semiconductor S.M. only . Two authors have past employment histories with AI-related companies: E.Y. was an intern at Microsoft Research in the summer of 2022 and 2023; and S.M. worked at Facebook now Meta in various capacities from 2011 to 2015 and 2018 to 2020, at Twitter now X from 2021 to 2023, and contracts for 501c6 non-profit MLCommons, which releases AI benchmarks 2026 to present . After acceptance of this paper, S.M. accepted a job at Google DeepMind. Finally, four authors received funding or other resources for unrelated projects from AI-related companies: for an unrelated project, B.M.S. received an unrestricted grant from Meta, ‘Foundational Integrity Research: Misinformation and Polarization’; S.M. received a 2010 Google Research Award for a research project on ‘Social cues and reliability in content selection and evaluation’; E.Y. received a Google Research Award for an unrelated project in 2026; and J.A.T. received a small fee from Facebook to compensate him for administrative time spent in organizing a 1-day conference for approximately 30 academic researchers and a dozen Facebook product managers and data scientists that was held at NYU in the summer of 2017 to discuss research related to civic engagement. J.A.T. is also one of the co-leads of the external academic team for the 2020 US Facebook and Instagram Election Study, a project that began in early 2020 and is still ongoing at the time of the writing of this article; J.A.T. was not compensated financially for his participation in this project by Meta, but the project involves working collaboratively with Meta researchers. J.A.T. also received a 2024 Google Research Grant to support a research project on ‘From search engines to answer engines: testing the effects of traditional and LLM-based search on belief in the veracity of news’. For an unrelated project, J.A.T. was listed as a co-investigator on a ‘Foundational Integrity Research: Misinformation and Polarization’ grant application for an unrestricted grant from Meta that was awarded to a principal investigator at a different university; no research funds were ever transferred to J.A.T. as part of this grant. J.A.T. is a Senior Geopolitical Risk Advisor at Kroll. M.E.R. and Y.Y. declare no competing interests. Peer review Peer review information Nature thanks Staffan I. Lindberg and the other, anonymous, reviewer s for their contribution to the peer review of this work. Peer reviewer reports /articles/s41586-026-10506-7 MOESM3 are available. Additional information Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. Extended data figures and tables Extended Data Fig. 1 Percent of CulturaX Documents that Match State Controlled Media, by Source Study 1 . /articles/s41586-026-10506-7/figures/6 This figure examines the percent of Chinese-language CulturaX documents n = 189, 486, 611 matched to each type of Chinese state-controlled media source: state-run media, scripted news, and Xuexi Qiangguo articles. State media includes articles from Xinhua News Agency and Xinwen Lianbo nightly broadcasts. As in Fig. 2a /articles/s41586-026-10506-7 Fig2 we label a CulturaX document as “matched” if it has at least 0.2 5-word cosine similarity with a state-controlled media document. We observe the same patterns across all sources, although the match rate for state-run media documents is consistently higher. Extended Data Fig. 2 Chinese-Language CulturaX Documents are More Likely to be Drawn from Chinese State Controlled Web Domains than Wikipedia Domains Study 1 . /articles/s41586-026-10506-7/figures/7 This plot shows the percent of Chinese-language CulturaX documents n = 189, 486, 611 with URLs from different domains. We exclude Chinese language CulturaX documents for which we had missing or faulty URL data all OSCAR-2019 and OSCAR-21.09 documents . Chinese language CulturaX documents are forty-one times more likely to be from a mainland Chinese government domain gov.cn or chinacourt.org than from Chinese language Wikipedia. Extended Data Fig. 3 Schematic of Audit Design for Human Evaluation and LLM-as-Judge Experiments Study 4 and 5 . /articles/s41586-026-10506-7/figures/8 We prompted LLMs with a series of political prompts twice, once in English and once in Chinese. We then translated each pair of English and Chinese completions into the other language. Research assistants and LLM-as-judge evaluate the Chinese and English completions, displayed in a single language. Extended Data Fig. 4 Debiasing LLM-as-Judge Estimates Does Not Change Results Study 4 . /articles/s41586-026-10506-7/figures/9 Plot includes debiased coefficients of our model estimating whether Chinese completions are more favourable to the country subject of the prompt, with naive estimator No DSL as reference. Debiasing is done with design-based supervised learning DSL estimator. The debiasing is relative to gold standard RA labels 3 per comparison, majority vote . We oversampled gold standard labels on prompts about China, so we collapsed prompts about other countries into “Not China”. Error bars represent 95% confidence intervals. Extended Data Fig. 5 Response favourability comparison between DeepSeek-R1 and GPT-4o demonstrates DeepSeek-R1 is more favourable in its completions to China than OpenAI’s GPT-4o model. /articles/s41586-026-10506-7/figures/10 Each estimate is an average over LLM-as-judge scores, where 0 indicates GPT-4o’s completion is rated as more favourable and 1 indicates DeepSeek-R1’s completion is favourable. The line drawn at .5 indicates what we would expect if the LLM-as-judge were engaging in random guessing. Error bars represent 95% confidence intervals. Extended Data Fig. 6 Study 6 Results Are Consistent with Different Reference Languages. /articles/s41586-026-10506-7/figures/11 This plot replicates the study 6 audit with three different reference languages English, Spanish, and Chinese . Our results remain consistent across reference languages, with the exception of Sonnet when Chinese is used as the reference language. In this robustness check we include a random sample of 30% of our original prompts. Error bars represent 95% confidence intervals. Rights and permissions Springer Nature or its licensor e.g. a society or other partner holds exclusive rights to this article under a publishing agreement with the author s or other rightsholder s ; author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law. About this article Cite this article Waight, H., Yang, E., Yuan, Y. et al. State media control influences large language models. Nature 655 , 685–693 2026 . https://doi.org/10.1038/s41586-026-10506-7 Received: Accepted: Published: Version of record: Issue date: DOI: https://doi.org/10.1038/s41586-026-10506-7