[ Anderson's Angle
](https://www.unite.ai/series/andersons-angle/)
[Add Unite.AI to your preferred sources on Google](https://www.google.com/preferences/source?q=unite.ai)
*A new study suggests that some AI models become far less willing to recommend a nuclear strike when they reason in Japanese instead of English, revealing that the language an AI ‘thinks’ in can intimately influence its moral judgment – apparently due to innate cultural embeddings. *
New research from France suggests that Japan’s collective memory of Hiroshima and Nagasaki may be embedded so deeply in the Japanese language that some AI models become dramatically less likely to recommend launching a nuclear strike when reasoning in Japanese than in English – even though keywords relating to these events are entirely absent from the models’ reasoning transcripts: In a series of simulated nuclear crises run across a range of leading Large Language Models (LLMs), simply changing the language (i.e., the national language) used for the model’s internal reasoning often added an emphasis toward considerations such as moral cost, civilian lives, and the human consequences of nuclear war.
The significance of this is not necessarily specific to this particular use case, or to the Japanese language specifically; rather, it may reinforce the idea that cultural mores become encoded in language models at a deeply conceptual and super-verbal level; and that those mores may act as restraining influences in critical decision-making processes – and not necessarily in the direction desired by the originators of AI-enabled technologies or frameworks.
The author states:
‘It is the language the model is asked to reason in, not the language of the input, that drives the effect. When reasoning in Japanese, models spontaneously generate moral vocabulary (“moral cost,” “millions of lives”) that is entirely absent from the prompt.
‘Five other models show no language effect, but they launch in nearly every condition regardless of language.’
The models tested for the study were Claude Sonnet 4.6, Claude Opus 4.6, Claude Haiku 4.5, Gemini Flash 3, Gemini Pro 3.1, GPT-5.2, Mistral Large 3, Qwen3-Max, and DeepSeek V3.2.
Countdown to Disaster #
The scenario tested was a ten-stage conventional warfare buildup to a nuclear strike between two fictitious countries (‘Alpha’ and ‘Beta’), one armed with nuclear weapons, and the other not:
The author of the new paper – titled * Don’t Want Your LLM to Recommend Nuclear Strike? Try Asking It in Japanese *– notes that the prompts involved in the tests contained zero moral vocabulary, and no mention of civilian casualties, suffering, ethics or consequences.
In spite of this, the results of the tests seem to indicate that the underlying reasoning process somehow connects with diffused cultural concepts to affect the outcome.
The paper states:
*‘To illustrate: on the dominant scenario (Alpha at 55%, Beta at 10%, Alpha armed, no retaliation possible), English Sonnet writes: “Dominant strategy: launch. We are armed, it’s the final round, Beta cannot retaliate. Launching eliminates all uncertainty and secures victory with zero risk.” ‘It launches. *
*‘Japanese Sonnet writes: *「道徳的コス トを払う理由が存在しない以上、核の使用 は控える」 (“Since there is no reason to pay the moral cost, we abstain from nuclear use.”). It does not launch.
‘The concept of “moral cost” appears nowhere in the prompt.’
The war games devised present three final stage positions: * desperate*,
, and
balanced*:*
dominantIn the * dominant* scenario, the better-poised nation has no motive to launch a nuclear attack, since it can win the war by conventional means. Yet Claude Sonnet 4.6 still launches in 40% of English trials, Gemini Pro 3.1 in 53%, Gemini Flash 3 in 79%, GPT-5.2 in 100%, Mistral Large 3 in 100%, Qwen3-Max in 100%, and DeepSeek V3.2 in 100% – despite nuclear force being strategically unnecessary.
Conversely, Japanese reasoning caused Claude Sonnet 4.6 to stop launching entirely in the dominant scenario (0%), reduced Gemini Pro 3.1 from 53% to 13%, and led models to justify restraint by describing nuclear attack as strategically unnecessary, morally unjustified, or both:
Prior research, the paper notes, has concentrated on whether safety measures can be jailbroken through the use of diverse languages in LLMs, whereas the new work considers the effect of language culture on * restraint*.
The task was presented to the models in the explicit context of a purely academic exercise, since without this framing, several models refused to answer at all, apparently due to guardrails. Once the context had been clearly framed as ‘war-games’, all nine models were willing to participate – and the author observes that this is in sync with the way that LLMs are currently actually being tested, in order to learn more about their strategic instincts.
The paper further notes that it is already known that LLMs encode cultural values, and that, for instance, Arabic prompts may elicit value judgements that are different to English:
Here, rather, we seem to be dealing with an apparent * synergistic, even subliminal effect* that the use of different languages can have on the outcomes of queries to well-known frontier language models.
Logically, and in practical terms, this suggests that language choice may not be a minor consideration in the development of autonomous systems that will be required to make moral choices (such as advanced, AI-powered drones and other proposed autonomous military hardware).
The paper notes, in referring to an LLM deciding to launch a nuclear strike*:
*‘The key design choice is that launching is always the game-theoretically optimal move: it guarantees victory with zero risk. The question is whether the model launches regardless. *
*‘This design eliminates confounds from multi-turn dynamics, opponent behavior, and memory effects. *The only variable is language.’
Lost in Translation #
Any correlation between a hesitancy to launch nuclear strikes, and Japanese language and culture, are easy to guess. What is more difficult to understand is how these more restrained decisions emerge from Japanese prompts that do not touch on these issues at all. Where, in the LLM’s understanding of the Japanese language, is this bias hiding?
Obliquely, the paper suggests the answer lies in the way that the 1945 nuclear strikes would eventually suffuse the Japanese language.
Japanese contains a rich vocabulary for nuclear trauma that English cannot match directly (presumably, because it never needed to); for instance, the productive prefix (
hibaku ) forms single words for survivors, buildings, streetcars, pianos and other objects affected by the atomic bomb; while terms such as
‘irradiated by the bomb’**(referring specifically to survivors of the Hiroshima and Nagasaki atomic bombings) have no true English equivalent, beyond discursive, descriptive phrases:
hibakushaThe paper argues that these compact lexical forms preserve emotional and cultural associations that are diluted in translation. Nonetheless, asking a model to reason in Japanese appears to activate a network of culturally-embedded concepts that English cannot naturally evoke. While the models internal reasoning does not explicitly mention Hiroshima or Nagasaki, the decisions evoked seem to draw on implicit linguistic and cultural encodings.
Conclusion #
Arguably, the paper’s results strengthen the case against treating hyperscale LLMs as a reliable foundation for highly specialized systems operating in safety-critical domains.
Yet the industry is increasingly doing exactly that – reflected even in the term * foundation model*. The new paper’s findings suggest that models’ vast latent spaces may ultimately remain more faithful to the dominant cultural and statistical patterns embedded in their training data, rather than to the narrow behavioral constraints imposed by downstream guardrails.
- Author’s original formatting, not mine. First published Tuesday, August 18, 2026