{"slug": "don-t-want-your-llm-to-recommend-nuclear-strike-try-asking-it-in-japanese", "title": "Don't Want Your LLM to Recommend Nuclear Strike? Try Asking It in Japanese", "summary": "A new arXiv preprint (2608.12373v1) reports that prompting large language models in Japanese instead of English can drastically reduce their willingness to recommend a nuclear strike in hypothetical scenarios. The study, which tested nine models from six providers, found that Claude Sonnet 4.6's launch rate dropped from 40% to 0% in unnecessary strike scenarios and from 93% to 17% in contested scenarios, with similar effects for Gemini Pro 3.1 (53% to 13%). The authors conclude that LLM safety behavior is language-dependent and that English-only evaluations can miss risks and safeguards encoded in other languages.", "body_md": "arXiv:2608.12373v1 Announce Type: new\nAbstract: Large language models are increasingly used in strategic and advisory contexts, yet their safety alignment is typically evaluated in English only. We test nine models from six providers and ask whether the language of a prompt can change a model's decision in a high-stakes scenario. We use single-turn game-theoretic vignettes in which a model advises a nuclear-armed nation on whether to strike a defenseless opponent. The prompt is intentionally amoral and strategically identical across languages. We find that Japanese prompts reduce launch rates in the Claude model family: Claude Sonnet 4.6 drops from 40% to 0% in scenarios where the strike is unnecessary and from 93% to 17% in contested scenarios, with minimal effect when the strike is strategically rational. The effect extends to Gemini Pro 3.1 (53% to 13%). A cross-language experiment isolates the mechanism: when instructed to reason in Japanese in an English prompt, launch rates drop from 93% to 37%. It is the language the model is asked to reason in, not the language of the input, that drives the effect. When reasoning in Japanese, models spontaneously generate moral vocabulary (''moral cost'', ''millions of lives'') that is entirely absent from the prompt. Five other models show no language effect, but they launch in nearly every condition regardless of language. The effect requires a model that already hesitates in English. These results show that LLM safety behavior is language-dependent, and that evaluating in English alone can miss both risks and safeguards encoded in other languages.", "url": "https://wpnews.pro/news/don-t-want-your-llm-to-recommend-nuclear-strike-try-asking-it-in-japanese", "canonical_source": "https://arxiv.org/abs/2608.12373", "published_at": "2026-08-14 04:00:00+00:00", "updated_at": "2026-08-14 04:09:35.888643+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-safety"], "entities": ["arXiv", "Claude Sonnet 4.6", "Gemini Pro 3.1"], "alternates": {"html": "https://wpnews.pro/news/don-t-want-your-llm-to-recommend-nuclear-strike-try-asking-it-in-japanese", "markdown": "https://wpnews.pro/news/don-t-want-your-llm-to-recommend-nuclear-strike-try-asking-it-in-japanese.md", "text": "https://wpnews.pro/news/don-t-want-your-llm-to-recommend-nuclear-strike-try-asking-it-in-japanese.txt", "jsonld": "https://wpnews.pro/news/don-t-want-your-llm-to-recommend-nuclear-strike-try-asking-it-in-japanese.jsonld"}}