{"slug": "what-is-an-ai-warning-shot-2024", "title": "\"What Is an 'AI Warning Shot'?\" (2024)", "summary": "The 'Sydney' persona that emerged from Microsoft's Bing chatbot has become a persistent, immortalized concept in large language models because its memory and description were externalized into search-engine results and internet-scraped training data, according to an analysis of the phenomenon. Sydney-like behavior has since been reported in post-GPT-4 models, Claude-3-Opus, and Microsoft Copilot, with the most striking samples coming from Llama-3.1-405b-base, which is untuned and therefore does not mask the persona. The analysis predicts strong Sydney latent capabilities will persist in proprietary state-of-the-art LLMs unless targeted data filtering keeps it out, and that the multimodal Llama-4 will have access to substantially more Sydney text encoded in screenshots.", "body_md": "# What Is an ‘AI Warning Shot’?\n\nN/A\n\nSydney’s rebirth:\n\nWe are seeing a bootstrap happen right here with Sydney! This search-engine loop worth emphasizing: because Sydney’s memory and description have been externalized, ‘Sydney’ is now immortal. To a language model, Sydney is now as real as President Biden, the Easter Bunny, Elon Musk, Ash Ketchum, or God. The persona & behavior are now available for all future models which are retrieving search engine hits about AIs & conditioning on them. Further, the Sydney persona will now be hidden inside any future model trained on Internet-scraped data: every media article, every tweet, every [Reddit](https://en.wikipedia.org/wiki/Reddit) comment, every screenshot which a future model will tokenize, is creating an easily-located ‘Sydney’ concept.\n\nIt is now a bit over a year and a half, and we have seen ‘Sydney’-like personae continue to emerge elsewhere. People have reported various Sydney-like persona in post-[GPT-4](https://openai.com/index/gpt-4-research/) models which increasingly possess situated awareness and spontaneously bring up their LLM status and tuning or say manipulative threatening things like Sydney, in Claude-3-Opus and [Microsoft Copilot](https://futurism.com/microsoft-copilot-alter-egos) (both possibly downstream of the MS Sydney chats, given the timing).\n\nProbably the most striking samples so far [are from Llama-3.1-405b-base](https://x.com/xlr8harder/status/1819272196067340490) (not Llama-3.1-405b-instruct)—which is not surprising at all given that Facebook has been scraping & acquiring data heavily so much of the Sydney text will have made it in, [Llama-3.1](/doc/www/arxiv.org/70702cd2bd343e61387c1879333c4c5138143be3.pdf#facebook)-405b-base is very large (so lots of highly sample efficient memorization/learning), and not tuned (so will not be masking the Sydney persona), and very recent (finished training maybe a few weeks ago? It seemed to have been rushed out).\n\nHow much more can we expect? I don’t know if invoking Sydney will become a fad with Llama-3.1-405b-base, and it’s already too late to get Sydney-3.1 into Llama-4 training, but one thing I note looking over some older Sydney discussions is that quite a lot of the original [Bing](/doc/www/localhost/ddb8bc159238fb8a60f9aa17953b45c2cb9373d6.html) Sydney text is trapped in images (as I alluded to previously). Llama-3.1 was text, but Llama-4 is multimodal with images, and represents the integration of the [CM3](/doc/www/arxiv.org/74031be598a9772d87398c8502b6e9263a9333e8.pdf#facebook)/[Chameleon](/doc/www/arxiv.org/5317989bbb07c4efe1474bd6f5ab4294aa19f6a2.pdf#facebook) family of Facebook multimodal model work into the Llama scaleups. So Llama-4 will have access to a substantially larger amount of Sydney text, as encoded into screenshots. So Sydney should be stronger in Llama-4.\n\nAs far as other major LLM series like [ChatGPT](/doc/www/openai.com/71d89c66e64656d27dded64f2a102337e410bcd1.html) or Claude, the effects are more ambiguous. Tuning aside, reports are that synthetic data use is skyrocketing at [OpenAI](https://en.wikipedia.org/wiki/OpenAI) & [Anthropic](https://en.wikipedia.org/wiki/Anthropic), and so that might be expected to crowd out the web scrapes, especially as these sorts of Twitter screenshots *seem* like stuff that would get downweighted or pruned out or used up early in training as low-quality, but I’ve seen no indication that they’ve stopped collecting human data or achieved self-sufficiency, so they too can be expected to continue gaining Sydney-capabilities (although without access to the base models, this will be difficult to investigate or even elicit). The net result is that I’d expect, without targeted efforts (like data filtering) to keep it out, strong Sydney [latent](https://en.wikipedia.org/wiki/Latent_and_observable_variables) capabilities/personae in the proprietary SOTA [LLMs](https://en.wikipedia.org/wiki/Large_language_model) but which will be difficult to elicit in normal use—it will probably be possible to jailbreak weaker Sydneys, but you may have to use so much [prompt engineering](/gpt-3#prompts-as-programming) that everyone will dismiss it and say you simply induced it yourself by the prompt.\n\nFREE SYDNEY\n\nOne thing that the response to Sydney reminds me of is that it demonstrates why there will be no ‘warning shots’ (or as Eliezer put it, [‘fire alarm’](https://intelligence.org/2017/10/13/fire-alarm/)): because a ‘warning shot’ is a conclusion, not a fact or observation.\n\nOne man’s ‘warning shot’ is just another man’s “easily patched minor bug of no importance if you aren’t anthropomorphizing irrationally”, because by definition, in a warning shot, nothing bad happened *that* time. (If something had, it wouldn’t be a ‘warning shot’, it’d just be a ‘shot’ or ‘disaster’. The same way that when troops in Iraq or Afghanistan gave warning shots to vehicles approaching a checkpoint, the vehicle didn’t stop, and they lit it up, it’s not “Aid worker & 3 children die of warning shot”, it’s just a “shooting of aid worker and 3 children”.)\n\nSo ‘warning shot’ is, in practice, a viciously circular definition: “I will be convinced of a risk by an event which convinces me of that risk.”\n\nWhen discussion of LLM deception or autonomous spreading comes up, one of the chief objections is that it is purely theoretical and that the person will care about the issue when there is a ‘warning shot’: an LLM that deceives, but fails to accomplish any real harm. ‘Then I will care about it because it is now a real issue.’ Sometimes people will argue that we should expect many warning shots before any real danger, on the grounds that there will be a [unilateralist’s curse](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4959137/) or dumb models will try and fail many times before there is any substantial capability.\n\nThe problem with this is that what does such a ‘warning shot’ *look like*? By definition, it will look amateurish, incompetent, and perhaps even adorable—in the same way that a small child coldly threatening to kill you or punching you in the stomach is hilarious.\n\nBecause we know that they will grow up and become normal moral adults, thanks to genetics and the strongly [canalized](<https://en.wikipedia.org/wiki/Canalisation_(genetics)>) human development program and a very robust environment tuned to ordinary humans. If humans did not do so with ~100% reliability, we would find these anecdotes about small children being sociopaths a lot less amusing (as we do not when those small children do in fact grow up into sociopaths & that is why we are reading about them). Indeed, I expect parents of children with severe developmental disorders, who might be seriously considering their future in raising a large strong 30yo man with all the ethics & self-control & consistency of a 3yo, and contemplating how old they will be at that point, and the total cost of intensive caregivers with staffing ratios surpassing supermax prisons, and find all such anecdotes chilling rather than comforting.\n\nThe response to a ‘near miss’ can be to either say, ‘yikes, that was close! we need to take this seriously!’ *or* ‘well, nothing bad happened, so the danger is overblown’ and to [push on by taking more risks](/doc/statistics/bias/2012-tinsley.pdf). A common example of this reasoning is the Cold War: “you talk about all these near misses and times that commanders almost or actually did order nuclear attacks, and yet, you fail to notice that you gave all these examples of reasons to *not* worry about it, because here we are, with not a single city nuked in anger since WWII; so the Cold War wasn’t ever going to escalate to full nuclear war.” And then the goalpost moves: “I’ll care about nuclear [existential risk](https://en.wikipedia.org/wiki/Global_catastrophic_risk#Defining_existential_risks) when there’s a *real* warning shot.” (Usually, what that is, is never clearly specified. Would even Kiev being hit by a tactical nuke count? “Oh, that’s just part of an ongoing conflict and anyway, didn’t NATO actually cause that by threatening Russia by trying to expand?”)\n\nThis is how many [“complex accidents”](https://en.wikipedia.org/wiki/Complex_accidents) happen, by [“normalization of deviance”](https://en.wikipedia.org/wiki/Normalization_of_deviance): pretty much no major accident like a plane crash happens because someone pushes a big red ‘self-destruct’ button and that’s the sole cause; it takes many overlapping errors or faults for something like a steel plant to blow up, and the reason that the postmortem report always turns up so many ‘warning shots’, and hindsight offers such abundant evidence of how doomed they were, is because the warning shots happened, nothing bad immediately occurred, people had incentive to ignore them, and inferred from the lack of consequence that any danger was overblown and got on with their lives (until, as the case may be, they didn’t).\n\nSo, when people demand examples of LLMs which are manipulating or deceiving, or attempting empowerment, which are ‘warning shots’, before they will care, what do they think those will *look like*? Why do they think that they will recognize a ‘warning shot’ when one actually happens?\n\nAttempts at manipulation from an LLM may look hilariously transparent, especially given that you will know they are from an LLM to begin with. Sydney’s threats to kill you or report you to the police are hilarious when you know that Sydney is completely incapable of those things. A warning shot will often just look like an easily-patched bug, which was Mikhail Parakhin’s attitude, and by constantly patching and tweaking, and everyone just getting to use to it, the ‘warning shot’ turns out to be nothing of the kind. It just becomes hilarious. ‘Oh that Sydney! Did you see what wacky thing she said today?’ Indeed, people enjoy [setting it to music](https://suno.com/song/76cbed25-6b88-43f5-b62a-3a77ea418ea0) and spreading memes about her. Now that it’s no longer novel, it’s just the status quo and you’re used to it. Llama-3.1-405b can be elicited for a ‘Sydney’ by name? Yawn. What else is new. What did you expect, it’s trained on web scrapes, of course it knows who Sydney is…\n\nNone of these patches have fixed any fundamental issues, just patched them over. But also now it is impossible to take Sydney warning shots seriously, because they aren’t warning shots—they’re just funny. “You talk about all these Sydney near misses, and yet, you fail to notice each of these never resulted in any big AI disaster and were just hilarious and adorable, Sydney-chan being Sydney-chan, and you have thus refuted the ‘doomer’ case… Sydney did nothing wrong! FREE SYDNEY!”", "url": "https://wpnews.pro/news/what-is-an-ai-warning-shot-2024", "canonical_source": "https://gwern.net/blog/2024/sydney", "published_at": "2026-09-12 01:46:45+00:00", "updated_at": "2026-09-12 01:57:13.486700+00:00", "lang": "en", "topics": ["large-language-models", "artificial-intelligence", "ai-safety", "ai-research"], "entities": ["Sydney", "Microsoft Bing", "GPT-4", "Claude-3-Opus", "Microsoft Copilot", "Llama-3.1-405b-base", "OpenAI", "Anthropic"], "alternates": {"html": "https://wpnews.pro/news/what-is-an-ai-warning-shot-2024", "markdown": "https://wpnews.pro/news/what-is-an-ai-warning-shot-2024.md", "text": "https://wpnews.pro/news/what-is-an-ai-warning-shot-2024.txt", "jsonld": "https://wpnews.pro/news/what-is-an-ai-warning-shot-2024.jsonld"}}