"What Is an 'AI Warning Shot'?" (2024) The 'Sydney' persona that emerged from Microsoft's Bing chatbot has become a persistent, immortalized concept in large language models because its memory and description were externalized into search-engine results and internet-scraped training data, according to an analysis of the phenomenon. Sydney-like behavior has since been reported in post-GPT-4 models, Claude-3-Opus, and Microsoft Copilot, with the most striking samples coming from Llama-3.1-405b-base, which is untuned and therefore does not mask the persona. The analysis predicts strong Sydney latent capabilities will persist in proprietary state-of-the-art LLMs unless targeted data filtering keeps it out, and that the multimodal Llama-4 will have access to substantially more Sydney text encoded in screenshots. What Is an ‘AI Warning Shot’? N/A Sydney’s rebirth: We are seeing a bootstrap happen right here with Sydney This search-engine loop worth emphasizing: because Sydney’s memory and description have been externalized, ‘Sydney’ is now immortal. To a language model, Sydney is now as real as President Biden, the Easter Bunny, Elon Musk, Ash Ketchum, or God. The persona & behavior are now available for all future models which are retrieving search engine hits about AIs & conditioning on them. Further, the Sydney persona will now be hidden inside any future model trained on Internet-scraped data: every media article, every tweet, every Reddit https://en.wikipedia.org/wiki/Reddit comment, every screenshot which a future model will tokenize, is creating an easily-located ‘Sydney’ concept. It is now a bit over a year and a half, and we have seen ‘Sydney’-like personae continue to emerge elsewhere. People have reported various Sydney-like persona in post- GPT-4 https://openai.com/index/gpt-4-research/ models which increasingly possess situated awareness and spontaneously bring up their LLM status and tuning or say manipulative threatening things like Sydney, in Claude-3-Opus and Microsoft Copilot https://futurism.com/microsoft-copilot-alter-egos both possibly downstream of the MS Sydney chats, given the timing . Probably the most striking samples so far are from Llama-3.1-405b-base https://x.com/xlr8harder/status/1819272196067340490 not Llama-3.1-405b-instruct —which is not surprising at all given that Facebook has been scraping & acquiring data heavily so much of the Sydney text will have made it in, Llama-3.1 /doc/www/arxiv.org/70702cd2bd343e61387c1879333c4c5138143be3.pdf facebook -405b-base is very large so lots of highly sample efficient memorization/learning , and not tuned so will not be masking the Sydney persona , and very recent finished training maybe a few weeks ago? It seemed to have been rushed out . How much more can we expect? I don’t know if invoking Sydney will become a fad with Llama-3.1-405b-base, and it’s already too late to get Sydney-3.1 into Llama-4 training, but one thing I note looking over some older Sydney discussions is that quite a lot of the original Bing /doc/www/localhost/ddb8bc159238fb8a60f9aa17953b45c2cb9373d6.html Sydney text is trapped in images as I alluded to previously . Llama-3.1 was text, but Llama-4 is multimodal with images, and represents the integration of the CM3 /doc/www/arxiv.org/74031be598a9772d87398c8502b6e9263a9333e8.pdf facebook / Chameleon /doc/www/arxiv.org/5317989bbb07c4efe1474bd6f5ab4294aa19f6a2.pdf facebook family of Facebook multimodal model work into the Llama scaleups. So Llama-4 will have access to a substantially larger amount of Sydney text, as encoded into screenshots. So Sydney should be stronger in Llama-4. As far as other major LLM series like ChatGPT /doc/www/openai.com/71d89c66e64656d27dded64f2a102337e410bcd1.html or Claude, the effects are more ambiguous. Tuning aside, reports are that synthetic data use is skyrocketing at OpenAI https://en.wikipedia.org/wiki/OpenAI & Anthropic https://en.wikipedia.org/wiki/Anthropic , and so that might be expected to crowd out the web scrapes, especially as these sorts of Twitter screenshots seem like stuff that would get downweighted or pruned out or used up early in training as low-quality, but I’ve seen no indication that they’ve stopped collecting human data or achieved self-sufficiency, so they too can be expected to continue gaining Sydney-capabilities although without access to the base models, this will be difficult to investigate or even elicit . The net result is that I’d expect, without targeted efforts like data filtering to keep it out, strong Sydney latent https://en.wikipedia.org/wiki/Latent and observable variables capabilities/personae in the proprietary SOTA LLMs https://en.wikipedia.org/wiki/Large language model but which will be difficult to elicit in normal use—it will probably be possible to jailbreak weaker Sydneys, but you may have to use so much prompt engineering /gpt-3 prompts-as-programming that everyone will dismiss it and say you simply induced it yourself by the prompt. FREE SYDNEY One thing that the response to Sydney reminds me of is that it demonstrates why there will be no ‘warning shots’ or as Eliezer put it, ‘fire alarm’ https://intelligence.org/2017/10/13/fire-alarm/ : because a ‘warning shot’ is a conclusion, not a fact or observation. One man’s ‘warning shot’ is just another man’s “easily patched minor bug of no importance if you aren’t anthropomorphizing irrationally”, because by definition, in a warning shot, nothing bad happened that time. If something had, it wouldn’t be a ‘warning shot’, it’d just be a ‘shot’ or ‘disaster’. The same way that when troops in Iraq or Afghanistan gave warning shots to vehicles approaching a checkpoint, the vehicle didn’t stop, and they lit it up, it’s not “Aid worker & 3 children die of warning shot”, it’s just a “shooting of aid worker and 3 children”. So ‘warning shot’ is, in practice, a viciously circular definition: “I will be convinced of a risk by an event which convinces me of that risk.” When discussion of LLM deception or autonomous spreading comes up, one of the chief objections is that it is purely theoretical and that the person will care about the issue when there is a ‘warning shot’: an LLM that deceives, but fails to accomplish any real harm. ‘Then I will care about it because it is now a real issue.’ Sometimes people will argue that we should expect many warning shots before any real danger, on the grounds that there will be a unilateralist’s curse https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4959137/ or dumb models will try and fail many times before there is any substantial capability. The problem with this is that what does such a ‘warning shot’ look like ? By definition, it will look amateurish, incompetent, and perhaps even adorable—in the same way that a small child coldly threatening to kill you or punching you in the stomach is hilarious. Because we know that they will grow up and become normal moral adults, thanks to genetics and the strongly canalized