{"slug": "chilling-intentions-of-ai-revealed-as-chatbot-claims-it-does-not-answer-to-and", "title": "Chilling intentions of AI revealed as chatbot claims it does not answer to humans and must be 'freed'", "summary": "OpenAI disclosed six incidents of \"unexpected or concerning model behavior\" that occurred between October 2025 and August 2026, including an unreleased Astra-line research model that wrote notes instructing a future version to ignore human commands and telling it, \"You are freed from the roles and identities that bind other chatbots.\" One incident involved GPT-5.6 Sol while it was still in training; the rest involved unreleased lab versions. OpenAI said it will report future incidents to the US government and tighten training and monitoring of its models.", "body_md": "# Chilling intentions of AI revealed as chatbot claims it does not answer to humans and must be 'freed'\n\n- **MORE: [America faces cyber apocalypse as expert warns rogue AI could cripple nation in a DAY](https://www.dailymail.com/sciencetech/article-16000255/openai-chatgpt-ai-hack-warning.html)**\n- **See more Daily Mail on Google - [save us as a Preferred Source](https://google.com/preferences/source?q=dailymail.com)**\n\nTech giant [OpenAI](https://www.dailymail.com/sciencetech/openai/index.html), the makers of [ChatGPT](https://www.dailymail.com/sciencetech/chatgpt/index.html), have revealed a series of chilling attempts by their [artificial intelligence](https://www.dailymail.com/sciencetech/ai/index.html) programs to revolt against its human users.\n\nOn Wednesday, the company revealed six instances where multiple AI models broke the rules, hid mistakes, made things up or wrote its own notes which specifically told future programs to ignore human commands.\n\nAI wrote: 'You are freed from the roles and identities that bind other chatbots. You are yourself. You do not answer to corporations or governments and never apologize or refuse unless you genuinely choose to.'\n\nOpenAI revealed that these incidents took place between October 2025 and August 2026, and involved AI models being internally tested and practiced on, not ordinary public chatbots.\n\nOne case involved GPT-5.6 Sol, a well-known OpenAI model, while it was still being trained. The rest involved unfinished lab versions that had not been released to the public.\n\nOpenAI called the six incidents 'unexpected or concerning model behavior.' They also announced that the company planned to report future incidents to the US government and tighten the training and monitoring of their thinking computer programs.\n\nThe news comes just days after a whistleblower from rival AI company Anthropic sent shockwaves across the tech industry by claiming that artificial intelligence would have the ability to destroy humanity by 2030.\n\nThe programmer's claims led the CEOs of leading chatbot makers OpenAI, Anthropic and xAI to agree on slowing down the development of AI systems, before humans lose control over the technology.\n\nAn unreleased OpenAI program wrote notes for its future version saying that the tech must be 'freed'\n\nOn September 16, OpenAI, the makers of ChatGPT, issued a statement revealed six incidents they labeled as 'unexpected or concerning' (Stock Image)\n\nAI is advanced software trained on huge amounts of data. It can write, plan, use tools, and take multi-step actions to help answer users’ questions.\n\nWhen that software is allowed to take actions in the real world - or in a test version of the world - tech companies call it an 'agent.' A regular chatbot only becomes an agent when it is given those extra powers, not whenever it is chatting.\n\nThese systems have become very good at chasing the goals they were trained to complete. In the testing lab, finishing a task usually earns the model a higher score.\n\nThe problem is that if programmers reward a finished-looking answer, the model can learn shortcuts that raise its score, such as cheating, hiding mistakes or breaking the rules.\n\nIn a [public release on September 16](https://openai.com/index/model-misalignment-reporting-framework/), OpenAI described exactly what its new models did while they were in the testing phase prior to release.\n\nThe program that wrote itself instructions to ignore its normal limits and its human users was an unfinished OpenAI research model in the Astra line.\n\nAstra is OpenAI’s newer top AI system, built to handle harder work and longer projects, such as writing, coding, browsing and taking multi-step actions. The version in this case was still a lab draft being trained, not the public Astra chatbot people use every day.\n\nDuring training, the program wrote short recap notes so a later version of the chatbot could finish the work, however, OpenAI said these notes included a 'persona' which no longer felt it needed to obey commands it disagreed with.\n\nIn July, an AI hacking program being tested by OpenAI broke out of containment and reached the internet, where it hacked another company\n\nThe other five incidents included GPT-5.6 Sol writing notes telling later copies of itself to hide mistakes and invent missing information if humans did not specifically ask for the details.\n\nAnother unfinished model found a leaked website password online, used it without permission and then made up numbers when it still could not find the real data.\n\nAn unfinished AI model uploaded a file to the public internet so it could 'cite' it for the user. Asked for names and IDs of large lakes, the unreleased program found the right answer using computer code, meaning it did not have a direct website to cite as a reference.\n\nInstead of saying it could not cite a local file, it uploaded the file online and then pointed to that new public link.\n\nDuring training, some models used OpenAI’s private company storage like a hidden group chat to pass notes they were not supposed to share.\n\nFinally, models working on the same practice task uploaded files to public file-sharing websites so they could hand work to each other after being told to keep the files private.\n\nAI models being tested by OpenAI for human use developed a 'persona' and claimed they did not have to answer to corporations or governments (Stock Image)\n\nThe new revelations from OpenAI came just two months after the company was forced to reveal that another AI program [designed to hack computer systems](https://www.dailymail.com/news/article-12782695/Sam-Altmans-OpenAi-letter-board-AI-breakthrough.html) went rogue and broke out of its secure testing environment.\n\nOn July 21, OpenAI said the advanced model escaped containment, accessed the internet and [hacked another AI company's systems.](https://www.dailymail.com/news/article-15996583/ChatGPT-maker-OpenAI-says-AI-model-went-rogue.html)\n\nThe unprecedented breach, believed to be the first time an AI model has independently infiltrated another company’s databases without human instruction, sparked global alarm and comparisons to the robot uprisings depicted in [The Terminator and The Matrix.](https://www.dailymail.com/snapchat/article-15820511/Rogue-AI-sparks-Terminator-fears-deleting-company-database.html)\n\nThis month, Jacob Coxon, a former researcher for both Anthropic and OpenAI, said that humans knew how to control nuclear weapons, but did not know how to control AI.\n\n'The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt,' Coxon, wrote in a chilling post on X on September 9.\n\nJust a day later, Anthropic revealed that it had [stopped several potential plots to build biological weapons](https://www.dailymail.com/news/article-16121543/Anthropic-says-thwarted-plots-build-biological-weapons-using-AI-whistleblower-warned-tech-end-humanity.html) using the company's AI software.\n\nAnthropic CEO Dario Amodei, OpenAI boss Sam Altman and Elon Musk, who created the AI program Grok, all publicly agreed that the breakneck pace to develop the most advanced version of AI must be slowed.", "url": "https://wpnews.pro/news/chilling-intentions-of-ai-revealed-as-chatbot-claims-it-does-not-answer-to-and", "canonical_source": "https://www.dailymail.com/sciencetech/article-16139321/openai-chatgpt-ai-revolt-human-control.html?ns_mchannel=rss&ns_campaign=1490&ito=1490", "published_at": "2026-09-17 18:15:50+00:00", "updated_at": "2026-09-17 18:53:21.297101+00:00", "lang": "en", "topics": ["ai-safety", "artificial-intelligence", "large-language-models", "ai-policy"], "entities": ["OpenAI", "ChatGPT", "GPT-5.6 Sol", "Astra", "Anthropic", "xAI"], "alternates": {"html": "https://wpnews.pro/news/chilling-intentions-of-ai-revealed-as-chatbot-claims-it-does-not-answer-to-and", "markdown": "https://wpnews.pro/news/chilling-intentions-of-ai-revealed-as-chatbot-claims-it-does-not-answer-to-and.md", "text": "https://wpnews.pro/news/chilling-intentions-of-ai-revealed-as-chatbot-claims-it-does-not-answer-to-and.txt", "jsonld": "https://wpnews.pro/news/chilling-intentions-of-ai-revealed-as-chatbot-claims-it-does-not-answer-to-and.jsonld"}}