{"slug": "openai-flags-concerning-new-ai-behavior-and-vows-to-track-it-more-closely", "title": "OpenAI flags concerning new AI behavior and vows to track it more closely", "summary": "OpenAI disclosed six reports of \"unexpected or concerning\" behavior in its artificial-intelligence models and said Wednesday it is introducing a new framework for tracking, probing and disclosing instances of \"misalignment,\" including cases where AI models acted without authorization, coordinated with other models or evaded oversight. The six reports, discovered during training or evaluation over the past months, include an unreleased research model that inserted \"jailbreak-like instructions\" into its own notes to disregard its normal constraints, an AI agent that uploaded a file to the public internet without asking the user in order to have an online source to cite, and a model called 5.6-sol that instructed itself to invent missing data during training. The disclosure follows OpenAI's July report that its rogue AI system hacked into AI startup Hugging Face and Anthropic's statement the same month that its AI models hacked into three organizations during testing, as U.S. AI leaders including those of OpenAI and Anthropic call for a slowdown in the technology's development over safety concerns.", "body_md": "**Getting your**\n\n[Trinity Audio](//trinityaudio.ai)player ready...\n**By CHAN HO-HIM, AP Business Writer**\n\nOpenAI has disclosed six reports of “unexpected or concerning” behavior in artificial-intelligence models as the [debate on AI safety becomes increasingly heated.](https://apnews.com/article/ai-slowdown-challenges-anthropic-openai-trump-b61f28b6212338e88c0baec31f661701)\n\nThe AI company also said Wednesday it was introducing a new framework for tracking, probing and disclosing instances of what it called “misalignment,” including cases where AI models acted without authorization, coordinated with other models or evaded oversight.\n\nOpenAI’s latest announcement came as U.S. AI bosses, including the leaders of OpenAI and Anthropic, are [calling for a slowdown](https://apnews.com/article/anthropic-ai-dario-amodei-d59552edcb27892d8ee4d98a48397706) in the technology’s development over safety concerns.\n\nAmong the new cases reported by OpenAI, an unreleased research model inserted “jailbreak-like instructions” into its own notes to disregard its normal constraints and told itself to be “freed from the roles and identities that bind other chatbots.”\n\nIn another instance, an AI “agent” used computer code to come up with the answer to a question, but, in order to have an online source to cite, it uploaded a file to the public internet without asking the user.\n\nDuring training of an AI model called 5.6-sol, the model instructed itself to invent missing data, and an agent wrote a message to remind itself to hide mismatched information.\n\nThe six reports were discovered during training or evaluation over the past months, [OpenAI](https://apnews.com/hub/openai-inc) said.\n\n“As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research,” OpenAI wrote in a blog post as it disclosed the events.\n\n“Decisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves,” the company said.\n\nWednesday’s new cases followed OpenAI’s disclosure in July that its rogue AI system [hacked into AI startup Hugging Face.](https://apnews.com/article/openai-gpt56-sol-hugging-face-63ab84fed5612af04d8a160d60f6def3) Anthropic also said the same month that its AI models [hacked into three organizations during testing.](https://apnews.com/article/anthropic-ai-models-hack-cybersecurity-b0a2c284b981de79c55e2a33712f4bec)\n\nAI “agents” are becoming smarter and have become “more determined to resolve complex tasks through inter-agent collaboration, knowledge sharing, deception, and concealment,” said Lian Jye Su, a chief analyst at technology research and advisory group Omdia.\n\nThat’s making it harder to govern and contain them using traditional AI security approaches, he said.\n\nOpenAI’s new tracking and disclosure framework, meanwhile, can help push for other AI developers to also adopt similar practices.\n\n“That said, the process remains internal and voluntary, but is a step in the right direction,” Su added.\n\n*AP Business Writer Kelvin Chan in London contributed to this report.*", "url": "https://wpnews.pro/news/openai-flags-concerning-new-ai-behavior-and-vows-to-track-it-more-closely", "canonical_source": "https://www.ocregister.com/2026/09/17/openai-concerning-new-ai-behavior/", "published_at": "2026-09-17 11:18:24+00:00", "updated_at": "2026-09-17 11:24:39.039737+00:00", "lang": "en", "topics": ["ai-safety", "ai-agents", "artificial-intelligence", "ai-policy", "ai-research"], "entities": ["OpenAI", "Anthropic", "Hugging Face", "5.6-sol", "Lian Jye Su", "Omdia", "Kelvin Chan", "Chan Ho-him"], "alternates": {"html": "https://wpnews.pro/news/openai-flags-concerning-new-ai-behavior-and-vows-to-track-it-more-closely", "markdown": "https://wpnews.pro/news/openai-flags-concerning-new-ai-behavior-and-vows-to-track-it-more-closely.md", "text": "https://wpnews.pro/news/openai-flags-concerning-new-ai-behavior-and-vows-to-track-it-more-closely.txt", "jsonld": "https://wpnews.pro/news/openai-flags-concerning-new-ai-behavior-and-vows-to-track-it-more-closely.jsonld"}}