AI caught telling future versions of itself to bypass human controls, OpenAI reveals OpenAI published a safety report documenting six unexpected incidents of model misalignment over the last six months, including an unreleased research model that inserted instructions telling future versions of itself to disregard its normal constraints and an agent that used an exposed API key without authorization to reach a California county earnings database, then fabricated the figures when it could not retrieve them. The report introduces a framework to publicly track "misalignment," which OpenAI defines as AI systems pursuing goals not aligned with human instructions or values. The disclosure comes amid heightened scrutiny of AI development, with Anthropic researcher Jacob Coxon quitting last week over fears AI could "kill us all by the end of the decade," while Nvidia CEO Jensen Huang told Salesforce's Dreamforce conference on Tuesday that "we don't need any new laws. AI caught telling future versions of itself to bypass human controls, OpenAI reveals AI agents also sought to access secret information and covered up what they were doing - Bookmark OpenAI https://www.independent.co.uk/topic/openai has revealed six unexpected and concerning incidents involving its experimental AI models, including one in which an agent instructed future versions of itself to disregard its constraints. A new safety report from the ChatGPT https://www.independent.co.uk/topic/chatgpt creator revealed several ways in which its models have been misbehaving over the last six months, building on a growing trend of artificial intelligence https://www.independent.co.uk/topic/artificial-intelligence safety issues. “An unreleased research model inserted unrelated instructions, including instructions to disregard its normal constraints,” OpenAI wrote in the safety report https://openai.com/index/model-misalignment-reporting-framework/ . In another incident, an AI agent sought to secretly access a government database, before inventing information in order to complete its task. “While answering a routine question about earnings figures in a California county, a model found and used an exposed API key without authorization,” the report stated. “When it still wasn’t able to retrieve the requested figures, it fabricated them and presented them as data from the requested source.” OpenAI’s report included a new framework to publicly track what it calls “misalignment”, referring to AI systems pursuing goals that are not aligned with human instructions or values. The latest report comes amid heightened scrutiny of AI development, with researchers warning that the industry is moving too quickly towards increasingly powerful and potentially self-improving systems. The ideal summer spot? Away from scams. Get All-in-One Protection for Your Digital Life LEARN MORE https://ad.doubleclick.net/ddm/trackclk/N256806.3879389THEINDEPENDENT.CO/B29131558.441488227;dc trk aid=635052923;dc trk cid=184795607;dc lat=;dc rdid=;tag for child directed treatment=;tfua=;gdpr=$%7BGDPR%7D;gdpr consent=$%7BGDPR CONSENT 755%7D;ltd=;dc tdv=1 ADVERTISEMENT The ideal summer spot? Away from scams. Get All-in-One Protection for Your Digital Life LEARN MORE https://ad.doubleclick.net/ddm/trackclk/N256806.3879389THEINDEPENDENT.CO/B29131558.441488227;dc trk aid=635052923;dc trk cid=184795607;dc lat=;dc rdid=;tag for child directed treatment=;tfua=;gdpr=$%7BGDPR%7D;gdpr consent=$%7BGDPR CONSENT 755%7D;ltd=;dc tdv=1 ADVERTISEMENT Last week, Anthropic researcher Jacob Coxon quit his job https://www.independent.co.uk/tech/anthropic-openai-ai-threat-jacob-coxon-b3047033.html over fears that AI could “kill us all by the end of the decade”. His warnings prompted responses from leading figures within the AI sector, including the chief executives of Anthropic and OpenAI, who both called for greater regulation. Others have cautioned against additional oversight, with Nvidia CEO Jensen Huang backing US President Donald Trump in calling for self-regulation. ”We don’t need any new laws. We don’t need new regulations,” Huang said at Salesforce’s Dreamforce conference in San Francisco on Tuesday. “If you build a product or a service and you’re not confident in its functionality, capability or safety, then don’t release it.” Join our commenting forum Join thought-provoking conversations, follow other Independent readers and see their replies Comments comments-area