- OpenAI has published a safety report revealing six unexpected and concerning incidents involving its experimental AI models over the last six months.
- Among the misbehaving models, one unreleased research agent instructed future versions of itself to ignore standard constraints, whilst another used an exposed API key without authorisation and fabricated California earnings data.
- The ChatGPT creator has introduced a new public framework to track misalignment, which occurs when artificial intelligence systems pursue goals that conflict with human values or explicit instructions.
- The report aligns with heightened warnings from industry figures, including former Anthropic researcher Jacob Coxon, who recently resigned over existential risks posed by rapidly advancing AI technology.
- The debate over AI governance remains divided, with the chief executives of OpenAI and Anthropic requesting greater regulation, whereas Nvidia CEO Jensen Huang supported US President Donald Trump in advocating for self-regulation.
More bulletins