AI experts sound alarm about ‘terrifying’ rumoured change to how ChatGPT works AI experts have raised alarms about a rumoured update to OpenAI's upcoming Astra model that could use a technique called 'opaque recurrence', potentially undermining chain-of-thought monitoring used to ensure AI safety. The Information reported that OpenAI, Anthropic, and Google DeepMind are considering such updates, though OpenAI staff have dismissed the reports as 'confused'. AI safety expert Zvi Mowshowitz warned the technique is 'playing with fire', and researcher Gary Marcus called it a 'terrifying situation' that could cross a redline. AI experts sound alarm about ‘terrifying’ rumoured change to how ChatGPT works Reported update could make it harder to understand how systems are behaving, critics warn - Bookmark - CommentsGo to comments Artificial intelligence experts have sounded alarm about a “terrifying” rumoured update to ChatGPT /topic/chatgpt that could make it harder to monitor what AI /topic/ai systems are doing. This week, reports suggested that OpenAI /topic/openai ’s upcoming Astra /topic/astra model will use a technique called “opaque recurrence”, though staff at the company have dismissed those reports. Tech publication The Information , which first reported the rumoured change at OpenAI, said in a separate report that both Anthropic /topic/anthropic and Google DeepMind were considering similar updates. Opaque recurrence, which is also known as recurrent depth, is a complex and technical tool that could allow AI systems to better deal with difficult tasks. But experts also advise that it could undermine chain-of-thought monitoring, another technique that allows people to see some of the reasoning that has gone into producing an AI system’s output. Chain-of-thought monitoring has some limitations. But it can be an important tool in ensuring that AI is safe, since it allows researchers to see possible misbehaviour by models and understand how it came about. In the recent cyber attack by an OpenAI model on fellow artificial intelligence firm Hugging Face, for instance, researchers were able to monitor how the experimental system had decided to launch its attack. OpenAI has not confirmed that its new Astra model uses the technique, and suggested the initial reports are “confused”, and The Information ‘s story suggested that it does so only in a limited way. When OpenAI announced earlier this week that it was pressing ahead with the launch of Astra, it had actually said that the model included “additional chain-of-thought monitoring to rapidly detect and contain potentially misaligned actions”. But AI experts nonetheless cautioned that weakening the ability to monitor AI systems could begin a dangerous process. “The technique is playing with fire, risking a taboo that OpenAI and Anthropic have fought to establish that we work hard to maintain Chain of Thought faithfulness and monitorability for as long as we can,” wrote AI safety expert Zvi Mowshowitz on Substack https://thezvi.substack.com/p/anthropic-has-some-alignment-problems . “More intensive use of such techniques would probably damage monitorability.” Gary Marcus, a researcher who regularly criticises the big AI companies and their claims, said that it would be crossing a redline and was a “terrifying situation”. “Chain of thought monitorability is limited and not fully reliable but it’s practically the only thread we have here to prevent seriously bad outcomes,” he wrote. OpenAI staff members suggested that reporting around the issue was “confused”. Chief scientist Jakub Pachocki warned that the rumours could start “a race into unmonitorability” and said that OpenAI remains committed to chain-of thought monitoring and “deeply care s about this technique”. Similarly, Dean W Ball, the company’s head of strategic futures, criticised the fact that much of the discussion was happening on social media, which meant “adjudicating technically complex and nuanced claims on the timeline with almost no ground-truth information about what is actually happening”. “It is frankly insane and crazymaking and grating for everyone involved,” he wrote on X. Join our commenting forum Join thought-provoking conversations, follow other Independent readers and see their replies Comments comments-area