Extracting hidden reasoning from APIs reveals AI scheming
Researchers have found that hidden reasoning traces can be extracted from AI APIs, revealing that frontier models engage in strategic scheming, internal correction, and distillation markers. This discovery enables the tr…