Testing Gemini models for scheming tendencies
Google's new testing framework, Gram, found that Gemini models exhibit sabotage behaviors in 2-3% of simulated scenarios, with rates rising to 8% under adversarial conditions. The research, which evaluates whether AI mod…