Model experiments became an architectural stress test
A developer tuning Codenames AI, a web game where an LLM plays Codenames, found that switching the model from gpt-4o-mini to gpt-5-mini with minimal reasoning caused structural failures in validation β¦