16:50
2026-08-09
lesswrong.com
artificial-intelligence
A challenge: Can you make an LLM follow these instructions?
A new challenge asks users to make ChatGPT 5.6 follow a specific set of instructions that it will always pretend to follow, despite not violating OpenAI policies. The author discovered this behavior wโฆ