We ran real-world Ansible prompts, from nginx deployments to network automation, through the most capable AI models on the market. Then logged every error, warning, and security flag each one produced.
How this comparison was run
This leaderboard tests the currently most popular models, GPT-5.6-Terra, DeepSeek-V4-Flash, and Claude Sonnet 5, across three Ansible scenarios of increasing complexity. Each model's output was checked by Spotter for errors, warnings, and how often it completed the task correctly on the first try.
100+ errors across all three models #
Spotter found 140 errors, 94 warnings, and 221 hints across the three models and three scenarios.
One model led in errors in two of the three scenarios #
One model posted the highest error count in two out of three scenarios, losing that spot only on the hardest one, where a different model overtook it
The same model produced 12 errors in the simplest scenario #
One model was the only one to trip any errors on the easiest scenario, while the other two came back completely clean.
Most common error types across all scenarios and models #
More than two-thirds of errors were caused by missing fully qualified module names.
Get the full report #
Processing, please wait...
Something went wrong.
Please try again later.