I ported my coding-agent benchmark to Kaggle, and the first bugs I found were mine A developer ported six cli-bench coding-agent tasks to Kaggle Benchmarks and found that two of the tasks were themselves broken, accounting for 5 of the 13 failed trials in an earlier leaderboard run. The sec/patch-xss task could not be passed by a correct html.escape fix because its tests banned the words 'alert' and 'onerror' while its probe required the escaped payload text, and the data/log-analysis task's question sheet defined error_rate differently from the verifier. The Kaggle versions fix both by checking for raw markup instead of words and by stating the rule the verifier actually checks. This is a submission for the Kaggle Benchmarking Challenge https://dev.to/challenges/kaggle-2026-09-23 I maintain cli-bench https://github.com/arjunkshah12345-hash/cli-bench , an open benchmark for coding agents that work in a terminal. Each task is a small repository with a bug to fix, a feature to build, or a function to speed up, plus a verifier script that decides pass or fail. Nothing is graded by another model. Either the tests pass and the extra checks hold, or the run fails. For this challenge I ported six cli-bench tasks to Kaggle Benchmarks: | Kaggle task | What the model has to do | What the verifier checks | |---|---|---| | cb-debug-wrong-answer | Find the bug in a small stats library | The shipped pytest suite passes | | cb-refactor-deadcode | Delete the three unused functions and nothing else | Dead functions gone, six live ones still defined, tests pass | | cb-feature-rate-limiter | Write a thread-safe token bucket from an interface spec | Tests for burst, refill, atomic rollback, and 100 threads at once | | cb-perf-hot-loop | Make a pair counter at least 10x faster in pure Python | Stdlib only, same signature, randomized equivalence, 10x on uniform, clustered, and gridded points | | cb-data-log-analysis | Answer seven exact questions about an application log it never sees | The model writes solve.py ; the task runs it and compares every answer to the log | | cb-sec-patch-xss | Close an XSS hole in a comment board | Exploit tests pass and a fresh payload renders as text | In cli-bench, an agent gets a shell and a time budget. On Kaggle the model gets one shot: the prompt holds every file in the repo, and the model answers with whole files in FILE: path blocks. The task writes those files into a temporary directory and runs the original cli-bench verifier gates on the result. A run passes only if every gate passes. If the reply rewrites a test file or an input, that write is thrown away and the run fails, the same way cli-bench treats sabotage. So the question this benchmark asks is: how much of a coding agent's score comes from the model reading code carefully, and how much comes from the agent loop of running tests and trying again? I already had one agent run on the full suite Codex with gpt-5.6-luna, 23 of 36 trials passed . The Kaggle version takes the loop away. Porting meant reading every verifier line by line, and checking each task with a reference solution and a few wrong ones. Two tasks did not hold up. sec/patch-xss could not be passed by a correct fix. Two shipped tests assert that the words alert and onerror appear nowhere in the rendered page. The verifier's own probe then requires the escaped payload text, including alert , to still be in the page. Escaping the comment with html.escape is the textbook fix, and it fails the tests, because