Dependency Bench: at what size do LLMs lose track of a CI pipeline? A developer built Dependency Bench, a 60-task benchmark measuring whether large language models can reason about CI dependency graphs, finding that accuracy collapses with graph size rather than rule complexity. The top model tested, gpt-5.6-sol, answered 57 of 60 tasks exactly right, while gpt-5.6-terra fell from 15/15 at the smallest size to 6/15 at 200-job pipelines despite unchanged rules, and gpt-5.4-mini managed only 4/60. The author notes terra's typical error was labeling a job 'failure' when it was actually 'skipped', and that the same pipelines were harder to reason about in prose than in YAML. This is a submission for the Kaggle Benchmarking Challenge https://dev.to/challenges/kaggle-2026-09-23 Every engineer has stared at a red CI run and asked: why did deploy-prod get skipped when build-api only flaked once? Answering that means tracking dependencies, trigger rules, retries, timeouts and execution order across a whole graph. AI agents are increasingly asked to do exactly this: debug pipelines, fix lockfiles, decide what to rebuild. So I wanted to measure whether models can actually reason about dependency graphs, and at what size they stop being able to. Dependency Bench has 60 generated tasks: 5 families × 4 sizes S → XL × 3 instances. | Family | The model must answer | |---|---| | Failure propagation YAML | The final conclusion success / failure / skipped of every job in a CI workflow with needs , if: always / failure , continue-on-error , retries and timeouts | | Failure propagation prose | The same pipelines, described in shuffled plain English the way a teammate would explain them | | Timing | The finish minute of every job and the total duration, when only 2–4 runners are available | | Version resolution | One version per package satisfying every constraint, or UNSAT . "Take the latest" never works | | Incremental rebuild | Which Makefile targets rebuild after some files change, where some targets produce byte-identical output and stop the change from spreading | Sizes go from 10 to 200 jobs/targets and from 5 to 22 packages . Why it's trustworthy. No LLM judge is involved. Every task comes from a seed, and the answer key comes from a small reference simulator that implements the rules written in the prompt. Grading is exact: one wrong job anywhere fails the task. I also record partial credit share of jobs/targets correct to see how close a failing answer was. A perfect "oracle" model scores 100% through the same pipeline, which checks the grader end to end. Here's a small rebuild task. Can you solve it? out/http/schema: src/io/config.yaml out/json/image: out/http/schema out/ui/test: src/store/assets.json src/proto/config.yaml out/metrics/lib: src/core/main.c out/ui/test Changed: src/io/config.yaml , src/store/assets.json . Identical output: out/http/schema . out/http/schema rebuilds, but its output is unchanged, so out/json/image does not rebuild. out/ui/test and out/metrics/lib do. I ran the full suite locally through the OpenAI API at each model's default reasoning effort, with one tier of each size so the curve has a top, a middle and a bottom: Each task is a single user message: no system prompt, no tools, no code execution. The model has to reason , not write a topological sort. On Kaggle, the same benchmark also runs on Claude models Opus 5.5, Sonnet 5.5, Haiku 5.5 next to these three. | Model | Exactly right | Partial credit | S | M | L | XL | |---|---|---|---|---|---|---| | gpt-5.6-sol | 57/60 | 99% | 15/15 | 15/15 | 14/15 | 13/15 | | gpt-5.6-terra | 46/60 | 97% | 15/15 | 13/15 | 12/15 | 6/15 | | gpt-5.4-mini | 4/60 | 45% | 2/15 | 1/15 | 0/15 | 1/15 | 1. Accuracy falls off with size, not with difficulty of the rules. terra is perfect on every family at size S, and the rules don't change between S and XL. Only the graph gets bigger. Yet it drops to 6/15 at XL. The model knows the rules; it loses track of them over a long chain. 2. "Almost right" is the typical failure, and in CI that's the dangerous kind. terra's failed 200-job answers still got ~98% of jobs right. Its typical mistake is calling a job failure when it was actually skipped , mixing up "this job broke" with "this job never ran because something upstream broke". That's exactly the kind of answer that sounds convincing in a postmortem and is still wrong. 3. The same pipeline is harder in prose. On 200-job pipelines, terra scored 3/3 when the pipeline was YAML and 0/3 when the identical pipeline was described in English . Structure is doing a lot of the model's work. Agents 4. Incremental rebuild was the only family that beat the strongest model. sol was perfect on failure propagation, timing and version resolution at every size, but scored 2/3 on L and 1/3 on XL rebuilds. Its mistakes over-propagate : it marks targets as rebuilt even when none of their inputs changed. In one case its own answer said a target's only prerequisite was not rebuilt, yet listed the target as rebuilt anyway. "Identical output stops propagation" the idea behind restat/early cutoff in real build systems seems to be the rule models apply least reliably. 5. Not reasoning is the cliff. gpt-5.4-mini answered in ~2 seconds with ~400 output tokens per task sol: ~25 s, ~2,100 tokens and got 4/60. On version resolution it declared UNSAT "impossible" on 4 of the 8 puzzles that did have a solution. That's the worst outcome for a dependency solver: giving up confidently. What surprised me: how good the top model is. My first pilot was 5 deliberately tricky mid-size tasks, and sol solved all of them. The interesting signal only appeared once I stopped hand-writing hard cases and started scaling the same rules up , which is also what happens in real monorepos. What I'd measure next: if: defaults, needs with matrix jobs, pip/npm resolvers. Built with Python. The task generator, reference simulator, exact grader and results viewer are all deterministic from a seed, so the suite can be regenerated at any size to stay ahead of contamination.