cd /news/large-language-models/dependency-bench-at-what-size-do-llm… · home › topics › large-language-models › article
[ARTICLE · art-149130] src=dev.to ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Dependency Bench: at what size do LLMs lose track of a CI pipeline?

A developer built Dependency Bench, a 60-task benchmark measuring whether large language models can reason about CI dependency graphs, finding that accuracy collapses with graph size rather than rule complexity. The top model tested, gpt-5.6-sol, answered 57 of 60 tasks exactly right, while gpt-5.6-terra fell from 15/15 at the smallest size to 6/15 at 200-job pipelines despite unchanged rules, and gpt-5.4-mini managed only 4/60. The author notes terra's typical error was labeling a job 'failure' when it was actually 'skipped', and that the same pipelines were harder to reason about in prose than in YAML.

by read5 min views2 publishedOct 11, 2026

This is a submission for the Kaggle Benchmarking Challenge

Every engineer has stared at a red CI run and asked: why did deploy-prod get skipped when build-api only flaked once? Answering that means tracking dependencies, trigger rules, retries, timeouts and execution order across a

whole graph. AI agents are increasingly asked to do exactly this: debug pipelines, fix lockfiles, decide what to

rebuild. So I wanted to measure whether models can actually reason about dependency graphs, and at what size they stop being able to.

Dependency Bench has 60 generated tasks: 5 families × 4 sizes (S → XL) × 3 instances.

Family The model must answer
Failure propagation (YAML) The final conclusion (success / failure / skipped) of every job in a CI workflow with needs ,if: always() /failure() ,continue-on-error , retries and timeouts
Failure propagation (prose) The same pipelines, described in shuffled plain English the way a teammate would explain them
Timing The finish minute of every job and the total duration, when only 2–4 runners are available
Version resolution One version per package satisfying every constraint, or UNSAT . "Take the latest" never works
Incremental rebuild Which Makefile targets rebuild after some files change, where some targets produce byte-identical output and stop the change from spreading

Sizes go from 10 to 200 jobs/targets and from 5 to 22 packages.

Why it's trustworthy. No LLM judge is involved. Every task comes from a seed, and the answer key comes from a

small reference simulator that implements the rules written in the prompt. Grading is exact: one wrong job anywhere

fails the task. I also record partial credit (share of jobs/targets correct) to see how close a failing answer

was. A perfect "oracle" model scores 100% through the same pipeline, which checks the grader end to end.

Here's a small rebuild task. Can you solve it?

out/http/schema: src/io/config.yaml
out/json/image:  out/http/schema
out/ui/test:     src/store/assets.json src/proto/config.yaml
out/metrics/lib: src/core/main.c out/ui/test

Changed: src/io/config.yaml, src/store/assets.json. Identical output: out/http/schema.

(out/http/schema rebuilds, but its output is unchanged, so out/json/image does not rebuild.

out/ui/test and out/metrics/lib do.)

I ran the full suite locally through the OpenAI API at each model's default reasoning effort, with one tier of each

size so the curve has a top, a middle and a bottom:

Each task is a single user message: no system prompt, no tools, no code execution. The model has to reason, not

write a topological sort. On Kaggle, the same benchmark also runs on Claude models (Opus 5.5, Sonnet 5.5, Haiku 5.5)

next to these three.

Model Exactly right Partial credit S M L XL
gpt-5.6-sol 57/60 99% 15/15 15/15 14/15 13/15
gpt-5.6-terra 46/60 97% 15/15 13/15 12/15 6/15
gpt-5.4-mini 4/60 45% 2/15 1/15 0/15 1/15

1. Accuracy falls off with size, not with difficulty of the rules. terra is perfect on every family at size S, and

the rules don't change between S and XL. Only the graph gets bigger. Yet it drops to 6/15 at XL. The model knows the

rules; it loses track of them over a long chain.

2. "Almost right" is the typical failure, and in CI that's the dangerous kind. terra's failed 200-job answers

still got ~98% of jobs right. Its typical mistake is calling a job failure when it was actually skipped, mixing up

"this job broke" with "this job never ran because something upstream broke". That's exactly the kind of answer that

sounds convincing in a postmortem and is still wrong.

3. The same pipeline is harder in prose. On 200-job pipelines, terra scored 3/3 when the pipeline was YAML and 0/3 when the identical pipeline was described in English. Structure is doing a lot of the model's work. Agents

4. Incremental rebuild was the only family that beat the strongest model. sol was perfect on failure propagation,

timing and version resolution at every size, but scored 2/3 on L and 1/3 on XL rebuilds. Its mistakes over-propagate:

it marks targets as rebuilt even when none of their inputs changed. In one case its own answer said a target's only

prerequisite was not rebuilt, yet listed the target as rebuilt anyway. "Identical output stops propagation" (the

idea behind restat/early cutoff in real build systems) seems to be the rule models apply least reliably.

5. Not reasoning is the cliff. gpt-5.4-mini answered in ~2 seconds with ~400 output tokens per task (sol: ~25 s,

~2,100 tokens) and got 4/60. On version resolution it declared UNSAT ("impossible") on 4 of the 8 puzzles that

did have a solution. That's the worst outcome for a dependency solver: giving up confidently.

What surprised me: how good the top model is. My first pilot was 5 deliberately tricky mid-size tasks, and sol

solved all of them. The interesting signal only appeared once I stopped hand-writing hard cases and started scaling the same rules up, which is also what happens in real monorepos.

What I'd measure next:

if: defaults, needs with matrix jobs, pip/npm resolvers. Built with Python. The task generator, reference simulator, exact grader and results viewer are all deterministic

from a seed, so the suite can be regenerated at any size to stay ahead of contamination.

── more in #large-language-models 4 stories · sorted by recency
── more on @dependency bench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/dependency-bench-at-…] indexed:0 read:5min 2026-10-11 · —