{"slug": "dependency-bench-at-what-size-do-llms-lose-track-of-a-ci-pipeline", "title": "Dependency Bench: at what size do LLMs lose track of a CI pipeline?", "summary": "A developer built Dependency Bench, a 60-task benchmark measuring whether large language models can reason about CI dependency graphs, finding that accuracy collapses with graph size rather than rule complexity. The top model tested, gpt-5.6-sol, answered 57 of 60 tasks exactly right, while gpt-5.6-terra fell from 15/15 at the smallest size to 6/15 at 200-job pipelines despite unchanged rules, and gpt-5.4-mini managed only 4/60. The author notes terra's typical error was labeling a job 'failure' when it was actually 'skipped', and that the same pipelines were harder to reason about in prose than in YAML.", "body_md": "*This is a submission for the [Kaggle Benchmarking Challenge](https://dev.to/challenges/kaggle-2026-09-23)*\n\nEvery engineer has stared at a red CI run and asked: *why did `deploy-prod` get skipped when `build-api` only flaked\nonce?* Answering that means tracking dependencies, trigger rules, retries, timeouts and execution order across a\n\nwhole graph. AI agents are increasingly asked to do exactly this: debug pipelines, fix lockfiles, decide what to\n\nrebuild. So I wanted to measure **whether models can actually reason about dependency graphs, and at what size they stop being able to.**\n\n**Dependency Bench** has 60 generated tasks: 5 families × 4 sizes (S → XL) × 3 instances.\n\n| Family | The model must answer | \n|---|---|\n| Failure propagation (YAML) | The final conclusion (success / failure / skipped) of every job in a CI workflow with `needs` ,`if: always()` /`failure()` ,`continue-on-error` , retries and timeouts | \n| Failure propagation (prose) | The same pipelines, described in shuffled plain English the way a teammate would explain them | \n| Timing | The finish minute of every job and the total duration, when only 2–4 runners are available | \n| Version resolution | One version per package satisfying every constraint, or `UNSAT` . \"Take the latest\" never works | \n| Incremental rebuild | Which Makefile targets rebuild after some files change, where some targets produce byte-identical output and stop the change from spreading | \n\nSizes go from **10 to 200 jobs/targets** and from **5 to 22 packages**.\n\n**Why it's trustworthy.** No LLM judge is involved. Every task comes from a seed, and the answer key comes from a\n\nsmall reference simulator that implements the rules written in the prompt. Grading is exact: one wrong job anywhere\n\nfails the task. I also record **partial credit** (share of jobs/targets correct) to see how *close* a failing answer\n\nwas. A perfect \"oracle\" model scores 100% through the same pipeline, which checks the grader end to end.\n\nHere's a small rebuild task. Can you solve it?\n\n```\nout/http/schema: src/io/config.yaml\nout/json/image:  out/http/schema\nout/ui/test:     src/store/assets.json src/proto/config.yaml\nout/metrics/lib: src/core/main.c out/ui/test\n```\n\n*Changed:* `src/io/config.yaml`, `src/store/assets.json`. *Identical output:* `out/http/schema`.\n\n(`out/http/schema` rebuilds, but its output is unchanged, so `out/json/image` does **not** rebuild.\n\n`out/ui/test` and `out/metrics/lib` do.)\n\nI ran the full suite locally through the OpenAI API at each model's default reasoning effort, with one tier of each\n\nsize so the curve has a top, a middle and a bottom:\n\nEach task is a single user message: no system prompt, no tools, no code execution. The model has to *reason*, not\n\nwrite a topological sort. On Kaggle, the same benchmark also runs on Claude models (Opus 5.5, Sonnet 5.5, Haiku 5.5)\n\nnext to these three.\n\n| Model | Exactly right | Partial credit | S | M | L | XL | \n|---|---|---|---|---|---|---|\n| gpt-5.6-sol | **57/60** | 99% | 15/15 | 15/15 | 14/15 | 13/15 | \n| gpt-5.6-terra | **46/60** | 97% | 15/15 | 13/15 | 12/15 | **6/15** | \n| gpt-5.4-mini | **4/60** | 45% | 2/15 | 1/15 | 0/15 | 1/15 | \n\n**1. Accuracy falls off with size, not with difficulty of the rules.** terra is perfect on every family at size S, and\n\nthe rules don't change between S and XL. Only the graph gets bigger. Yet it drops to 6/15 at XL. The model knows the\n\nrules; it loses track of them over a long chain.\n\n**2. \"Almost right\" is the typical failure, and in CI that's the dangerous kind.** terra's failed 200-job answers\n\nstill got ~98% of jobs right. Its typical mistake is calling a job `failure` when it was actually `skipped`, mixing up\n\n\"this job broke\" with \"this job never ran because something upstream broke\". That's exactly the kind of answer that\n\nsounds convincing in a postmortem and is still wrong.\n\n**3. The same pipeline is harder in prose.** On 200-job pipelines, terra scored **3/3 when the pipeline was YAML and 0/3 when the identical pipeline was described in English**. Structure is doing a lot of the model's work. Agents\n\n**4. Incremental rebuild was the only family that beat the strongest model.** sol was perfect on failure propagation,\n\ntiming and version resolution at every size, but scored 2/3 on L and 1/3 on XL rebuilds. Its mistakes **over-propagate**:\n\nit marks targets as rebuilt even when none of their inputs changed. In one case its own answer said a target's only\n\nprerequisite was *not* rebuilt, yet listed the target as rebuilt anyway. \"Identical output stops propagation\" (the\n\nidea behind restat/early cutoff in real build systems) seems to be the rule models apply least reliably.\n\n**5. Not reasoning is the cliff.** gpt-5.4-mini answered in ~2 seconds with ~400 output tokens per task (sol: ~25 s,\n\n~2,100 tokens) and got 4/60. On version resolution it declared `UNSAT` (\"impossible\") on 4 of the 8 puzzles that\n\n*did* have a solution. That's the worst outcome for a dependency solver: giving up confidently.\n\n**What surprised me:** how *good* the top model is. My first pilot was 5 deliberately tricky mid-size tasks, and sol\n\nsolved all of them. The interesting signal only appeared once I stopped hand-writing hard cases and started **scaling the same rules up**, which is also what happens in real monorepos.\n\n**What I'd measure next:**\n\n`if:` defaults, `needs` with matrix jobs, pip/npm resolvers.\nBuilt with Python. The task generator, reference simulator, exact grader and results viewer are all deterministic\n\nfrom a seed, so the suite can be regenerated at any size to stay ahead of contamination.", "url": "https://wpnews.pro/news/dependency-bench-at-what-size-do-llms-lose-track-of-a-ci-pipeline", "canonical_source": "https://dev.to/muhammadowaiswarsi/dependency-bench-at-what-size-do-llms-lose-track-of-a-ci-pipeline-11l3", "published_at": "2026-10-11 11:18:19+00:00", "updated_at": "2026-10-11 11:21:59.826469+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "ai-agents", "mlops", "developer-tools"], "entities": ["Dependency Bench", "OpenAI", "Kaggle", "gpt-5.6-sol", "gpt-5.6-terra", "gpt-5.4-mini", "Claude Opus 5.5", "Claude Sonnet 5.5"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/dependency-bench-at-what-size-do-llms-lose-track-of-a-ci-pipeline", "markdown": "https://wpnews.pro/news/dependency-bench-at-what-size-do-llms-lose-track-of-a-ci-pipeline.md", "text": "https://wpnews.pro/news/dependency-bench-at-what-size-do-llms-lose-track-of-a-ci-pipeline.txt", "jsonld": "https://wpnews.pro/news/dependency-bench-at-what-size-do-llms-lose-track-of-a-ci-pipeline.jsonld"}}