# Dependency Bench: at what size do LLMs lose track of a CI pipeline?

> Source: <https://dev.to/muhammadowaiswarsi/dependency-bench-at-what-size-do-llms-lose-track-of-a-ci-pipeline-11l3>
> Published: 2026-10-11 11:18:19+00:00

*This is a submission for the [Kaggle Benchmarking Challenge](https://dev.to/challenges/kaggle-2026-09-23)*

Every engineer has stared at a red CI run and asked: *why did `deploy-prod` get skipped when `build-api` only flaked
once?* Answering that means tracking dependencies, trigger rules, retries, timeouts and execution order across a

whole graph. AI agents are increasingly asked to do exactly this: debug pipelines, fix lockfiles, decide what to

rebuild. So I wanted to measure **whether models can actually reason about dependency graphs, and at what size they stop being able to.**

**Dependency Bench** has 60 generated tasks: 5 families × 4 sizes (S → XL) × 3 instances.

| Family | The model must answer | 
|---|---|
| Failure propagation (YAML) | The final conclusion (success / failure / skipped) of every job in a CI workflow with `needs` ,`if: always()` /`failure()` ,`continue-on-error` , retries and timeouts | 
| Failure propagation (prose) | The same pipelines, described in shuffled plain English the way a teammate would explain them | 
| Timing | The finish minute of every job and the total duration, when only 2–4 runners are available | 
| Version resolution | One version per package satisfying every constraint, or `UNSAT` . "Take the latest" never works | 
| Incremental rebuild | Which Makefile targets rebuild after some files change, where some targets produce byte-identical output and stop the change from spreading | 

Sizes go from **10 to 200 jobs/targets** and from **5 to 22 packages**.

**Why it's trustworthy.** No LLM judge is involved. Every task comes from a seed, and the answer key comes from a

small reference simulator that implements the rules written in the prompt. Grading is exact: one wrong job anywhere

fails the task. I also record **partial credit** (share of jobs/targets correct) to see how *close* a failing answer

was. A perfect "oracle" model scores 100% through the same pipeline, which checks the grader end to end.

Here's a small rebuild task. Can you solve it?

```
out/http/schema: src/io/config.yaml
out/json/image:  out/http/schema
out/ui/test:     src/store/assets.json src/proto/config.yaml
out/metrics/lib: src/core/main.c out/ui/test
```

*Changed:* `src/io/config.yaml`, `src/store/assets.json`. *Identical output:* `out/http/schema`.

(`out/http/schema` rebuilds, but its output is unchanged, so `out/json/image` does **not** rebuild.

`out/ui/test` and `out/metrics/lib` do.)

I ran the full suite locally through the OpenAI API at each model's default reasoning effort, with one tier of each

size so the curve has a top, a middle and a bottom:

Each task is a single user message: no system prompt, no tools, no code execution. The model has to *reason*, not

write a topological sort. On Kaggle, the same benchmark also runs on Claude models (Opus 5.5, Sonnet 5.5, Haiku 5.5)

next to these three.

| Model | Exactly right | Partial credit | S | M | L | XL | 
|---|---|---|---|---|---|---|
| gpt-5.6-sol | **57/60** | 99% | 15/15 | 15/15 | 14/15 | 13/15 | 
| gpt-5.6-terra | **46/60** | 97% | 15/15 | 13/15 | 12/15 | **6/15** | 
| gpt-5.4-mini | **4/60** | 45% | 2/15 | 1/15 | 0/15 | 1/15 | 

**1. Accuracy falls off with size, not with difficulty of the rules.** terra is perfect on every family at size S, and

the rules don't change between S and XL. Only the graph gets bigger. Yet it drops to 6/15 at XL. The model knows the

rules; it loses track of them over a long chain.

**2. "Almost right" is the typical failure, and in CI that's the dangerous kind.** terra's failed 200-job answers

still got ~98% of jobs right. Its typical mistake is calling a job `failure` when it was actually `skipped`, mixing up

"this job broke" with "this job never ran because something upstream broke". That's exactly the kind of answer that

sounds convincing in a postmortem and is still wrong.

**3. The same pipeline is harder in prose.** On 200-job pipelines, terra scored **3/3 when the pipeline was YAML and 0/3 when the identical pipeline was described in English**. Structure is doing a lot of the model's work. Agents

**4. Incremental rebuild was the only family that beat the strongest model.** sol was perfect on failure propagation,

timing and version resolution at every size, but scored 2/3 on L and 1/3 on XL rebuilds. Its mistakes **over-propagate**:

it marks targets as rebuilt even when none of their inputs changed. In one case its own answer said a target's only

prerequisite was *not* rebuilt, yet listed the target as rebuilt anyway. "Identical output stops propagation" (the

idea behind restat/early cutoff in real build systems) seems to be the rule models apply least reliably.

**5. Not reasoning is the cliff.** gpt-5.4-mini answered in ~2 seconds with ~400 output tokens per task (sol: ~25 s,

~2,100 tokens) and got 4/60. On version resolution it declared `UNSAT` ("impossible") on 4 of the 8 puzzles that

*did* have a solution. That's the worst outcome for a dependency solver: giving up confidently.

**What surprised me:** how *good* the top model is. My first pilot was 5 deliberately tricky mid-size tasks, and sol

solved all of them. The interesting signal only appeared once I stopped hand-writing hard cases and started **scaling the same rules up**, which is also what happens in real monorepos.

**What I'd measure next:**

`if:` defaults, `needs` with matrix jobs, pip/npm resolvers.
Built with Python. The task generator, reference simulator, exact grader and results viewer are all deterministic

from a seed, so the suite can be regenerated at any size to stay ahead of contamination.
