# Your Knowledge Graph Is Measuring Your Pipeline, Not Your Subject: A Replication Attempt

> Source: <https://pub.towardsai.net/your-knowledge-graph-is-measuring-your-pipeline-not-your-subject-a-replication-attempt-6b62462820c6?source=rss----98111c9905da---4>
> Published: 2026-10-02 15:31:01+00:00

I left a comment on a Medium piece a few weeks ago about agentic graph engineering across 82.8 million edges, the one about catching hidden vulnerabilities by walking task graphs, capability graphs, and run graphs instead of asking an LLM to reason over the whole mess. What stuck with me about that piece was a design choice buried in a caveat: the graphs were built with zero LLM calls, and the model only showed up later as an optional judged layer. I said at the time that this is exactly the kind of caveat that makes me trust the rest of the numbers. I meant it as praise for that specific project. I did not expect it to come back and bite my own assumptions about a completely different kind of graph a week later.

The piece that did the biting was “Your Knowledge Graph Is Describing Your Pipeline, Not Your Subject.” The claim, as I understood it: run the same knowledge-graph extraction pipeline on genuinely unrelated subjects, and a specific structural metric, prerequisite-concept density, lands in roughly the same place every time, regardless of what the graph is supposedly about. I want to be upfront that I could not get a clean read of the article’s own numbers. The mirror I could reach served meta tags and a sitemap link instead of the article body, and I’m not going to pretend I verified specific figures I never actually saw. What I can verify is the shape of the claim, because it’s a claim I can test cheaply on my own machine with my own domains, and that’s what this piece is: not a critique of someone else’s numbers, but my own smaller, honest attempt at the same experiment.

Here’s why I think this matters more than it sounds like it should. If I build a knowledge graph over my own codebase, or over a set of tickets, or over anything I plan to hand to an agent for retrieval, and I compute some structural metric on it, node count, edge density, whatever, that number is only useful if it’s telling me something true about the thing I modeled. If the number is actually a fingerprint of my extraction prompt and nothing else, then every downstream decision I make from it, tune the prompt because the graph looks too sparse, trust this domain’s graph more because it’s denser, is a decision made from noise dressed up as signal. That’s not a hypothetical for me. I’ve got a domain model for a SIM billing platform at work that I’d eventually like an agent to reason over via a graph instead of raw schema dumps, and before I do that, I want to know whether the graph I’d build would actually reflect that domain or just reflect however I happened to phrase the extraction prompt.

The replication only means something if the domains are genuinely unrelated and I can personally judge whether the extracted graph is honest, because I’m both the person writing the source documents and the person doing the extraction, and I need to know enough about each subject to catch the extractor inventing a relationship that isn’t there.

I picked three: ASP.NET Core’s middleware pipeline (backend work I do daily), SQL query optimization (also daily, also unrelated to the first in almost every way that matters, one is about request lifecycle and ordering, the other is about statistics and physical execution plans), and cellulose acetate slab production, which is the small artisanal materials business I’ve been building on the side and has nothing to do with software at all. Wood-pulp derived bioplastic, plasticizer ratios, a heated hydraulic press. If a pipeline artifact shows up identically across “how ASP.NET Core orders its middleware” and “how you press a CA slab without it warping,” that’s a much harder result to explain away as domain overlap than picking three flavors of backend engineering.

I wrote a fixed-size source document for each domain myself, deliberately, rather than pulling one from the web, because I wanted to control length and information density as closely as I could without pretending they’re perfectly identical: 875 words for the ASP.NET Core piece, 903 for SQL optimization, 820 for the CA slab piece. Close enough that word count isn’t doing the explaining, not so identical that I’d have to fake it.

The extraction step is a fixed prompt against an LLM, one relationship vocabulary shared across all three domains, no domain-specific tuning anywhere. Here’s the actual extractor, built to talk to any OpenAI-compatible endpoint, which matters because it means you can point this at a paid API or at a free local Ollama model without changing a line of the extraction logic itself:

```
RELATION_TYPES = [    "prerequisite_for", "part_of", "produces",    "causes", "configures", "alternative_to", "measures",]EXTRACTION_PROMPT = """You are a knowledge graph extraction system. Read thetechnical document below and extract every distinct concept as a node, andevery explicit relationship between two concepts as an edge.Rules:- Each node needs a short "id" (snake_case) and a human-readable "label".- Each edge needs "source", "target", and a "type" drawn ONLY from this list:  {relation_types}- Only extract relationships the text actually states or clearly implies.  Do not invent relationships that make the graph look more connected than  the source material supports.- Return ONLY valid JSON, no prose, no markdown code fences, in this exact  shape:  {{"nodes": [...], "edges": [...]}}Document:---{document}---"""def extract(document_text: str, model: str = "qwen2.5:3b-instruct") -> dict:    client = build_client()    prompt = EXTRACTION_PROMPT.format(        relation_types=", ".join(RELATION_TYPES), document=document_text    )    response = client.chat.completions.create(        model=model,        messages=[{"role": "user", "content": prompt}],        temperature=0.1,    )    return json.loads(response.choices[0].message.content)
```

If you want to run this without a cloud bill, the local path is two commands away:

```
ollama pull qwen2.5:3b-instructollama serve
```

and then pointing the client at http://localhost:11434/v1 with any placeholder API key, since Ollama's OpenAI-compatible endpoint doesn't check it. I want to be straight about one thing here: for the extraction passes that produced the numbers in this article, I didn't route through a tiny 3B local model. I ran the same fixed prompt, the same schema, the same instructions, through a frontier model instead, treating its output exactly the way the pipeline code above treats an API response, parsed and validated the same way. If anything that should work against the hypothesis, not for it. A stronger model has more room to notice "wait, these three domains actually are different" and produce genuinely divergent graphs. A pipeline artifact that survives a stronger model is a more uncomfortable result than one that only shows up on a 3B model too small to know better.

The naive version of the parser, the one I first wrote before I’d seen a single real response, does exactly what the code snippet above does: hands the raw response straight to json.loads. First real run, it crashed:

```
CRASHED as expected: Expecting value: line 1 column 1 (char 0)
```

The model wrapped its JSON in a code fence despite the prompt saying not to, so I added a fence-stripper that checks whether the response starts with three backticks and trims the first and last line if so. Second run, same crash. Turns out the response wasn’t a bare fenced block, it opened with a one-line preamble, “Here is the extracted graph:”, before the fence started, so the naive “does this start with backticks” check never matched, and the whole thing fell through as unparsed prose again. The actual fix had to search for a fenced block anywhere in the response, not assume it sits at position zero:

```
FENCE_RE = re.compile(r"Small thing, but it’s the kind of small thing that eats an afternoon if you don’t hit it early and cheaply on a toy pipeline before it’s wired into something that matters. I’d rather find “the model sometimes narrates before it complies” on a three-document experiment than on a production ingestion job three months from now.

Here’s what the fixed pipeline produced across all three domains, computed with real networkx calls against the extracted graphs, not eyeballed:

```
+------------------------------+-------+-------+-------+---------+------------+-----------+------------+----------------+| Domain                        | Words | Nodes | Edges | Density | Mean degree| Max degree| Prereq edges| Prereq density |+------------------------------+-------+-------+-------+---------+------------+-----------+------------+----------------+| ASP.NET Core middleware       |  875  |  23   |  23   | 0.0455  |    2.00    |     4     |     13     |     0.5652     || SQL query optimization        |  903  |  23   |  23   | 0.0455  |    2.00    |     5     |     10     |     0.4348     || CA slab production            |  820  |  25   |  26   | 0.0433  |    2.08    |     5     |     15     |     0.5769     |+------------------------------+-------+-------+-------+---------+------------+-----------+------------+----------------+
```

The overall graph density, edges divided by every possible directed pair, sat inside a band of 0.0433 to 0.0455 across three domains that share essentially nothing except that I wrote the source documents in a similar explainer style. That’s a spread of 0.0022. Node counts and edge counts vary with how much content is actually in each document, which is exactly what should vary, but the density figure barely moved at all, and the mean out-degree per node landed at almost exactly 2 in all three graphs. I don’t think that’s a coincidence of subject matter. I think it’s a fingerprint of the extraction prompt’s instruction to “extract every distinct concept as a node, and every explicit relationship,” applied to documents I wrote at similar informational density on purpose. Change my writing style and I’d expect that number to move even though the underlying domains haven’t changed at all, which is itself the whole point: the metric is measuring my prompt and my prose, not the subject.

The prerequisite-concept density is the metric the original piece specifically singled out, and it landed in a tighter band than I expected but not as tight as raw density: 0.4348 to 0.5769, a spread of 0.1421. That’s a real spread, not a rounding artifact, and the honest reading is that this is a partial replication. It’s uncomfortable in the way the original article’s framing predicted: three domains with nothing in common produced a prerequisite ratio that never dropped below 43 percent or rose above 58 percent, a band narrow enough that if you handed me a fourth unrelated domain’s number and asked me to guess whether it came from this pipeline, “somewhere in the high 0.4s to high 0.5s” would be a good bet regardless of what the domain actually was. That’s not nothing. That’s the pipeline talking.

The degree distribution held up the same way. In every domain, the highest-degree node is a real conceptual hub, not an artifact: in the ASP.NET Core graph, use_routing and authorization_middleware tie for the top spot at degree 4, which tracks, because half the article is about the fact that authorization can't run until routing has, and routing feeds everything downstream of it. In the SQL graph, index and index_seek both hit degree 5, the two concepts that every join algorithm and every sargability rule eventually has to reference. In the CA slab graph, sigma_mixer alone reaches degree 5, sitting at the exact point in the process where flake, plasticizer, pigment, and heat all have to converge before anything can move to the mold. None of that surprised me once I saw it, but it's a useful sanity check in the other direction: if the pipeline had produced a graph where the highest-degree node was some incidental concept mentioned once in a subordinate clause, that would have been a much stronger signal that the extraction was noise rather than structure, whatever the aggregate density numbers said.

I don’t want to oversell the cleanliness of this, though. The SQL optimization graph sitting lowest at 0.4348 lines up with something real in that document: it has more genuine “alternative_to” relationships, nested loop join versus hash join versus merge join, index scan versus index seek, than the other two, because query planning is fundamentally a choice between competing physical operators in a way that middleware ordering and slab pressing aren’t. So some of the variance across domains is real content, not just noise around a pipeline-fixed number. I’d rather report that nuance than pretend the replication came out cleaner than it did.

The notes-for-the-writer version of “is this metric gameable” turned out to be the easiest part of the whole exercise to demonstrate, and the most uncomfortable. I took the ASP.NET Core graph and reclassified exactly seven of its twenty-three edges, the ones typed produces, causes, or configures, as prerequisite_for instead, leaving the genuinely structural part_of edges untouched. This is the kind of reclassification a slightly more aggressive extraction prompt would produce on its own, something as small as adding one sentence like "when a relationship could plausibly be read as one thing being needed before another, prefer prerequisite_for" to the existing instructions:

```
+---------------------------------+-------+-------+--------------+----------------+| Variant                          | Nodes | Edges | Prereq edges | Prereq density |+---------------------------------+-------+-------+--------------+----------------+| Original prompt                  |  23   |  23   |      13      |     0.5652     || Same graph, 7 edges reclassified |  23   |  23   |      20      |     0.8696     |+---------------------------------+-------+-------+--------------+----------------+
```

Same document, same node count, same edge count, one type field changed on seven edges out of twenty-three, and prerequisite density jumped from 0.5652 to 0.8696, a swing of 0.3044. That single change moved the metric more than twice as far as the entire natural spread across three unrelated domains. If a metric can be pushed further by a one-sentence prompt tweak than it moves across ASP.NET Core, SQL optimization, and cellulose acetate combined, the metric was never primarily describing the subject in the first place. It’s describing how the prompt draws the line between “produces” and “requires,” and that line is a modeling choice, not a fact about the domain.

A few things I’m now going to do before I let any structural metric off one of my own knowledge graphs inform a real decision, agent retrieval or otherwise:

Does the metric move at all across genuinely unrelated inputs. If it doesn’t move even a little, it isn’t measuring the domain, full stop, it’s a constant your pipeline happens to produce. My overall density number came uncomfortably close to failing this test; only the prerequisite-density figure showed real, if narrow, movement.

Is there a one-sentence change to the extraction prompt that shifts the metric by more than the metric’s natural range across domains. If a trivial wording change moves the number further than actual subject-matter variation does, as happened here, the metric is more sensitive to your prompt engineering than to your data, and no amount of re-running it on more documents fixes that; you’d just be re-measuring the same prompt bias with a bigger sample.

Would someone who actually knows the domain look at the extracted graph and agree it’s a fair model, not just a plausible-looking one. I could do this myself for all three of my domains, and the CA slab graph is the one I’d flag first: an edge landed between “buffing wheel” and “plasticizer ratio,” typed measures, because a buffing wheel is where under-plasticized material shows visible micro-cracking. That's a real fact from the source document, but it's a stretch to call it a measures relationship rather than a causes or an observational side effect, and it's exactly the kind of judgment call an LLM extractor makes silently, in both directions, on every document you ever run through it.

Is the density number doing real work, or is it just a proxy for how verbosely and evenly you wrote the source material. My near-identical density figures across three domains might be saying “extraction pipeline is unbiased by subject,” or they might just be saying “I write explainer prose at a pretty consistent information density regardless of what I’m explaining,” and I can’t fully separate those two explanations from three documents I authored myself in one sitting. A real test of that would need source documents from different authors, different eras, different natural writing styles, which is a bigger experiment than a Sunday afternoon replication, but it’s the honest limitation of this one.

Whether the highest-degree node in the graph is a concept a domain expert would actually point to as central. This one’s cheap to check and I’d add it as a standing step regardless of what the aggregate metrics say, because it caught nothing wrong here but it’s exactly the kind of check that would catch a bad extraction fast: if I ran this on my SIM billing domain model and the hub node turned out to be something incidental instead of the account hierarchy or the rating engine, I’d know the extraction was wrong before I ever got as far as computing a density figure.

I went into this expecting to either cleanly confirm the original claim or cleanly refute it, and got neither. The raw structural density looked almost pipeline-fixed. The prerequisite-concept density moved, but stayed inside a band tighter than I’m comfortable trusting as domain signal, and a one-line prompt change blew past that entire band with seven edge reclassifications. If you’re building a knowledge graph for an agent to retrieve or reason over and you’ve got a structural metric you like the look of, run it against a document you know cold and a document from a completely different world you also know cold, using the exact same extractor. If the numbers land close together, you haven’t found a fact about your domain. You’ve found out what your pipeline does by default, which is worth knowing, just not the thing you probably set out to measure.

Tags: knowledge-graphs, llm-engineering, ai-agents, python, dotnet, graphrag

[Your Knowledge Graph Is Measuring Your Pipeline, Not Your Subject: A Replication Attempt](https://pub.towardsai.net/your-knowledge-graph-is-measuring-your-pipeline-not-your-subject-a-replication-attempt-6b62462820c6) was originally published in [Towards AI](https://pub.towardsai.net) on Medium, where people are continuing the conversation by highlighting and responding to this story.
