# Extracting metadata with an LLM is data cleaning

> Source: <https://blog.kolen.dev/LLM/metadata-extraction.html>
> Published: 2026-10-02 00:00:00+00:00

Suppose you have a few thousand documentation files, and you are building a pipeline that uses an LLM to extract metadata from them: a title, an instrument, whether the file is about climate change. And you find the outputs aren’t reproducible: run it twice and some of the answers change. How should you approach this? Or, if you are an RSE and a researcher brings you this, what would you advise?

As in [my last post](../LLM/refactoring.html), I’d start from the [XY problem](https://xyproblem.info/). They ask about reproducibility, X. What they almost certainly need is metadata they can trust, Y. And here Y is easier to find than usual, because the context leaves only a few possibilities.

I think reproducibility and uncertainty are two separate problems. Both can be solved, but they are solved differently, and the second one is the one that matters. Seen that way, I’d treat it as data analysis, specifically the data cleaning step (or wrangling, or munging, depending on who you ask), which is a statistical problem, and an LLM is one tool to do it. But first, reproducibility, since that is what they asked about.

Yes, we can, but that’s beside the point. Here’s how anyway.

Some people say that an LLM is stochastic in nature, so it can’t be deterministic. The first half is true, the second is not. An LLM gives a probability for every possible next token, and the randomness comes from sampling from it. Turn the temperature to zero and it takes the most likely token at every step instead, a (greedy) maximum likelihood estimate, which should be deterministic.

However, it is not this simple. Floating-point addition is not associative, so adding the same numbers in a different order can give a different answer, and a deterministic result needs a deterministic order of execution. On a hosted model, that order is not up to you. The server batches your request with other people’s, and how many there are changes the order of the sums inside the model. Earlier, people blamed the mixture-of-experts architecture instead, where tokens from different requests in a batch compete for the same experts (Chann 2023). Either way, what else is in the batch leaks into your answer, so even at temperature zero a hosted model is not deterministic.

But it can be made so. With kernels that keep the same order whatever the batch size, He (2025) got 1000 identical completions out of 1000 at temperature zero, on an open-weight model served with vLLM. So if we want reproducible output, we can have it.

Then the question is, does reproducibility here mean what we want? I’d argue the reruns that disagree are the extraction’s uncertainty showing, and they are reporting the symptom as the disease. A deterministic model hides the uncertainty rather than removing it: it always returns its most likely answer, and a reproducible wrong answer arrives with a certificate.

So, to make the research reproducible, pin the model version, record every output, and cache it by a hash of the input.<sup>1</sup> Then the downstream analysis reruns, up to the equivalence that matters: for LLM output, the same answer up to formatting, just as a numerical pipeline reproduces up to floating point (see my [talk on reproducibility](../RSE/UoE/2025-11-26-reproducibility-article.html)). If what they need is provenance for a paper, that is it: the model and its version, the prompt, the settings, and the recorded outputs as a data artefact.

This doesn’t contradict the uncertainty above. What the pin defines is an equivalence class, and if the problem is probabilistic in nature, the equivalence is, roughly, up to the statistical distribution. So rather than one run, a point estimate, standing in for the whole class, run it many times, a Monte Carlo of sorts, and keep the sample distribution. Then calibrate it against what a human says.

Whether the problem is probabilistic depends on the field, on what kind of answer it has. So here is the flowchart I’d walk them through, and the rest of this post is why.

Where the field is regular, say an XML field or a convention every file follows, an LLM makes writing a bespoke parser cheap, which it never used to be. At a few thousand files the pre-LLM route works too: a regex with every edge case handled, as long as the files share a convention. Either way, run the parser, not the LLM.

The human checks the parser if it is short, or its output on a few dozen files if not, and an LLM can flag suspicious outputs for the human to look at. The answer is as close to perfect as the checking makes it. This is the same principle as in the refactoring post: keep what a human must check by eye as small as possible.

If the field isn’t regular but there is a right answer and it is in the file, I’d treat the LLM output as a measurement with an error rate. Then the treatment is data cleaning as it was done before LLMs, with the models as annotators: several models, several runs each, consensus where they agree, and a human on the disagreements.

The one thing new here is that the models’ errors are correlated, because they are trained on much the same data. So their agreement overstates our confidence, and a human should check a sample of the agreements too, not only the disagreements.

Either there is no ground truth (“is this file about climate change?”, where two people would disagree), or it is not in the document (a title that never appears in the file). Then the outcome is inherently probabilistic, and aiming at a deterministic answer is the wrong angle. It is a data analysis problem: quantify the uncertainty and make sure it is unbiased.

First, the labels. Tags and booleans become one problem once tags have a vocabulary. To get one, let the LLM tag each file in a sample freely, and take the union of all the tags. Many of them will mean the same thing, and finding which ones do is what embeddings are good at, the same similarity search that RAG is built on. Collapse each group into one tag, and what is left is a smaller set of tags with disjoint meanings. The final set is the scientist’s decision, because it is part of their research question, and it is the set every file then gets tagged against again. After that, both are classification over a known set, plus three labels I’d keep apart:

Second, ask for probabilities, not labels. What we want here is what [TypeSafe](https://docs.typesafe.ai/introduction) calls a “System One” model, after Kahneman’s fast, intuitive System One: unstructured text in, a probability over a typed, closed set out, rather than generated text. Their *Jev*, in early access since September 2026, is the first purpose-built one I know of, and I think it is a promising tool for exactly this. But the idea doesn’t depend on it. Any open-weight LLM can be repurposed to do it, because it already computes log-probabilities, and what it lacks is the output structure. Constrain its output to the closed set, or score each label’s log-probability and normalize over the set. Failing that, use the frequency across repeated runs. As of September 2026 Jev is API only, so the repurposed open model is also the one that can be self-hosted, if the files can’t leave the institution.

Whichever it is, check the calibration on the hand-labelled sample: when it says 80%, is it right 80% of the time? A reliability diagram shows it, and if it is off, rescale. I’d do this even when the vendor says the model is calibrated, Jev included. Calibration depends on the data, so a claim that a model is calibrated is a claim about the data it was measured on, and a model can be calibrated on its training data and overconfident on ours (AgentConn 2026). I think saying that an ML model’s output probability is calibrated, full stop, is too strong a statement to be possible. An [independent test of Jev](https://github.com/scienthoon/jev-ood-calibration) on a task it cannot have seen found it underconfident on yes/no questions and overconfident on choices and scores, so even the direction of the correction depends on the question. And at worst, if it was calibrated all along, we have shown that it is, which is due diligence we would want on record anyway.

Third, aggregate before deciding. For “what fraction of the files are about X?”, sum the probabilities. Taking the most likely label per file and then counting biases the total. Take the most likely label only where a per-file decision is really needed.

And last, bias. The errors can follow the files’ source, age or length, so the aggregate can be biased even at high accuracy. Check the error rate per group. Prediction-powered inference (Angelopoulos et al. 2023) combines the small labelled sample with the model’s predictions to give unbiased estimates with valid confidence intervals, without assuming anything about the model that made the predictions.

Every branch of the flowchart ends on a small hand-labelled sample, a few hundred files at most. That is how a few thousand files get validated without anyone reading them all: the labelled sample, the disagreements, and a sample of the agreements.

So that’s how I’d use an LLM to extract metadata. The specific tools will change. Jev is a few weeks old, and there may well be something better next year. What I think doesn’t change is the rest of it: it is data cleaning, which is a statistical problem, and the LLM is one way of making the measurement. The judgement on what the labels are, and the calibration of the statistics against what a human sees, stays with the human in the loop.

This is how Nix thinks about a reproducible environment (Dolstra 2006). [Reproducible builds](https://reproducible-builds.org/) are hard to achieve, and identical output is not a property we can assert in the wild, whatever the reason. So think about it functionally instead. If the build is a pure function, depending on its inputs and the function alone, then whatever comes out is treated as one equivalence class, addressed by a hash of the inputs. Reproducibility then means we won’t recompute it if none of those have changed. That assumes every input is captured and the function is pure, which can’t be verified, so the human in the loop is responsible for designing the pipeline so that it is. Here, that means the model version, the prompt and the settings all go into the hash along with the file.↩︎
