Extracting metadata with an LLM is data cleaning A technical blog post argues that non-reproducible LLM metadata extraction should be treated as a data-cleaning and statistical problem rather than a determinism problem, citing He (2025), who obtained 1000 identical completions out of 1000 at temperature zero on an open-weight model served with vLLM using batch-size-invariant kernels. The post recommends pinning the model version, recording every output, and caching by input hash for provenance, while running many samples to capture the output distribution and calibrating it against human labels. It notes that even at temperature zero hosted models are non-deterministic because floating-point addition is not associative and server-side batching changes summation order, an issue previously attributed to mixture-of-experts routing (Chann 2023). Suppose you have a few thousand documentation files, and you are building a pipeline that uses an LLM to extract metadata from them: a title, an instrument, whether the file is about climate change. And you find the outputs aren’t reproducible: run it twice and some of the answers change. How should you approach this? Or, if you are an RSE and a researcher brings you this, what would you advise? As in my last post ../LLM/refactoring.html , I’d start from the XY problem https://xyproblem.info/ . They ask about reproducibility, X. What they almost certainly need is metadata they can trust, Y. And here Y is easier to find than usual, because the context leaves only a few possibilities. I think reproducibility and uncertainty are two separate problems. Both can be solved, but they are solved differently, and the second one is the one that matters. Seen that way, I’d treat it as data analysis, specifically the data cleaning step or wrangling, or munging, depending on who you ask , which is a statistical problem, and an LLM is one tool to do it. But first, reproducibility, since that is what they asked about. Yes, we can, but that’s beside the point. Here’s how anyway. Some people say that an LLM is stochastic in nature, so it can’t be deterministic. The first half is true, the second is not. An LLM gives a probability for every possible next token, and the randomness comes from sampling from it. Turn the temperature to zero and it takes the most likely token at every step instead, a greedy maximum likelihood estimate, which should be deterministic. However, it is not this simple. Floating-point addition is not associative, so adding the same numbers in a different order can give a different answer, and a deterministic result needs a deterministic order of execution. On a hosted model, that order is not up to you. The server batches your request with other people’s, and how many there are changes the order of the sums inside the model. Earlier, people blamed the mixture-of-experts architecture instead, where tokens from different requests in a batch compete for the same experts Chann 2023 . Either way, what else is in the batch leaks into your answer, so even at temperature zero a hosted model is not deterministic. But it can be made so. With kernels that keep the same order whatever the batch size, He 2025 got 1000 identical completions out of 1000 at temperature zero, on an open-weight model served with vLLM. So if we want reproducible output, we can have it. Then the question is, does reproducibility here mean what we want? I’d argue the reruns that disagree are the extraction’s uncertainty showing, and they are reporting the symptom as the disease. A deterministic model hides the uncertainty rather than removing it: it always returns its most likely answer, and a reproducible wrong answer arrives with a certificate. So, to make the research reproducible, pin the model version, record every output, and cache it by a hash of the input.