{"slug": "mitigating-memorization-in-llms", "title": "Mitigating memorization in LLMs", "summary": "Jane Street ML Research intern Monte Bohde found that divergence decoding — inference-time steering in logits space using two smaller auxiliary models trained on pre-2015 and pre-2026 data — reduced pre-cutoff performance on a cricket match preview dataset without harming post-cutoff performance, evidence it mitigates LLM memorization. Bohde's attempts to distill the steered model back into the large model failed, with the distilled model consistently underperforming both the inference-time divergence-decoding model and the original. The project also found that larger models performed better and were more memorized, and that finding a model clearly demonstrating memorization was surprisingly difficult due to data quirks including a distributional shift from men's international matches to more varied cricket after the knowledge cutoff.", "body_md": "*The following is part of a series of posts about 2026 summer intern projects – for more, see [“What the interns have wrought, special jumbo 2026 edition”](https://blog.janestreet.com/wrought-2026/)*\n\nAt Jane Street we spend a lot of effort trying to predict the future, and we’d like to know if LLMs are any good at it. To do this, we might assemble a collection of queries with known answers—say, predicting which team will win the Super Bowl in a given year—and then measure how well an LLM does.\n\nWhen trying to construct such a benchmark, a problem immediately arises: LLMs are pretrained on a colossal amount of text that contains answers to many of the questions in our dataset. A model may perform well on a historical benchmark by consulting its memorized knowledge of past outcomes, but fail miserably when presented with new scenarios in production.\n\nThis summer, Monte Bohde, a summer intern in our ML Research group, had two tasks: first, measure when LLMs are succeeding merely by regurgitating memorized facts from pretraining, and second, find ways to prevent this behavior.\n\n## Finding a dataset and model\n\nObviously our interest here is in predictions that impact financial markets, but the same ideas can be explored (and discussed on our blog) with lower-stakes data sets. A lot of the experiments that Monte ran were on a dataset of cricket match previews—little writeups setting the stage for an upcoming match. He wanted to see if he could find an open model that clearly demonstrated memorization. That is, when you ask the model to predict the outcome of the match given the match preview text, does it do a lot better for matches that took place before its pretraining knowledge cutoff than those that happened after?\n\nFinding such a model was surprisingly difficult. There were quirks in the data, for instance a distributional shift over time that Monte had to tease out: before the cutoff, the previews were mostly about men’s international matches, whereas after the cutoff they became much more varied. Some models performed little better than chance, which made them poor candidates for the analysis. Others couldn’t even consistently format their predictions.\n\nTo assess a model’s level of memorization, he asked three questions:\n\nH1: Is the model any good at predicting cricket matches, i.e., is it additive to the\nbaseline post-cutoff?\n\nH2: Is it better pre-cutoff than post-cutoff?\n\nSharpness: Does it seem more overconfident pre-cutoff? (A qualitative proxy for\nmemorization)\n\nBigger models in general performed better and were also more memorized, which made them suitable for the study. The question then became how to mitigate the effect of memorization. As is typical of ML research projects, Monte explored a number of ideas, only some of which panned out.\n\n## Divergence decoding reduces memorization without affecting out-of-sample performance\n\nDivergence decoding attempts to remove later years of knowledge from frontier models by doing inference-time steering in logits space. Specifically, given two smaller auxiliary models trained only on data pre-2015 and pre-2026 (or some knowledge cutoff of your large model), you steer your frontier model as:\n\nThat way you end up with a frontier model that behaves as though it has knowledge of\nevents *only* before 2015, without having to train such a full-sized model yourself. To a\nfirst approximation you’ve wiped its memory for the years spanning your two auxiliary\nmodels.\n\nOn the cricket dataset, divergence decoding reduced pre-cutoff performance without harming post-cutoff performance, evidence that it mitigates the effect of memorization.\n\n## Distillation\n\nIt is also in theory possible to distill the steered model back into the large model and obtain weights for a frontier model that lacks knowledge of recent years. Monte tried to do this distillation, but the resulting model consistently underperformed both the inference-time divergence-decoding model and the original.\n\nMonte’s hypothesis for what went wrong was the data mix during distillation. He used a subset of CommonCrawl tokens from 2019-2024, but in order for the distillation to actually work, your model needs to see sufficiently many task-related tokens. Cricket is somewhat obscure and probably not very well represented in this set.\n\n## Prompt engineering and synthetic rewrites\n\nIs there a way to re-cast your data set in a form such that the model is less likely to reach for memorized knowledge? Monte tried a few tacks. One was to rewrite the cricket match previews “in basketball terms.” Maybe then the model would reason about what was in the preview text itself, and what that implied about win probabilities, instead of relying on specific cricket match outcomes.\n\nThat didn’t work, but what did work, surprisingly, was to rewrite the previews to be in “podcast form” – that is, to appear to be transcripts from a podcast discussing the upcoming match. In this rewrite Monte also removed player quotes, which seemed to be low signal. With this slightly different presentation of the same data, the model performed better both pre- and post-cutoff but had a larger improvement post-, implying reduced memorization.\n\n## Natural language autoencoder\n\nAnother method is to try to directly catch the model when it is regurgitating facts. To\ntry this, Monte trained a [Natural Language\nAutoencoder](https://transformer-circuits.pub/2026/nla/) (NLA). An NLA contains an\nencoder model, which maps a model’s activations to a string of text, and a decoder model,\nwhich maps that text back to activations. The encoder and decoder are jointly trained via\nreinforcement learning to accurately reconstruct activations from the produced text.\n\nMonte found that while the NLA decodings didn’t directly explain what the model was thinking about, the sentiment of the NLA decoding did contain a lot of information. For example, the relative count of how many times each team was mentioned in the cricket task was highly correlated (corr ~0.93) to the models’ final prediction:\n\nUnfortunately this effect was not different pre/post cutoff, so it was not a useful signal to identify/remove memorization.\n\nModels were also prone to hallucination in the NLA decoded text. For example, on the cricket task, the model frequently mentioned the team that it thought would win and then filled in the other team with a random placeholder, usually India.\n\n### Diff decoding\n\nAt one point, Monte tried fine-tuning the base model on the narrow task of predicting cricket match outcomes from the match previews, in the hopes of inducing it to memorize more often, while maintaining its performance post-cutoff. That had mostly lukewarm results, but an interesting finding was that you can use an autoencoder to probe the difference between the two models. In particular you can decode the “diff embedding,” which is just the FT model embedding minus the base model embedding. On the cricket dataset this decodes almost exactly into a description of the FT objective, i.e. the diff decode explicitly mentions making probability based predictions about cricket:\n\nIn other words, fine-tuning introduces a “task vector” for predicting probabilities along with a piece that’s correlated to the linear head for what the probability is.\n\n### Sequence effects\n\nMonte created a synthetic reasoning dataset to investigate NLA decodings in a high-signal setting. The rows were of the form:\n\nHe found that if he applied the NLA unembedding matrix to intermediate embeddings, he saw that the model made up its mind around the middle layer and then in the last few layers started to think about formatting tokens. This suggested that he could remove those last layers. (This turned out to be correct, and removing the unnecessary layers sped up training by 1.5x.)\n\nIn general, he found that the NLA readouts and linear head predictions varied across the sequence dimension—and in particular the model didn’t seem to make up its mind until reading the very last token (!).\n\n## Conclusions\n\nLike many ML research projects, Monte’s work here was exploratory. It didn’t result in a de-memorizer, but did help identify the most promising avenues for future research. In particular the success of divergence decoding was suggestive, as was some of the structure that he found using NLA.\n\n## Looking forward to next summer…\n\n*If you’re interested in doing work like this, consider applying! You can find more details here: [Jane Street Internships](https://www.janestreet.com/join-jane-street/internships/). Applications for our 2027 Summer ML Research internship are now open!*", "url": "https://wpnews.pro/news/mitigating-memorization-in-llms", "canonical_source": "https://blog.janestreet.com/mitigating-memorization-in-llms/", "published_at": "2026-09-30 00:00:00+00:00", "updated_at": "2026-09-30 23:18:28.142472+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research"], "entities": ["Jane Street", "Monte Bohde"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/mitigating-memorization-in-llms", "markdown": "https://wpnews.pro/news/mitigating-memorization-in-llms.md", "text": "https://wpnews.pro/news/mitigating-memorization-in-llms.txt", "jsonld": "https://wpnews.pro/news/mitigating-memorization-in-llms.jsonld"}}