cd /news/large-language-models/secret-dates-in-system-prompts-under… · home › topics › large-language-models › article
[ARTICLE · art-148326] src=unite.ai ↗ pub= topic=large-language-models verified=true sentiment=↓ negative

Secret Dates in System Prompts Undermine Language Model Evaluation

A new academic collaboration between Germany, Mexico and the USA found that the hidden automatic injection of the current date into system prompts introduces non-determinism in LLM evaluation, altering model performance and reshuffling leaderboard rankings across 9 models, 6 datasets and four tasks. The paper reports performance varying by up to 6% on multiple-choice questions, 14% on mathematical reasoning and 7% on code generation, with GPT-5.1 showing accuracy fluctuations of up to 4% across three multiple-choice benchmarks during a week of testing in December 2025 even with an empty system prompt, reasoning disabled and randomness set to zero. The researchers recommend removing the date from system prompts where possible, or fixing and documenting it, to ensure fair comparisons.

read6 min views1 publishedOct 9, 2026
Secret Dates in System Prompts Undermine Language Model Evaluation
Image: Unite (auto-discovered)

[Anderson's Angle](https://www.unite.ai/series/andersons-angle/)


[Add Unite.AI to your preferred sources on Google](https://www.google.com/preferences/source?q=unite.ai)

Without the ability to benchmark Large Language Models (LLMs), it is difficult for consumers and businesses to understand what progress a model has made over recent versions, and how it stands up to its competitors:

Since LLMs are non-deterministic (i.e., they will not always produce consistent outputs given the same inputs), evaluating them is tricky. Even when researchers use identical prompts, model settings and benchmark datasets, seemingly minor differences in the execution environment can produce different answers and significantly alter performance scores.

Factors such as hardware configuration, numerical precision, inference batch size, and even the ordering of multiple-choice answers can influence results. This can make it hard to see whether a reported improvement reflects genuine progress, or merely some semi-random variation in the conditions under which the model was tested.

To a certain extent, one can account for some of these variables, or at least establish what margin-of-error they generate, so that comparative evaluation becomes meaningful. It’s important to try, since a lot of money, and a lot of reputation depends on being able to benchmark AI systems of this kind with some degree of accuracy.

However, according to new research, one particular variable can not only be destructive to benchmarks, but is also very difficult to eliminate from the prompts that define them – today’s date.

Times Change #

The new paper*, an academic collaboration between Germany, Mexico and the USA, asserts that the fact that the current date is automatically and secretly included in the system prompt of all frontier and many deployments of open-weight models means that reproducibility could be nigh-on impossible:

‘We identify a critical, often overlooked source of non-determinism in LLM evaluation: the hidden injection of the current date into system prompts.

‘Across 9 models, 6 datasets, and four tasks, this dynamic metadata alters model performance and reshuffles leaderboard rankings, surpassing the variance introduced by other system-level factors such as batch size or numerical precision.’

The researchers also note that the identified effect is larger for tasks requiring generated answers, with performance varying by up to 6% on multiple-choice questions, 14% on mathematical reasoning, and 7% on code generation – and with machine translation scores also varying significantly.

Providing example answers and  encouraging step-by-step reasoning did not resolve the problem – in fact, the latter made it worse:

The effect was also identified in proprietary models, with GPT-5.1 showing accuracy fluctuations of up to 4% across three multiple-choice benchmarks during a week of testing in December 2025. The researchers used an empty system prompt, disabled reasoning and set randomness to zero – but the date was still inserted automatically by the provider:

The researchers recommend removing the date from system prompts where possible, or fixing and documenting it, to ensure fair comparisons. However, they note that date-centric influence may be more deeply-ingrained:

‘One possible reason for the date sensitivity is that the system prompt might be fixed during supervised fine-tuning (SFT) and reinforcement learning from human feedback (RLHF), making the model brittle to any slight modification.’

To investigate the issue, the researchers tried six different wordings for the system prompt, and found that these changes affected accuracy just as much as changing the date (0.78%). Therefore the date appears to have as much influence as deliberate prompt engineering – except that it changes automatically, without the user even knowing

Other approaches were tried to mitigate the ‘current date’ problem, including few-shot learning (where the model got five example answers before responding). This helped a little, bringing the average variation down from 2.52% to 2.27%, but didn’t fix the problem.

Changes in GPU hardware, batch size, numerical precision, answer order and system-prompt wording were also tested. The biggest effects were seen with answer-order and prompt wording, which came closest to the impact of changing the date.

Last Days

It’s reasonable to expect that an LLM/VLM will understand the date from the very beginning of a chat; however, there seems to be no explicit reason to build it into the system prompt (the unseen rubric that imposes guardrails, and conditions the LLM’s behaviors in ways inaccessible to the user) when the LLM could routinely make a sub-Kb RAG call to ingest the latest date, as a minor housekeeping routine prior to engagement with the user.

No doubt other possibilities exist to resolve the matter; however, since the imposition of the current date into the system prompt has not hitherto been seen as a problem, there has presumably been little or no investigation in regard to this.

Online Dating #

For the tests, identical prompts were used for every model, with only the date in the system prompt being changed. Every day of 2024 was tested, from January 1st through to December 31st, with all other settings kept the same. Six benchmarks were used: MMLU; GPQA; and ARC-Challenge, for multiple-choice questions (scored by answer-token probability). GSM8K for step-by-step math (final answer checked); HumanEval for Python code generation (unit-tested); and WMT for English-to-German; English-to-Finnish; and English-to-Czech translation (full output evaluated). Time-dependent questions were excluded.

Nine models were tested: Llama 3.1 Instruct (8B and 70B); Gemma 3 Instruct (4B and 27B); Qwen3 (4B); Qwen3-Next (80B); Phi-4 (14B); and GPT-OSS (20B and 120B).

Accuracy (the percentage of questions answered correctly) was used to score the multiple-choice and math tests, while Expected Calibration Error (ECE) measured how well the models’ confidence matched their actual performance.

Code was checked using pass@1 (the percentage of generated code solutions that pass all tests on the first attempt); translations were scored using BLEU and chrF.

The results confirmed the researchers’ hypothesis: simply changing the date in the system prompt changed how accurately the models answered the same questions.

Along with other results shown earlier in the article, this is potentially bad news for LLM benchmarking, because a model could score better or worse depending on which day it was tested, possibly changing its position on a leaderboard, without any actual improvement or decline in its capabilities.

Conclusion #

This issue highlights the divide between the deterministic computing systems we have been used to prior to around 2023, and the very different nature of diffusion-based and similar AI systems that have evolved, and continue to evolve, since then.

Date resolution was essentially solved on January 1st 1970, but has, apparently, returned to haunt the world of computing in the form of cut-off dates, among other temporal concerns.

  • Titled ‘Dating the Model: Hidden Dates in System Prompts Affect LLM Evaluation’ First published Friday, October 9, 2026
── more in #large-language-models 4 stories · sorted by recency
── more on @gpt-5.1 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/secret-dates-in-syst…] indexed:0 read:6min 2026-10-09 · —