{"slug": "secret-dates-in-system-prompts-undermine-language-model-evaluation", "title": "Secret Dates in System Prompts Undermine Language Model Evaluation", "summary": "A new academic collaboration between Germany, Mexico and the USA found that the hidden automatic injection of the current date into system prompts introduces non-determinism in LLM evaluation, altering model performance and reshuffling leaderboard rankings across 9 models, 6 datasets and four tasks. The paper reports performance varying by up to 6% on multiple-choice questions, 14% on mathematical reasoning and 7% on code generation, with GPT-5.1 showing accuracy fluctuations of up to 4% across three multiple-choice benchmarks during a week of testing in December 2025 even with an empty system prompt, reasoning disabled and randomness set to zero. The researchers recommend removing the date from system prompts where possible, or fixing and documenting it, to ensure fair comparisons.", "body_md": "### \n[Anderson's Angle](https://www.unite.ai/series/andersons-angle/)\n\n# Secret Dates in System Prompts Undermine Language Model Evaluation\n\n[Add Unite.AI to your preferred sources on Google](https://www.google.com/preferences/source?q=unite.ai)\n\nWithout the ability to [benchmark](https://www.unite.ai/benchmarks-for-llms/) Large Language Models (LLMs), it is difficult for consumers and businesses to understand what progress a model has made over recent versions, and how it stands up to its competitors:\n\nSince LLMs are non-deterministic (i.e., they will *not* always produce consistent outputs given the same inputs), evaluating them is tricky. Even when researchers use identical prompts, model settings and benchmark datasets, seemingly minor differences in the execution environment can produce different answers and significantly alter performance scores.\n\nFactors such as hardware configuration, [numerical precision](https://www.unite.ai/rethinking-scaling-laws-in-ai-development/), [inference batch size](https://archive.is/8quL2#batchsize), and even [the ordering of multiple-choice answers](https://arxiv.org/pdf/2308.11483) can influence results. This can make it hard to see whether a reported improvement reflects genuine progress, or merely some semi-random variation in the conditions under which the model was tested.\n\nTo a certain extent, one can account for some of these variables, or at least establish what margin-of-error they generate, so that comparative evaluation becomes meaningful. It’s important to try, since a lot of money, and a lot of reputation depends on being able to benchmark AI systems of this kind with some degree of accuracy.\n\nHowever, according to new research, one particular variable can not only be destructive to benchmarks, but is also very difficult to eliminate from the prompts that define them – *today’s date*.\n\n## Times Change\n\nThe [new paper](https://arxiv.org/pdf/2609.36931)*, an academic collaboration between Germany, Mexico and the USA, asserts that the fact that the current date is automatically and secretly included in the system prompt of all frontier and many deployments of [open-weight models](https://www.unite.ai/the-rise-of-open-weight-models-how-alibabas-qwen2-is-redefining-ai-capabilities/) means that reproducibility could be nigh-on impossible:\n\n*‘We identify a critical, often overlooked source of non-determinism in LLM evaluation: the hidden injection of the current date into system prompts.* \n\n*‘Across 9 models, 6 datasets, and four tasks, this dynamic metadata alters model performance and reshuffles leaderboard rankings, surpassing the variance introduced by other system-level factors such as batch size or numerical precision.’*\n\nThe researchers also note that the identified effect is larger for tasks requiring generated answers, with performance varying by up to 6% on multiple-choice questions, 14% on mathematical reasoning, and 7% on code generation – and with machine translation scores also varying significantly.\n\nProviding example answers and  encouraging [step-by-step reasoning](https://www.unite.ai/what-is-chain-of-thought-cot-prompting-examples-benefits/) did not resolve the problem – in fact, the latter made it *worse*:\n\nThe effect was also identified in proprietary models, with [GPT-5.1](https://www.unite.ai/openai-releases-gpt-5-2-after-internal-code-red-over-googles-gemini-3/) showing accuracy fluctuations of up to 4% across three multiple-choice benchmarks during a week of testing in December 2025. The researchers used an empty system prompt, disabled reasoning and set [randomness](https://www.unite.ai/unveiling-the-control-panel-key-parameters-shaping-llm-outputs/#:~:text=1%2E%20Temperature) to zero – but the date was still inserted automatically by the provider:\n\nThe researchers recommend removing the date from system prompts where possible, or fixing and documenting it, to ensure fair comparisons. However, they note that date-centric influence may be more deeply-ingrained:\n\n*‘One possible reason for the date sensitivity is that the system prompt might be fixed during supervised fine-tuning (SFT) and reinforcement learning from human feedback (RLHF), making the model brittle to any slight modification.’*\n\nTo investigate the issue, the researchers tried six different wordings for the system prompt, and found that these changes affected accuracy just as much as changing the date (0.78%). Therefore the date appears to have as much influence as deliberate [prompt engineering](https://www.unite.ai/what-is-prompt-engineering-in-ai-why-it-matters/) – except that it changes automatically, without the user even knowing\n\nOther approaches were tried to mitigate the ‘current date’ problem, including [few-shot learning](https://www.unite.ai/few-shot-learning/) (where the model got five example answers before responding). This helped a little, bringing the average variation down from 2.52% to 2.27%, but didn’t fix the problem.\n\nChanges in GPU hardware, batch size, numerical precision, answer order and system-prompt wording were also tested. The biggest effects were seen with answer-order and prompt wording, which came closest to the impact of changing the date.\n\n### Last Days\n\nIt’s reasonable to expect that an LLM/VLM will understand the date from the very beginning of a chat; however, there seems to be no explicit reason to build it into the system prompt (the unseen rubric that imposes [guardrails](https://www.unite.ai/what-are-ai-guardrails-how-production-systems-control-model-behavior/), and conditions the LLM’s behaviors in ways inaccessible to the user) when the LLM could routinely make a sub-Kb RAG call to ingest the latest date, as a minor housekeeping routine prior to engagement with the user.\n\nNo doubt other possibilities exist to resolve the matter; however, since the imposition of the current date into the system prompt has not hitherto been seen as a problem, there has presumably been little or no investigation in regard to this.\n\n## Online Dating\n\nFor the tests, identical prompts were used for every model, with only the date in the system prompt being changed. Every day of 2024 was tested, from January 1st through to December 31st, with all other settings kept the same.\n\nSix benchmarks were used: [MMLU](https://arxiv.org/abs/2009.03300); [GPQA](https://arxiv.org/abs/2311.12022); and [ARC-Challenge](https://www.unite.ai/from-o1-to-o3-how-openai-is-redefining-complex-reasoning-in-ai/#:~:text=ARC%20Challenge%2C%20a%20benchmark%20designed%20to%20test%20reasoning%20and%20adaptability), for multiple-choice questions (scored by answer-token probability). [GSM8K](https://arxiv.org/abs/2509.25160) for step-by-step math (final answer checked); [HumanEval](https://arxiv.org/abs/2410.12381) for Python code generation (unit-tested); and [WMT](https://doi.org/10.18653/v1/W16-2301) for English-to-German; English-to-Finnish; and English-to-Czech translation (full output evaluated). Time-dependent questions were excluded.\n\nNine models were tested: Llama 3.1 Instruct ([8B](https://huggingface.co/meta-llama/Llama-3.1-8B-Instruct) and [70B](https://build.nvidia.com/meta/llama-3_1-70b-instruct)); Gemma 3 Instruct ([4B](https://huggingface.co/google/gemma-3-4b-it) and [27B](https://huggingface.co/google/gemma-3-27b-it)); Qwen3 ([4B](https://huggingface.co/Qwen/Qwen3-4B)); Qwen3-Next ([80B](https://huggingface.co/Qwen/Qwen3-Next-80B-A3B-Instruct)); Phi-4 ([14B](https://ollama.com/library/phi4:14b)); and GPT-OSS ([20B](https://huggingface.co/openai/gpt-oss-20b) and [120B](https://huggingface.co/openai/gpt-oss-120b)).\n\nAccuracy (the percentage of questions answered correctly) was used to score the multiple-choice and math tests, while [Expected Calibration Error](https://arxiv.org/pdf/2501.19047v2) (ECE) measured how well the models’ confidence matched their actual performance.\n\nCode was checked using pass@1 (the percentage of generated code solutions that pass all tests on the first attempt); translations were scored using [BLEU](https://www.ibm.com/docs/en/watsonx/saas?topic=metrics-bleu) and [chrF](https://machinetranslate.org/chrF).\n\nThe results confirmed the researchers’ hypothesis: simply changing the date in the system prompt changed how accurately the models answered the same questions.\n\nAlong with other results shown earlier in the article, this is potentially bad news for LLM benchmarking, because a model could score better or worse depending on which day it was tested, possibly changing its position on a leaderboard, without any actual improvement or decline in its capabilities.\n\n## Conclusion\n\nThis issue highlights the divide between the deterministic computing systems we have been used to prior to around 2023, and the very different nature of diffusion-based and similar AI systems that have evolved, and continue to evolve, since then.\n\nDate resolution was essentially solved [on January 1st 1970](https://dev.to/snappy_tools/unix-timestamps-explained-why-computers-count-from-january-1-1970-241p), but has, apparently, returned to haunt the world of computing in the form of [cut-off dates](https://aiknowledgecutoff.com/), among other [temporal concerns](https://www.unite.ai/fine-tuning-ai-can-lead-to-unexpected-time-travel/).\n\n* *Titled ‘Dating the Model: Hidden Dates in System Prompts Affect LLM Evaluation’*\n\n*First published Friday, October 9, 2026*", "url": "https://wpnews.pro/news/secret-dates-in-system-prompts-undermine-language-model-evaluation", "canonical_source": "https://www.unite.ai/secret-dates-in-system-prompts-undermine-language-model-evaluation/", "published_at": "2026-10-09 00:00:00+00:00", "updated_at": "2026-10-09 14:55:26.002335+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "machine-learning", "artificial-intelligence"], "entities": ["GPT-5.1", "OpenAI", "Unite.AI"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/secret-dates-in-system-prompts-undermine-language-model-evaluation", "markdown": "https://wpnews.pro/news/secret-dates-in-system-prompts-undermine-language-model-evaluation.md", "text": "https://wpnews.pro/news/secret-dates-in-system-prompts-undermine-language-model-evaluation.txt", "jsonld": "https://wpnews.pro/news/secret-dates-in-system-prompts-undermine-language-model-evaluation.jsonld"}}