Your LLM Benchmark Score Changes Every Day Without You Touching Anything
Your LLM Benchmark Score Changes Every Day Without You Touching Anything
The leaderboard you trusted to pick your model stack? It may have been wrong on Tuesday and right on Thursday—for no reason you can see or control. A new paper reveals that something as mundane as the calendar date, silently injected into the system prompt by the model provider, is quietly shifting performance by double digits. Your eval is lying to you, and it’s been doing it every single day.
What happened

Researchers tested 9 recent LLMs across 6 datasets covering multiple-choice QA, math reasoning, code generation, and machine translation—and found that model performance fluctuates purely as a function of the current date, because many providers silently inject the current date into system prompts in ways users cannot see or override. The swings are not rounding errors: up to 6% on multiple-choice QA, 14% on math reasoning, 7% on code generation, and 2.84 BLEU on machine translation. These date-driven deltas are larger than other known sources of non-determinism like batch size and numerical precision. Critically, chain-of-thought prompting—the standard “make it more reliable” move—doesn’t fix the problem and actually amplifies the sensitivity. Model rankings on leaderboards shift as a result, meaning the winner on the day someone ran their benchmark contamination check may not have been the winner yesterday.
Cold read
This is a real finding, but the leap from “evaluation is noisy” to “your production system is broken” is not supported by the paper. The study measures benchmark performance variance, not real-world task degradation—a 14% swing on a math reasoning benchmark does not mean your customer-facing math feature fluctuates by 14% day to day. The mechanism is also underspecified: the paper identifies the date injection as the vector, but doesn’t fully explain why the date drives these swings, which means you can’t predict which tasks or models will be most affected on your specific use case. The paper covers 9 models, which is a reasonable sample, but we don’t know how provider-specific the date injection behavior is—some APIs may not inject dates at all, or may let you override them. And if you’re already using fixed-date system prompts or a locked deployment configuration, much of this may simply not apply to you.
What it means for you
- Signal maturity: 3/5 — Empirically solid on evals, but translation to production impact is unproven
- Who gets hurt: ML engineers running recurring model comparison benchmarks to justify infrastructure spend or model-switching decisions; anyone using LLM-as-judge pipelines that run on a schedule without pinning date context
- What breaks if this is true: Any golden dataset evaluation you ran in Q1 to select your model vendor may have produced a result that wouldn’t replicate today—or even tomorrow—making model procurement decisions systematically unreliable
- Why it might not land: Most operators don’t run evaluations on bare API calls; wrappers, fixed system prompts, and orchestration layers may already neutralize or stabilize the date signal without anyone knowing it
- Watch for: Major model providers (OpenAI, Anthropic, Google) updating their API documentation or model cards to explicitly disclose whether and how date context is injected into system prompts—that would confirm the industry took this seriously
Forecast as of 2026-10-01
By Q2 2027, at least two of the top-five LLM API providers will add explicit documentation or an API parameter allowing users to suppress or pin the injected date in system prompts—driven by enterprise reproducibility complaints, not this paper alone.
Source: Dating the Model: Hidden Dates in System Prompts Affect LLM Evaluation — Mario Sanz-Guerrero, Minh Duc Bui, Manuel Mager, Katharina von der Wense. https://arxiv.org/abs/2609.36931v1
