Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints
🤯 Can LLMs Measure Themselves? Why ‘Frozen’ AI Judges Are Broken
The AI ecosystem is booming, and measurement—or ‘judging’—is at the core of it. From benchmarking training datasets to scoring generations and populating leaderboards, large language models (LLMs) are increasingly used as automatic judges for other LLMs. This seemingly robust system rests on a shaky premise: that sending the same prompt to a model today will yield the same performance score tomorrow.
Our latest work audits this critical assumption across massive-scale, shared service endpoints. The findings are sobering: The measurement itself is unstable.
In a detailed, preregistered investigation Clean Engineering, Unstable Measurement, we audited over 52,000 identical request attempts across multiple days and providers. Our core metrics showed significant divergence: same-window repeat rankings only agreed at a Spearman correlation of $0.400$, drastically short of the required $0.90$; even byte-identical next-day replays scored $0.78$ (compared to the expected $0.99$).
The ‘ceiling’ on our execution record—the consistency achieved when we controlled every variable—demonstrated a massive gap between idealized performance and real-world, shared infrastructure measurements. This instability isn’t theoretical; it affects published benchmarks and model rankings today.
🛠️ The Three Breakdown Mechanisms
We identify three major culprits for this ‘observational drift’:
- Label Drift: Simply mapping a label to its meaning can bias the readout as strongly as the underlying signal itself. The human element in judging isn’t pure data; it’s contextual.
- Quantization Noise: Candidate differences are incredibly small, falling seven orders of magnitude below the noise floor of the measurement instrument itself. This means the judge is essentially measuring its own static interference.
- Permutation Instability: Even when inputs are byte-identical, subsequent rankings can diverge due to systemic noise inherent in shared endpoints and exact-permutation readouts.
🌐 Beyond Model Names: The Systemic Problem
Our investigation shows that simply waiting or switching API providers does not resolve the fundamental measurement instability. Across four major providers sampled, the median reliability metrics consistently remained low (medians between $0.74$ and $0.88$), suggesting a systemic failure across commercial shared endpoints.
Furthermore, we show that remediation efforts are insufficient: self-hosting on stable kernels only provided stability when the server was idle; for active services, the problem persists. The inconsistency tracks with known gaps—the measurement readout is more sensitive to type of error than its size.
What does this mean for AI development?
The reliability of external measurements on shared endpoints cannot be taken for granted. A model name deployed today is not a frozen instrument. Before any gate can be placed on model performance, the measurement apparatus itself must be thoroughly audited. We propose a three-level snapshot-identity ladder and eight concrete design rules to safeguard against this ‘Measurement Drift.’
The takeaway: If you rely on external benchmarks or shared APIs to validate AI progress, you must first audit the reliability of your measuring tool.