The problem
LLM-as-a-Judge is used to compare models and assess generated answers. Even when answers have equal quality, presentation order or surface form can affect the verdict. This study starts by examining the evaluator itself.
The research idea
Drawing on dark current in a sensor with no incident light, we examine judgments when no quality difference is present. Empty inputs, identical answers, quality-preserving surface changes and reversed answer order are combined with a controlled quality ladder. The Judge Datasheet organizes these measurements into an instrument profile.
- Dark current
- Judgments without a quality difference
- Cross-sensitivity
- Responses to quality-preserving surface changes
- Position bias
- Effects of answer presentation order
- Quality sensitivity
- Discrimination on a controlled quality ladder
What was evaluated
Three public models are evaluated for judgments under null conditions, responses to surface form and presentation order, and sensitivity to controlled quality differences. Experiments also examine how prompts favoring ties change both bias measures and sensitivity. The design examines distinct judge properties rather than collapsing them into one accuracy score.
Considering applications
For answer ranking and quality assessment with LLMs, the study provides a way to examine bias, quality sensitivity and tie criteria separately. It motivates treating the model, prompt and verdict readout as one evaluation instrument, with explicit measurement conditions.
Original research
LLM Judges Have Dark Current: A Psychometric Datasheet for LLM-as-a-Judge Evaluation
Usami, Hiroyasu, Hara, Keisuke, Tsuboi, Ayato, Matsuda, Naohiko
arXiv, 2026 · arXiv:2606.15610 [cs.CL], 22 pages, 4 figures
References in subsequent research
Subsequent studies discuss what an LLM evaluator measures and how it should be characterized.
Stanford University
A Calibrated Instrument for Measuring How Inference Optimizations Affect Output Quality
Jerry Kaplan
A study measuring how inference optimizations affect output quality. It cites Judge Datasheet as prior work characterizing an LLM evaluator as a measurement instrument.
Read the citing paperUniversity of Oxford
Three Ways Classical Test Theory Misleads for LLM Judges
Louis Yiven Zhu
A study examining measurement design when classical test theory is applied to LLM judges. It cites the Judge Datasheet framework among measurement-theoretic approaches to judge evaluation.
Read the citing paperStony Brook University
The First Token Is Not the Verdict: Hidden Costs of Reading LLM Judges Without Generating
Gnaneswar Villuri, Hashmath Shaik, Alex Doboli
In a paper submitted to JUDGe — Can We Trust the Judge? @ NeurIPS 2026, this work is cited as an example of a judge audit that generates text and parses the verdict. The paper examines how that procedure differs from first-token readout.
Read the citing paperFeatured coverage
· Jake Handy · AI Weekly Update
AI Research of the Week
The June 22, 2026 edition featured the paper and its authors in AI Research of the Week.
Read the article