Skip to main content
Usami Laboratory
Research explainers

Examining the Instrument Behind LLM Evaluation

When AI evaluates AI, what does the judge measure? We propose a Judge Datasheet that treats an LLM as a measurement instrument and separates judgment bias from sensitivity to differences in answer quality.

Joint research with Mitsubishi Heavy Industries, Research & Innovation Center

arXiv · Preprint · 2026

The problem

LLM-as-a-Judge is used to compare models and assess generated answers. Even when answers have equal quality, presentation order or surface form can affect the verdict. This study starts by examining the evaluator itself.

The research idea

Drawing on dark current in a sensor with no incident light, we examine judgments when no quality difference is present. Empty inputs, identical answers, quality-preserving surface changes and reversed answer order are combined with a controlled quality ladder. The Judge Datasheet organizes these measurements into an instrument profile.

Dark current
Judgments without a quality difference
Cross-sensitivity
Responses to quality-preserving surface changes
Position bias
Effects of answer presentation order
Quality sensitivity
Discrimination on a controlled quality ladder

What was evaluated

Three public models are evaluated for judgments under null conditions, responses to surface form and presentation order, and sensitivity to controlled quality differences. Experiments also examine how prompts favoring ties change both bias measures and sensitivity. The design examines distinct judge properties rather than collapsing them into one accuracy score.

Considering applications

For answer ranking and quality assessment with LLMs, the study provides a way to examine bias, quality sensitivity and tie criteria separately. It motivates treating the model, prompt and verdict readout as one evaluation instrument, with explicit measurement conditions.

Original research

LLM Judges Have Dark Current: A Psychometric Datasheet for LLM-as-a-Judge Evaluation

Usami, Hiroyasu, Hara, Keisuke, Tsuboi, Ayato, Matsuda, Naohiko

arXiv, 2026 · arXiv:2606.15610 [cs.CL], 22 pages, 4 figures

References in subsequent research

Subsequent studies discuss what an LLM evaluator measures and how it should be characterized.

Stanford University

A Calibrated Instrument for Measuring How Inference Optimizations Affect Output Quality

Jerry Kaplan

A study measuring how inference optimizations affect output quality. It cites Judge Datasheet as prior work characterizing an LLM evaluator as a measurement instrument.

Read the citing paper

University of Oxford

Three Ways Classical Test Theory Misleads for LLM Judges

Louis Yiven Zhu

A study examining measurement design when classical test theory is applied to LLM judges. It cites the Judge Datasheet framework among measurement-theoretic approaches to judge evaluation.

Read the citing paper

Stony Brook University

The First Token Is Not the Verdict: Hidden Costs of Reading LLM Judges Without Generating

Gnaneswar Villuri, Hashmath Shaik, Alex Doboli

In a paper submitted to JUDGe — Can We Trust the Judge? @ NeurIPS 2026, this work is cited as an example of a judge audit that generates text and parses the verdict. The paper examines how that procedure differs from first-token readout.

Read the citing paper

Featured coverage

· Jake Handy · AI Weekly Update

AI Research of the Week

The June 22, 2026 edition featured the paper and its authors in AI Research of the Week.

Read the article