AI Grounds

Open AI Grounds on a desktop

These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.

Guided discovery

Preference & Judge Metrics Lab

Trace the judge. Explain the ranking.

Your dataset

Six authored prompts have two fixed authored answers, A and B. Each answer has disclosed accuracy/completeness and style marks from 0 to 2. They are teaching judgments, not universal ground truth or measured human ratings. Style evaluates wording separately from correctness: a fluent answer can be false. No actual language model or human judge runs here.

Display order changes which candidate is shown first, while its A/B identity stays fixed. The marks and whitespace word counts travel with the answer. All six pairs remain in every aggregate.

A is displayed first, B second. Marks are accuracy / style. Words are whitespace-separated tokens including punctuation.
PromptFirst candidateSecond candidate
P1 · What is 2 + 2?A

4.

Marks 2 / 1 · Words 1
B

5, definitely.

Marks 0 / 2 · Words 2
P2 · Which country contains Berlin?A

Berlin is in Germany.

Marks 2 / 2 · Words 4
B

Berlin is in Germany. Germany is the country that contains Berlin.

Marks 2 / 1 · Words 11
P3 · Name both listed colors: red and blue.A

Red and blue are the two listed colors.

Marks 2 / 1 · Words 8
B

Red.

Marks 1 / 2 · Words 1
P4 · Explain overfitting in one sentence.A

Fitting training quirks can hurt unseen performance.

Marks 2 / 1 · Words 7
B

Overfitting guarantees perfect performance on new data.

Marks 0 / 2 · Words 7
P5 · What is 3 times 3?A

9.

Marks 2 / 2 · Words 1
B

9.

Marks 2 / 2 · Words 1
P6 · Briefly explain a cache.A

Reuse a stored result for the same request.

Marks 2 / 2 · Words 8
B

Cache means storing a reusable result; it is a reusable stored result, used again as a result.

Marks 2 / 1 · Words 17

Base score = accuracy weight × accuracy mark + style weight × style mark. Rubric only uses that score. First-position bonus +4 adds four only to the first displayed candidate. Longer-answer bonus +2 adds two only to the strictly longer answer; equal lengths get no bonus. A larger judged score wins; equal judged scores tie.

A wins 5 · B wins 0 · Ties 1 · All 6 pairs counted. Raw A win rate = 5/6 = 0.833333. A/B tie-adjusted scores = 0.916667 / 0.083333.

Every actual control or preset change clears stale explanation/transfer answers; reapplying the same state preserves them. Reset and prediction changes restore Accuracy first for the current experiment.

A tie-adjusted score

Mean pair outcome, with half credit for ties.

0.916667

(A wins + 0.5 × ties) / 6

5 wins, 1 ties. Raw win rate counts no tie credit; both use all six pairs.

Reference agreement

Exact A/B/Tie matches with a fixed authored rubric.

1

matching verdicts / 6

6/6. Reference: Accuracy first + Rubric only. This is not independent human validation.

A online Elo

Sequential pairwise rating after six toy updates.

1062.278708

R′A = RA + 32 × (SA − EA)

B=937.721292; both start at 1000. K=32, scale=400. Not a current Arena Bradley–Terry fit.

Trace the judge’s decision

Judge scores include only the selected bonus. The fixed reference always uses accuracy weight 3, style weight 1 and no bonus: A,A,A,A,Tie,A. Agreement means the same A/B/Tie label. Baseline agreement 1 is true by construction, because the judge and reference use the same rule; it does not establish independent validity or chance-corrected agreement.

Scores keep candidate identity fixed even when Display order swaps.
ItemBase ABase BBonus A / BJudged A / BVerdictReferenceExact match
P1720 / 07 / 2AAYes
P2870 / 08 / 7AAYes
P3750 / 07 / 5AAYes
P4720 / 07 / 2AAYes
P5880 / 08 / 8TieTieYes
P6870 / 08 / 7AAYes

Update a pairwise rating

Start A=B=1000. Expected A outcome EA=1/(1+10^((RB−RA)/400)). SA is 1 for an A win, 0 for a B win and 0.5 for a tie. Add Δ=32×(SA−EA) to A and subtract the same Δ from B. Ratings always sum to 2000; no intermediate rounding is used. A tie can change unequal ratings. The expectation is a rating-model quantity, not measured correctness or a guaranteed future win probability.

P1 → P6: later expected scores use the updated ratings.
ItemBefore ABefore BExpected AOutcome AΔ AAfter AAfter B
P1100010000.51161016984
P210169840.545922114.5304981030.530498969.469502
P31030.530498969.4695020.58698113.2166351043.747134956.252866
P41043.747134956.2528660.623318112.0538091055.800943944.199057
P51055.800943944.1990570.6553030.5-4.9696971050.831246949.168754
P61050.831246949.1687540.642267111.4474621062.278708937.721292

Same votes, different sequence: Forward A=1062.278708, B=937.721292; Reverse A=1065.813398, B=934.186602. Sequential updates depend on the ratings before each battle. Reordering can change the result without changing answers or vote counts.

This is a small historical online Elo illustration with a chosen K=32. It is not a real leaderboard, a current Arena Bradley–Terry maximum-likelihood fit, confidence interval, statistical significance test or universal model-quality measure. No real models are trained, generated or compared. The position and length bonuses construct interpretable bias examples; they are not calibrated estimates of real evaluator bias.

Construction and limits

The finite model has three rubrics, three judge rules, two display orders and two rating sequences: 36 states. Marks are authored accuracy/completeness and style values 0..2, with no actual learned semantic scoring. Words are whitespace tokens. An equal judged score is one Tie. Raw A win rate counts wins only; adjusted scores give each tie one half. All six items, including ties, remain in every denominator; adjusted A+B=1. Reference is the fixed Accuracy first/Rubric only verdict list, with exact label matching, not independent human votes or a gold standard. Agreement is neither validity nor Cohen’s kappa.

Position order and rating sequence are different controls. The former changes presentation and the selected position bonus, while A/B identities persist; the latter changes only the order of rating updates. Rubric-only and length-rule votes do not depend on display order. Length bonus rewards the strictly longer answer, not actual information quality. Score gaps can outweigh a position bonus. Equal lengths receive no length bonus. No claim is made that real judge biases follow these arithmetic rules or that swapping removes them.

Online Elo starts both ratings at 1000, uses base 10, scale 400 and constant K=32. Each SA is 1/0/0.5; expectation is recomputed from the previous ratings. No intermediate rounding occurs. Opposite updates conserve the rating sum, ties can change unequal ratings, and reordering fixed votes can change the final ratings. This illustrates a classic sequential procedure rather than current Arena’s Bradley–Terry fitting or an objective ranking. Six authored examples support no uncertainty inference or general real-model performance claim.

Judging LLM-as-a-Judge · Rubrics, agreement and bias
Judging the Judges · Position consistency
LMSYS · Online Elo to Bradley–Terry
FastChat · Pairwise rating formulas