AI Grounds
Open AI Grounds on a desktop
These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.
AI Grounds
These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.
Guided discovery
Trace the judge. Explain the ranking.
Six authored prompts have two fixed authored answers, A and B. Each answer has disclosed accuracy/completeness and style marks from 0 to 2. They are teaching judgments, not universal ground truth or measured human ratings. Style evaluates wording separately from correctness: a fluent answer can be false. No actual language model or human judge runs here.
Display order changes which candidate is shown first, while its A/B identity stays fixed. The marks and whitespace word counts travel with the answer. All six pairs remain in every aggregate.
| Prompt | First candidate | Second candidate |
|---|---|---|
| P1 · What is 2 + 2? | A 4. Marks 2 / 1 · Words 1 | B 5, definitely. Marks 0 / 2 · Words 2 |
| P2 · Which country contains Berlin? | A Berlin is in Germany. Marks 2 / 2 · Words 4 | B Berlin is in Germany. Germany is the country that contains Berlin. Marks 2 / 1 · Words 11 |
| P3 · Name both listed colors: red and blue. | A Red and blue are the two listed colors. Marks 2 / 1 · Words 8 | B Red. Marks 1 / 2 · Words 1 |
| P4 · Explain overfitting in one sentence. | A Fitting training quirks can hurt unseen performance. Marks 2 / 1 · Words 7 | B Overfitting guarantees perfect performance on new data. Marks 0 / 2 · Words 7 |
| P5 · What is 3 times 3? | A 9. Marks 2 / 2 · Words 1 | B 9. Marks 2 / 2 · Words 1 |
| P6 · Briefly explain a cache. | A Reuse a stored result for the same request. Marks 2 / 2 · Words 8 | B Cache means storing a reusable result; it is a reusable stored result, used again as a result. Marks 2 / 1 · Words 17 |
Base score = accuracy weight × accuracy mark + style weight × style mark. Rubric only uses that score. First-position bonus +4 adds four only to the first displayed candidate. Longer-answer bonus +2 adds two only to the strictly longer answer; equal lengths get no bonus. A larger judged score wins; equal judged scores tie.
A wins 5 · B wins 0 · Ties 1 · All 6 pairs counted. Raw A win rate = 5/6 = 0.833333. A/B tie-adjusted scores = 0.916667 / 0.083333.
Every actual control or preset change clears stale explanation/transfer answers; reapplying the same state preserves them. Reset and prediction changes restore Accuracy first for the current experiment.
Mean pair outcome, with half credit for ties.
0.916667
(A wins + 0.5 × ties) / 6
5 wins, 1 ties. Raw win rate counts no tie credit; both use all six pairs.
Exact A/B/Tie matches with a fixed authored rubric.
1
matching verdicts / 6
6/6. Reference: Accuracy first + Rubric only. This is not independent human validation.
Sequential pairwise rating after six toy updates.
1062.278708
R′A = RA + 32 × (SA − EA)
B=937.721292; both start at 1000. K=32, scale=400. Not a current Arena Bradley–Terry fit.
Judge scores include only the selected bonus. The fixed reference always uses accuracy weight 3, style weight 1 and no bonus: A,A,A,A,Tie,A. Agreement means the same A/B/Tie label. Baseline agreement 1 is true by construction, because the judge and reference use the same rule; it does not establish independent validity or chance-corrected agreement.
| Item | Base A | Base B | Bonus A / B | Judged A / B | Verdict | Reference | Exact match |
|---|---|---|---|---|---|---|---|
| P1 | 7 | 2 | 0 / 0 | 7 / 2 | A | A | Yes |
| P2 | 8 | 7 | 0 / 0 | 8 / 7 | A | A | Yes |
| P3 | 7 | 5 | 0 / 0 | 7 / 5 | A | A | Yes |
| P4 | 7 | 2 | 0 / 0 | 7 / 2 | A | A | Yes |
| P5 | 8 | 8 | 0 / 0 | 8 / 8 | Tie | Tie | Yes |
| P6 | 8 | 7 | 0 / 0 | 8 / 7 | A | A | Yes |
Start A=B=1000. Expected A outcome EA=1/(1+10^((RB−RA)/400)). SA is 1 for an A win, 0 for a B win and 0.5 for a tie. Add Δ=32×(SA−EA) to A and subtract the same Δ from B. Ratings always sum to 2000; no intermediate rounding is used. A tie can change unequal ratings. The expectation is a rating-model quantity, not measured correctness or a guaranteed future win probability.
| Item | Before A | Before B | Expected A | Outcome A | Δ A | After A | After B |
|---|---|---|---|---|---|---|---|
| P1 | 1000 | 1000 | 0.5 | 1 | 16 | 1016 | 984 |
| P2 | 1016 | 984 | 0.545922 | 1 | 14.530498 | 1030.530498 | 969.469502 |
| P3 | 1030.530498 | 969.469502 | 0.58698 | 1 | 13.216635 | 1043.747134 | 956.252866 |
| P4 | 1043.747134 | 956.252866 | 0.623318 | 1 | 12.053809 | 1055.800943 | 944.199057 |
| P5 | 1055.800943 | 944.199057 | 0.655303 | 0.5 | -4.969697 | 1050.831246 | 949.168754 |
| P6 | 1050.831246 | 949.168754 | 0.642267 | 1 | 11.447462 | 1062.278708 | 937.721292 |
Same votes, different sequence: Forward A=1062.278708, B=937.721292; Reverse A=1065.813398, B=934.186602. Sequential updates depend on the ratings before each battle. Reordering can change the result without changing answers or vote counts.
This is a small historical online Elo illustration with a chosen K=32. It is not a real leaderboard, a current Arena Bradley–Terry maximum-likelihood fit, confidence interval, statistical significance test or universal model-quality measure. No real models are trained, generated or compared. The position and length bonuses construct interpretable bias examples; they are not calibrated estimates of real evaluator bias.
The finite model has three rubrics, three judge rules, two display orders and two rating sequences: 36 states. Marks are authored accuracy/completeness and style values 0..2, with no actual learned semantic scoring. Words are whitespace tokens. An equal judged score is one Tie. Raw A win rate counts wins only; adjusted scores give each tie one half. All six items, including ties, remain in every denominator; adjusted A+B=1. Reference is the fixed Accuracy first/Rubric only verdict list, with exact label matching, not independent human votes or a gold standard. Agreement is neither validity nor Cohen’s kappa.
Position order and rating sequence are different controls. The former changes presentation and the selected position bonus, while A/B identities persist; the latter changes only the order of rating updates. Rubric-only and length-rule votes do not depend on display order. Length bonus rewards the strictly longer answer, not actual information quality. Score gaps can outweigh a position bonus. Equal lengths receive no length bonus. No claim is made that real judge biases follow these arithmetic rules or that swapping removes them.
Online Elo starts both ratings at 1000, uses base 10, scale 400 and constant K=32. Each SA is 1/0/0.5; expectation is recomputed from the previous ratings. No intermediate rounding occurs. Opposite updates conserve the rating sum, ties can change unequal ratings, and reordering fixed votes can change the final ratings. This illustrates a classic sequential procedure rather than current Arena’s Bradley–Terry fitting or an objective ranking. Six authored examples support no uncertainty inference or general real-model performance claim.
Judging LLM-as-a-Judge · Rubrics, agreement and bias