AI Grounds
Open AI Grounds on a desktop
These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.
AI Grounds
These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.
Guided discovery
Reorder results; separate metric questions.
Six fixed IDs for each of two toy queries have authored relevance grades. Grade 0 means not relevant; any positive grade counts as one relevant item for binary metrics. Larger grades carry more gain for nDCG. Reorder results without changing these judgments. No learned score or real semantic model is evaluated.
The dashed boundary is after rank 3. Bars count gain only in the returned prefix; rows outside it have zero counted gain, not necessarily grade 0. Every exact ID, relevance grade and inclusion remains in the table. Fixed gain scale 0..7 permits direct comparisons.
Remove A from Q1 and insert at this rank; other IDs keep their relative order. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.
Shared cutoff for precision, recall and nDCG; full-list RR/MRR do not use this cutoff. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.
Binary relevance · Inspect Q1 · Item A at rank 2 · Top k 3 · Q1 B,A,D,C,F,E · Q2 A,B,C,D,E,F
Query inspects a different list; Item to move selects an ID. These selections preserve both orders and answers. Only Move to rank changes an order. Scenario changes restore named starting orders/k; prediction or Reset restores the current experiment.
| Rank | Item | Grade | Relevant? | Gain | Discount | Counted gain | Returned |
|---|---|---|---|---|---|---|---|
| 1 | B | 0 | No | 0.000000 | 1.000000 | 0.000000 | Included |
| 2 | A | 1 | Yes | 1.000000 | 0.630930 | 0.630930 | Included |
| 3 | D | 0 | No | 0.000000 | 0.500000 | 0.000000 | Included |
| 4 | C | 1 | Yes | 1.000000 | 0.430677 | 0.000000 | Outside |
| 5 | F | 0 | No | 0.000000 | 0.386853 | 0.000000 | Outside |
| 6 | E | 1 | Yes | 1.000000 | 0.356207 | 0.000000 | Outside |
Q1: found 1 / returned 3; found 1 / total relevant 3. DCG@3 = 0.630930; ideal DCG@3 = 2.130930 from all six grades sorted descending (1,1,1,0,0,0) before the same cutoff.
Relevant fraction of the returned prefix.
0.333333
1 found / 3 returned
Uses binary relevance; order inside the same set does not matter.
Fraction of all relevant IDs recovered.
0.333333
1 found / 3 total relevant
Complete six-item judged universe; zero denominator is undefined.
Mean of both full-list reciprocal ranks.
0.500000
(0.500000 + 0.500000) / 2
First hit only; includes no-hit RR=0. This is not MRR@k.
Discounted graded gain versus the ideal prefix.
0.296082
0.630930 / 2.130930
Same k/gain/discount in numerator and ideal denominator.
Q1 first relevant rank 2: RR = 0.500000. Q2 first relevant rank 2: RR = 0.500000. MRR = (RRQ1+RRQ2)/2 = 0.500000.
RR uses the first relevant item in the full six-item list; changing only k cannot change this untruncated RR or MRR. One query’s RR is not the two-query mean. Later relevant items do not affect first-hit RR. Other tasks may use MRR@k or different empty-query policies; state the convention.
Both queries have six unique IDs with fixed toy relevance grades. Query-specific judgments need not agree for an ID. Positive grades are binary relevant; nDCG uses gains 2ᵍʳᵃᵈᵉ−1 (grades 0,1,2,3 give gains 0,1,3,7), discounted by 1/log₂(rank+1). A grade is not a probability. There are no learned ranking scores or score ties: the manual order is strict. Moving one ID preserves all other relative order and every identity.
Precision@k=found/k. Recall@k=found/total relevant when the relevant count is positive, otherwise Undefined. DCG sums the top-k discounted gains; ideal DCG sorts every judged item by descending grade and then applies the same k and discounts. nDCG=DCG/ideal DCG when the ideal is positive, otherwise Undefined. Equal-grade ideal permutations give the same denominator. Different gain conventions give different values; for example, scikit-learn uses supplied true scores directly, rather than automatically applying this demo’s exponential gain.
MRR is the arithmetic mean of two full-list RR values. Each RR is 1/first relevant rank, or explicitly 0 if no relevant item exists. Empty queries remain in the mean. Skipping or scoring them differently would be another policy; undefined recall/nDCG does not imply perfect retrieval. Selected-query precision/recall/nDCG are not averages across queries. More results or a higher one-metric score does not establish overall task success.
Actual rank/k/scenario edits clear stale explanations; Query and Item to move only select evidence. Reset restores the current baseline, and free-exploration Reset starts Experiment 1. Native selects, number editors and sliders make every order reachable without dragging. This small evaluation demo excludes training, embeddings, ANN/indexing, unknown judgments, label editing, significance tests and production-quality claims. Compare only with the same query set, judgments, cutoff, gain/discount and aggregation conventions.
Stanford IR book · Ranked retrieval and exponential graded gain