AI Grounds

Open AI Grounds on a desktop

These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.

Guided discovery

Retrieval Ranking Metrics Lab

Reorder results; separate metric questions.

Your dataset

Six fixed IDs for each of two toy queries have authored relevance grades. Grade 0 means not relevant; any positive grade counts as one relevant item for binary metrics. Larger grades carry more gain for nDCG. Reorder results without changing these judgments. No learned score or real semantic model is evaluated.

Discounted gain by rank (Q1)

03.571. B0.0000002. A0.6309303. D0.0000004. C0.0000005. F0.0000006. E0.000000Counted gain (fixed 0..7)

The dashed boundary is after rank 3. Bars count gain only in the returned prefix; rows outside it have zero counted gain, not necessarily grade 0. Every exact ID, relevance grade and inclusion remains in the table. Fixed gain scale 0..7 permits direct comparisons.

Remove A from Q1 and insert at this rank; other IDs keep their relative order. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.

Shared cutoff for precision, recall and nDCG; full-list RR/MRR do not use this cutoff. Use arrow keys on the slider. Press Enter or leave the number field to apply an exact edit.

Binary relevance · Inspect Q1 · Item A at rank 2 · Top k 3 · Q1 B,A,D,C,F,E · Q2 A,B,C,D,E,F

Query inspects a different list; Item to move selects an ID. These selections preserve both orders and answers. Only Move to rank changes an order. Scenario changes restore named starting orders/k; prediction or Reset restores the current experiment.

Ranked results for Q1

Every ID appears exactly once. Relevant means grade>0. Gain = 2ᵍʳᵃᵈᵉ−1; discount = 1/log₂(rank+1). Counted gain is zero outside Top k. Values are computed at full precision, displayed rounded to six decimals.
RankItemGradeRelevant?GainDiscountCounted gainReturned
1B0No0.0000001.0000000.000000Included
2A1Yes1.0000000.6309300.630930Included
3D0No0.0000000.5000000.000000Included
4C1Yes1.0000000.4306770.000000Outside
5F0No0.0000000.3868530.000000Outside
6E1Yes1.0000000.3562070.000000Outside

Q1: found 1 / returned 3; found 1 / total relevant 3. DCG@3 = 0.630930; ideal DCG@3 = 2.130930 from all six grades sorted descending (1,1,1,0,0,0) before the same cutoff.

Precision@3 (Q1)

Relevant fraction of the returned prefix.

0.333333

1 found / 3 returned

Uses binary relevance; order inside the same set does not matter.

Recall@3 (Q1)

Fraction of all relevant IDs recovered.

0.333333

1 found / 3 total relevant

Complete six-item judged universe; zero denominator is undefined.

MRR (two queries)

Mean of both full-list reciprocal ranks.

0.500000

(0.500000 + 0.500000) / 2

First hit only; includes no-hit RR=0. This is not MRR@k.

nDCG@3 (Q1)

Discounted graded gain versus the ideal prefix.

0.296082

0.630930 / 2.130930

Same k/gain/discount in numerator and ideal denominator.

Q1 first relevant rank 2: RR = 0.500000. Q2 first relevant rank 2: RR = 0.500000. MRR = (RRQ1+RRQ2)/2 = 0.500000.

RR uses the first relevant item in the full six-item list; changing only k cannot change this untruncated RR or MRR. One query’s RR is not the two-query mean. Later relevant items do not affect first-hit RR. Other tasks may use MRR@k or different empty-query policies; state the convention.

Construction and limits

Both queries have six unique IDs with fixed toy relevance grades. Query-specific judgments need not agree for an ID. Positive grades are binary relevant; nDCG uses gains 2ᵍʳᵃᵈᵉ−1 (grades 0,1,2,3 give gains 0,1,3,7), discounted by 1/log₂(rank+1). A grade is not a probability. There are no learned ranking scores or score ties: the manual order is strict. Moving one ID preserves all other relative order and every identity.

Precision@k=found/k. Recall@k=found/total relevant when the relevant count is positive, otherwise Undefined. DCG sums the top-k discounted gains; ideal DCG sorts every judged item by descending grade and then applies the same k and discounts. nDCG=DCG/ideal DCG when the ideal is positive, otherwise Undefined. Equal-grade ideal permutations give the same denominator. Different gain conventions give different values; for example, scikit-learn uses supplied true scores directly, rather than automatically applying this demo’s exponential gain.

MRR is the arithmetic mean of two full-list RR values. Each RR is 1/first relevant rank, or explicitly 0 if no relevant item exists. Empty queries remain in the mean. Skipping or scoring them differently would be another policy; undefined recall/nDCG does not imply perfect retrieval. Selected-query precision/recall/nDCG are not averages across queries. More results or a higher one-metric score does not establish overall task success.

Actual rank/k/scenario edits clear stale explanations; Query and Item to move only select evidence. Reset restores the current baseline, and free-exploration Reset starts Experiment 1. Native selects, number editors and sliders make every order reachable without dragging. This small evaluation demo excludes training, embeddings, ANN/indexing, unknown judgments, label editing, significance tests and production-quality claims. Compare only with the same query set, judgments, cutoff, gain/discount and aggregation conventions.

Stanford IR book · Ranked retrieval and exponential graded gain
NIST TREC · Reciprocal rank and mean across queries
scikit-learn · nDCG and gain conventions