AI Grounds

Open AI Grounds on a desktop

These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.

Guided discovery

Reference Answer Metrics Lab

Compare answers. See what each metric counts.

Your dataset

One authored candidate answer is compared with one selected reference. No language model generates text here. Exact match, Token F1, BLEU and ROUGE use the selected lexical policy. Toy cosine always uses its own fixed normalized, hand-authored word vectors; it is not BERTScore or a correctness judge.

Reference · Original quote

the cat sat on the mat

Candidate · Original quote

The CAT sat on the mat!

Score scale 0..1 for these authored examples. Defined zero scores have dots. Undefined cosine has no bar or invented zero dot. The table provides exact data and formulas. These values are not correctness probabilities.

Raw preserves whole strings for Exact match and splits whitespace for token metrics. BLEU-2 is a fixed-order unsmoothed single-pair score. Toy cosine is 1; representation-dependent similarity.

Every actual selector or preset change clears stale explanation/transfer answers; reapplying the same state preserves them. Changing Text normalization does not rewrite the original quotes or change the toy vector preprocessing.

Matched tokens

Shared token occurrences, clipped to available counts.

3

sum(min(candidate count, reference count))

6 candidate / 6 reference tokens; order is ignored.

Matched bigrams

Shared adjacent token pairs, with clipped counts.

2

p2 = clipped matches / candidate pairs

5 candidate bigrams. Shown for comparison even when BLEU-1 ignores p2.

Ordered matches

Length of a longest common subsequence.

3

ROUGE-L F1 = 2×LCS / (candidate + reference length)

Matches keep sequence order but may have gaps; adjacency is not required.

Trace what each score counts

Processed Reference: the cat sat on the mat
Processed Candidate: The CAT sat on the mat!

One longest common subsequence: sat → on → the. Other longest choices can have the same length.

One pair · Raw lexical policy · fixed BLEU-2 · Toy cosine has separate fixed vector preprocessing. Display rounds to six decimals.
MetricConstructionValue
Exact matchProcessed strings differ; the policy does not edit the original quotes.0
Token F1Clipped overlap 3; precision=0.5, recall=0.5. F1=2×3/(6+6).0.5
BLEU-2p1=3/6, p2=2/5; missing denominators use precision 0. BP=1. BP×sqrt(p1×p2); no smoothing.0.447214
ROUGE-L F1Longest common subsequence length 3; 2×3/(6+6).0.5
Toy cosineDot=0.5; norms=0.707107,0.707107. Cosine = dot / product of norms.1

Clipped token and bigram counts

Repeated occurrences cannot claim more matches than the reference contains. All rows are shown, including unmatched reference items. A bigram is an adjacent pair.

UnitTextCandidate countReference countClipped matches
TokenThe100
TokenCAT100
Tokensat111
Tokenon111
Tokenthe121
Tokenmat!100
Tokencat010
Tokenmat010
BigramThe CAT100
BigramCAT sat100
Bigramsat on111
Bigramon the111
Bigramthe mat!100
Bigramthe cat010
Bigramcat sat010
Bigramthe mat010

Toy vector construction · Separate fixed normalization

The hand-authored dictionary maps cat/feline to (1,0), sat/sit/rested to (0,1), and mat/rug to (1,1); every other word, including not and did, to (0,0). The vector is the coordinatewise mean over all normalized word tokens, including zero-vector words. Order is discarded. This dictionary deliberately misses negation and is not a learned semantic model.

Candidate mean = (0.5, 0.5); reference mean = (0.5, 0.5). Dot = 0.5; cosine = dot / (candidate norm × reference norm) = 1.

General cosine ranges −1..1; this nonnegative toy dictionary produces 0..1 when defined. A zero mean vector makes cosine Undefined. Learned metrics such as BERTScore use actual contextual embeddings and different token matching; these toy values cannot stand in for their results. High overlap or similarity does not establish equivalent meaning or factual correctness.

Construction and limits

The finite model has three references and eight candidates, two lexical policies and fixed BLEU orders 1/2. The strings are authored English examples. Raw Exact match compares entire strings; Raw token metrics split on whitespace and preserve punctuation/case. Normalized lowercases, removes ASCII punctuation, removes English articles a/an/the, and collapses whitespace. These are explicit conventions, not a universal tokenizer or multilingual evaluation policy. One reference is selected; no multi-reference maximum or benchmark aggregation is used.

Token F1 uses multiset overlap, precision=matched/candidate length and recall=matched/reference length. If overlap is zero, F1 is 0, including two empty token lists. Two empty strings have Exact match 1. ROUGE-L here is sentence-level LCS F1 with beta=1, without stemming or summary-level union: 2×LCS/(candidate+reference length), with 0 for no ordered matches or empty input. These conventions must accompany comparisons.

BLEU uses clipped candidate n-gram precision and brevity penalty BP=exp(1−reference length/candidate length) for a nonempty shorter candidate, else BP=1 for a nonempty candidate at least as long as its reference. Empty candidate uses BP=0 and BLEU=0. BLEU-1=BP×p1; BLEU-2=BP×sqrt(p1×p2), with equal weights and no smoothing. If an order has no candidate n-grams, its precision is 0: an exact one-token pair still has BLEU-2=0 under this fixed-order convention. Other implementations may smooth or adjust effective order. This single-pair demonstration is not corpus BLEU-4; averaging sentence BLEU does not generally equal corpus BLEU. Punctuation, normalization, order, reference choice and aggregation all affect comparability.

Toy cosine always preprocesses original strings with Normalized, independently of Text normalization and BLEU order. It averages the disclosed dictionary vectors, including zero vectors, and divides their dot product by both vector norms; a zero norm is Undefined. It has no learned model, BERTScore computation, generation, calibrated correctness probability, factuality verifier, semantic guarantee or preferred universal score. Reference choice can penalize valid alternate wording. Presets restore their complete state; Reset and prediction changes restore Surface for the current experiment, while free Reset starts Experiment 1.

BLEU · Original clipped precision and brevity penalty
ROUGE · Longest common subsequence measures
SQuAD v1.1 evaluation script · Normalization and token F1 conventions
BERTScore · Actual contextual embedding metric, not this toy