AI Grounds
Open AI Grounds on a desktop
These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.
AI Grounds
These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.
Guided discovery
Compare answers. See what each metric counts.
One authored candidate answer is compared with one selected reference. No language model generates text here. Exact match, Token F1, BLEU and ROUGE use the selected lexical policy. Toy cosine always uses its own fixed normalized, hand-authored word vectors; it is not BERTScore or a correctness judge.
the cat sat on the mat
The CAT sat on the mat!
Score scale 0..1 for these authored examples. Defined zero scores have dots. Undefined cosine has no bar or invented zero dot. The table provides exact data and formulas. These values are not correctness probabilities.
Raw preserves whole strings for Exact match and splits whitespace for token metrics. BLEU-2 is a fixed-order unsmoothed single-pair score. Toy cosine is 1; representation-dependent similarity.
Every actual selector or preset change clears stale explanation/transfer answers; reapplying the same state preserves them. Changing Text normalization does not rewrite the original quotes or change the toy vector preprocessing.
Shared token occurrences, clipped to available counts.
3
sum(min(candidate count, reference count))
6 candidate / 6 reference tokens; order is ignored.
Shared adjacent token pairs, with clipped counts.
2
p2 = clipped matches / candidate pairs
5 candidate bigrams. Shown for comparison even when BLEU-1 ignores p2.
Length of a longest common subsequence.
3
ROUGE-L F1 = 2×LCS / (candidate + reference length)
Matches keep sequence order but may have gaps; adjacency is not required.
Processed Reference: the cat sat on the mat
Processed Candidate: The CAT sat on the mat!
One longest common subsequence: sat → on → the. Other longest choices can have the same length.
| Metric | Construction | Value |
|---|---|---|
| Exact match | Processed strings differ; the policy does not edit the original quotes. | 0 |
| Token F1 | Clipped overlap 3; precision=0.5, recall=0.5. F1=2×3/(6+6). | 0.5 |
| BLEU-2 | p1=3/6, p2=2/5; missing denominators use precision 0. BP=1. BP×sqrt(p1×p2); no smoothing. | 0.447214 |
| ROUGE-L F1 | Longest common subsequence length 3; 2×3/(6+6). | 0.5 |
| Toy cosine | Dot=0.5; norms=0.707107,0.707107. Cosine = dot / product of norms. | 1 |
Repeated occurrences cannot claim more matches than the reference contains. All rows are shown, including unmatched reference items. A bigram is an adjacent pair.
| Unit | Text | Candidate count | Reference count | Clipped matches |
|---|---|---|---|---|
| Token | The | 1 | 0 | 0 |
| Token | CAT | 1 | 0 | 0 |
| Token | sat | 1 | 1 | 1 |
| Token | on | 1 | 1 | 1 |
| Token | the | 1 | 2 | 1 |
| Token | mat! | 1 | 0 | 0 |
| Token | cat | 0 | 1 | 0 |
| Token | mat | 0 | 1 | 0 |
| Bigram | The CAT | 1 | 0 | 0 |
| Bigram | CAT sat | 1 | 0 | 0 |
| Bigram | sat on | 1 | 1 | 1 |
| Bigram | on the | 1 | 1 | 1 |
| Bigram | the mat! | 1 | 0 | 0 |
| Bigram | the cat | 0 | 1 | 0 |
| Bigram | cat sat | 0 | 1 | 0 |
| Bigram | the mat | 0 | 1 | 0 |
The hand-authored dictionary maps cat/feline to (1,0), sat/sit/rested to (0,1), and mat/rug to (1,1); every other word, including not and did, to (0,0). The vector is the coordinatewise mean over all normalized word tokens, including zero-vector words. Order is discarded. This dictionary deliberately misses negation and is not a learned semantic model.
Candidate mean = (0.5, 0.5); reference mean = (0.5, 0.5). Dot = 0.5; cosine = dot / (candidate norm × reference norm) = 1.
General cosine ranges −1..1; this nonnegative toy dictionary produces 0..1 when defined. A zero mean vector makes cosine Undefined. Learned metrics such as BERTScore use actual contextual embeddings and different token matching; these toy values cannot stand in for their results. High overlap or similarity does not establish equivalent meaning or factual correctness.
The finite model has three references and eight candidates, two lexical policies and fixed BLEU orders 1/2. The strings are authored English examples. Raw Exact match compares entire strings; Raw token metrics split on whitespace and preserve punctuation/case. Normalized lowercases, removes ASCII punctuation, removes English articles a/an/the, and collapses whitespace. These are explicit conventions, not a universal tokenizer or multilingual evaluation policy. One reference is selected; no multi-reference maximum or benchmark aggregation is used.
Token F1 uses multiset overlap, precision=matched/candidate length and recall=matched/reference length. If overlap is zero, F1 is 0, including two empty token lists. Two empty strings have Exact match 1. ROUGE-L here is sentence-level LCS F1 with beta=1, without stemming or summary-level union: 2×LCS/(candidate+reference length), with 0 for no ordered matches or empty input. These conventions must accompany comparisons.
BLEU uses clipped candidate n-gram precision and brevity penalty BP=exp(1−reference length/candidate length) for a nonempty shorter candidate, else BP=1 for a nonempty candidate at least as long as its reference. Empty candidate uses BP=0 and BLEU=0. BLEU-1=BP×p1; BLEU-2=BP×sqrt(p1×p2), with equal weights and no smoothing. If an order has no candidate n-grams, its precision is 0: an exact one-token pair still has BLEU-2=0 under this fixed-order convention. Other implementations may smooth or adjust effective order. This single-pair demonstration is not corpus BLEU-4; averaging sentence BLEU does not generally equal corpus BLEU. Punctuation, normalization, order, reference choice and aggregation all affect comparability.
Toy cosine always preprocesses original strings with Normalized, independently of Text normalization and BLEU order. It averages the disclosed dictionary vectors, including zero vectors, and divides their dot product by both vector norms; a zero norm is Undefined. It has no learned model, BERTScore computation, generation, calibrated correctness probability, factuality verifier, semantic guarantee or preferred universal score. Reference choice can penalize valid alternate wording. Presets restore their complete state; Reset and prediction changes restore Surface for the current experiment, while free Reset starts Experiment 1.
BLEU · Original clipped precision and brevity penalty