AI Grounds
Open AI Grounds on a desktop
These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.
AI Grounds
These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.
Guided discovery
Separate evidence, citations and coverage.
Retrieval-augmented generation (RAG) supplies retrieved source material to a generator before it answers. Here, authored source and answer bundles expose the evaluation steps.
Fictional question: When does the Harbor Museum open, and what is admission?
Authored current reference: Tuesday · 10:00 · 4 fictional credits. This is a teaching example, not information about a real venue.
The four sources have disclosed fact IDs. Support means the exact claim’s fact ID is listed by a retrieved source; no actual retriever, language model or semantic entailment judge runs here. “Entails” means the source supports that claim in this authored mapping. D4 is relevant but outdated; relevance does not imply freshness.
| Source | Exact authored text | Fact IDs | Retrieved? | Question-relevant? |
|---|---|---|---|---|
| D1 | Current schedule: the Harbor Museum opens Tuesday at 10:00. | day, time | Yes | Yes |
| D2 | Current admission: the Harbor Museum charges 4 fictional credits. | fee | Yes | Yes |
| D3 | Garden note: the garden has blue benches. | benches | No | No |
| D4 | Outdated schedule: the Harbor Museum opens Monday at 09:00. | oldDay, oldTime | No | Yes · Outdated |
Retrieved 2 documents · Relevant 2 · Answer claims 3 · Supported 3 · Present citations 3 · Valid citations 3 · Covered required facts 3/3.
Context relevance = relevant retrieved documents / all retrieved documents = 2/2 = 1. No retrieved documents makes this Undefined. This is a simple document fraction, not rank-aware context precision or an exact RAGAS implementation.
Each actual control/preset change clears stale explanation/transfer answers; reapplying the same state preserves them. Context, claims and citation labels can change independently.
Faithfulness uses any retrieved source. Citation support instead checks the source actually cited: it must be retrieved and individually entail the claim. A cited source may entail a claim in the full corpus but be absent from retrieved context; that citation is invalid in this toy convention.
| Claim / fact ID | Support in retrieved context | Cited source | Cited source retrieved? | Cited source entails? | Valid citation? |
|---|---|---|---|---|---|
| The Harbor Museum opens on Tuesday.day | D1 | D1 | Yes | Yes | Yes |
| The Harbor Museum opens at 10:00.time | D1 | D1 | Yes | Yes | Yes |
| Admission costs 4 fictional credits.fee | D2 | D2 | Yes | Yes | Yes |
Claims supported by any retrieved source.
1
supported answer claims / answer claims
3/3. Undefined for no claims. Relative to these sources, not guaranteed current truth.
Claims with a retrieved, individually entailing citation.
1
validly cited claims / answer claims
3/3. Coverage of all claims; support elsewhere does not validate the cited source.
Correct required current reference facts present.
1
covered required facts / 3
3/3. Empty answer is 0. Extra unsupported claims are not penalized by this coverage measure.
Citation precision = valid citations / present citations = 3/3 = 1. No citations makes precision Undefined. A nonempty answer with no citations still has Citation support 0. With one citation per claim, precision and claim coverage agree only when all claims have a citation.
| Required current fact | Fact ID | Present in answer? |
|---|---|---|
| The Harbor Museum opens on Tuesday. | day | Yes |
| The Harbor Museum opens at 10:00. | time | Yes |
| Admission costs 4 fictional credits. | fee | Yes |
A faithful fragment omits required facts. A faithful garden answer can be off-topic. An answer supported by outdated D4 can be well cited but disagree with the current reference. Conversely, a correct reference fact can appear without retrieved support. These scores are separate measurements; none is a complete model-quality verdict.
The finite model has five retrieved context bundles, six answer bundles and four citation rules: 120 states. The museum and credits are fictional. Claims are atomic disclosed fact IDs; entailment uses membership in a source’s authored fact list. No search, generation, natural-language inference or real-world fact verification runs. Contradiction versus missing support is not separately graded. The current authored reference requires Tuesday, 10:00 and 4 credits; outdated Monday/09:00 are supported by D4 but do not cover those current facts.
Context relevance counts question-relevant retrieved documents divided by retrieved documents, with Undefined for no documents. Relevance is not source freshness or truth and this is not rank-aware context precision. Faithfulness counts claims supported by any retrieved source divided by claims made, Undefined for an empty answer. Citation support counts claims with an optional citation that is both retrieved and individually entailing, divided by all claims. Precision uses present citations instead. Empty answers have Undefined coverage and precision; no citations gives Undefined precision and zero citation coverage for nonempty answers. Missing evidence never implies free admission. No fabricated perfect empty-answer score is supplied.
Matched sources are D1 for day/time, D2 for fee/free, D4 for old day/time and D3 for benches; D2 does not entail free admission. Misplaced cites D3 for all claims. Mixed matches the first two claims and cites D3 thereafter. There is at most one citation per claim; no multi-source joint entailment or redundant-citation evaluation is modeled. These are simplified teaching conventions rather than exact RAGAS or full ALCE metrics. Completeness is distinct correct required facts covered over three; it ignores unsupported extras and is not overall accuracy. Presets restore complete states; Reset/prediction restore Fully supported for the current exercise, free Reset starts Experiment 1.
RAGAs · Separate evaluation dimensions