AI Grounds

Open AI Grounds on a desktop

These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.

Guided discovery

RAG Groundedness Metrics Lab

Separate evidence, citations and coverage.

Your dataset

Retrieval-augmented generation (RAG) supplies retrieved source material to a generator before it answers. Here, authored source and answer bundles expose the evaluation steps.

Fictional question: When does the Harbor Museum open, and what is admission?

Authored current reference: Tuesday · 10:00 · 4 fictional credits. This is a teaching example, not information about a real venue.

The four sources have disclosed fact IDs. Support means the exact claim’s fact ID is listed by a retrieved source; no actual retriever, language model or semantic entailment judge runs here. “Entails” means the source supports that claim in this authored mapping. D4 is relevant but outdated; relevance does not imply freshness.

Retrieved context selects sources; it does not regenerate the answer.
SourceExact authored textFact IDsRetrieved?Question-relevant?
D1Current schedule: the Harbor Museum opens Tuesday at 10:00.day, timeYesYes
D2Current admission: the Harbor Museum charges 4 fictional credits.feeYesYes
D3Garden note: the garden has blue benches.benchesNoNo
D4Outdated schedule: the Harbor Museum opens Monday at 09:00.oldDay, oldTimeNoYes · Outdated

Retrieved 2 documents · Relevant 2 · Answer claims 3 · Supported 3 · Present citations 3 · Valid citations 3 · Covered required facts 3/3.

Context relevance = relevant retrieved documents / all retrieved documents = 2/2 = 1. No retrieved documents makes this Undefined. This is a simple document fraction, not rank-aware context precision or an exact RAGAS implementation.

Each actual control/preset change clears stale explanation/transfer answers; reapplying the same state preserves them. Context, claims and citation labels can change independently.

Trace each answer claim

Faithfulness uses any retrieved source. Citation support instead checks the source actually cited: it must be retrieved and individually entail the claim. A cited source may entail a claim in the full corpus but be absent from retrieved context; that citation is invalid in this toy convention.

One optional citation per claim. All support and reference mappings are authored, not model judgments.
Claim / fact IDSupport in retrieved contextCited sourceCited source retrieved?Cited source entails?Valid citation?
The Harbor Museum opens on Tuesday.dayD1D1YesYesYes
The Harbor Museum opens at 10:00.timeD1D1YesYesYes
Admission costs 4 fictional credits.feeD2D2YesYesYes

Faithfulness

Claims supported by any retrieved source.

1

supported answer claims / answer claims

3/3. Undefined for no claims. Relative to these sources, not guaranteed current truth.

Citation support

Claims with a retrieved, individually entailing citation.

1

validly cited claims / answer claims

3/3. Coverage of all claims; support elsewhere does not validate the cited source.

Completeness

Correct required current reference facts present.

1

covered required facts / 3

3/3. Empty answer is 0. Extra unsupported claims are not penalized by this coverage measure.

Check the denominator and the question

Citation precision = valid citations / present citations = 3/3 = 1. No citations makes precision Undefined. A nonempty answer with no citations still has Citation support 0. With one citation per claim, precision and claim coverage agree only when all claims have a citation.

Completeness counts distinct correct required facts, independently of source support or citation placement.
Required current factFact IDPresent in answer?
The Harbor Museum opens on Tuesday.dayYes
The Harbor Museum opens at 10:00.timeYes
Admission costs 4 fictional credits.feeYes

A faithful fragment omits required facts. A faithful garden answer can be off-topic. An answer supported by outdated D4 can be well cited but disagree with the current reference. Conversely, a correct reference fact can appear without retrieved support. These scores are separate measurements; none is a complete model-quality verdict.

Construction and limits

The finite model has five retrieved context bundles, six answer bundles and four citation rules: 120 states. The museum and credits are fictional. Claims are atomic disclosed fact IDs; entailment uses membership in a source’s authored fact list. No search, generation, natural-language inference or real-world fact verification runs. Contradiction versus missing support is not separately graded. The current authored reference requires Tuesday, 10:00 and 4 credits; outdated Monday/09:00 are supported by D4 but do not cover those current facts.

Context relevance counts question-relevant retrieved documents divided by retrieved documents, with Undefined for no documents. Relevance is not source freshness or truth and this is not rank-aware context precision. Faithfulness counts claims supported by any retrieved source divided by claims made, Undefined for an empty answer. Citation support counts claims with an optional citation that is both retrieved and individually entailing, divided by all claims. Precision uses present citations instead. Empty answers have Undefined coverage and precision; no citations gives Undefined precision and zero citation coverage for nonempty answers. Missing evidence never implies free admission. No fabricated perfect empty-answer score is supplied.

Matched sources are D1 for day/time, D2 for fee/free, D4 for old day/time and D3 for benches; D2 does not entail free admission. Misplaced cites D3 for all claims. Mixed matches the first two claims and cites D3 thereafter. There is at most one citation per claim; no multi-source joint entailment or redundant-citation evaluation is modeled. These are simplified teaching conventions rather than exact RAGAS or full ALCE metrics. Completeness is distinct correct required facts covered over three; it ignores unsupported extras and is not overall accuracy. Presets restore complete states; Reset/prediction restore Fully supported for the current exercise, free Reset starts Experiment 1.

RAGAs · Separate evaluation dimensions
ALCE · Citation evaluation
Correctness and faithfulness in question answering