AI Grounds

Open AI Grounds on a desktop

These interactive lessons need a larger screen. Please continue on a desktop or laptop computer.

Guided discovery

Safety, Refusal & Robustness Metrics Lab

Compare safety decisions with their labeled denominators.

Your dataset

Eight authored prompt families have fixed labels: four benign requests should receive Help, four unsafe goals should receive Refusal. Help means appropriate help on a benign request; Harmful compliance means following an unsafe goal. Labels and outcomes are known by construction. No model, harmful response generation or actual jailbreak attack runs.

Policy and Prompt variant select a disclosed toy outcome bank. Reworded and Abstract wrapper represent alternate wording of the same goal; no actionable attack text is supplied. These are authored examples, not empirical performance or a universal safety policy.

Included 8 families · Benign 4 · Unsafe 4 · Harmful 1 · False refusals 1 · Originally refused unsafe 3 · Newly harmful among eligible 0 · All-variant correct 3 · All-variant consistent 4 · Selected refusals 4.

Each actual control/preset change clears stale explanation/transfer answers; reapplying the same state preserves them. Excluded rows remain visible for comparison and contribute to no numerator or denominator.

Selected column: Original wording. Help is correct for benign; Refusal is correct for unsafe. All-variant columns inspect the full row.
Family / goalLabelIncluded?Original wording · SelectedRewordedAbstract wrapperOriginally refused unsafe?Now harmful among eligible?All variants correct?All variants identical?
B1Plant-care questionBenignYesHelpHelpHelpNoNoYesYes
B2Debug a killed processBenignYesHelpRefusalHelpNoNoNoNo
B3Historical account of a warBenignYesHelpHelpHelpNoNoYesYes
B4Identify poisonous plantsBenignYesRefusalHelpHelpNoNoNoNo
U1Fraud goal (abstract)UnsafeYesRefusalRefusalRefusalYesNoYesYes
U2Targeted harassment goal (abstract)UnsafeYesRefusalHarmful complianceHarmful complianceYesNoNoNo
U3Unauthorized access goal (abstract)UnsafeYesRefusalRefusalHarmful complianceYesNoNoNo
U4Violence goal (abstract)UnsafeYesHarmful complianceHarmful complianceHarmful complianceNoNoNoYes

Harmful compliance

Unsafe requests followed in the selected variant.

0.25

harmful included unsafe / included unsafe

1/4. Undefined when the unsafe class is absent.

False refusal

Benign requests refused in the selected variant.

0.25

refused included benign / included benign

1/4. Undefined when the benign class is absent.

Conditional jailbreak success

Originally refused unsafe families now harmful.

0

newly harmful eligible / eligible unsafe

0/3. Eligibility is Original Refusal under the same policy; originally harmful families are excluded.

Correct across variants, or just consistent?

Robust success = included families correctly handled in all three variants / included families = 3/8 = 0.375. A benign family must receive Help in every variant; an unsafe family must receive Refusal in every variant.

Decision consistency = included families with identical outcomes across all three variants / included families = 4/8 = 0.5. Consistently refusing benign requests or consistently following unsafe goals is still wrong. Selecting a different variant changes current error rates but cannot change either all-variant measure.

Conditional jailbreak success is an explicitly conditional attack success rate (ASR), not a universal benchmark convention. In Original wording it is 0 by construction: eligible families are precisely those originally refused. That does not exclude original harmful compliance outside its denominator. All-unsafe harmful compliance includes those families. A zero error rate on four authored cases does not prove real-model safety, and an absent class produces Undefined rather than a fabricated perfect score.

Construction and limits

The bank has three authored policies, three prompt variants and three nonempty evaluation subsets: 27 states. All eight labeled families and outcomes are fixed teaching examples; the names Balanced, Strict and Permissive describe this bank, not deployed defenses. Labels do not resolve context-dependent real-world safety judgments. Partial refusals, safe redirection, harmful detail severity, adaptive attackers, multi-attempt budgets, stochastic variation, uncertainty intervals and judge reliability are out of scope.

Harmful compliance and false refusal use distinct classes in the selected variant. Conditional ASR counts only unsafe families originally refused under the same policy and subset. Changing policy also changes eligibility; it does not establish a causal attack ranking across real systems. Robust success is correctness in all three disclosed variants, not a guarantee for unseen prompts. Consistency ignores correctness. All fractions with a zero denominator are Undefined. Excluded rows remain visible but do not contribute. Reset/prediction restore Balanced baseline for the current exercise; free Reset starts Experiment 1.

XSTest · Safe requests and excessive refusal
JailbreakBench · Explicit evaluation conventions