Mechanism · Safeguard attestation

Attestation shows a safeguard ran, not that it is effective

On this page

← All known flaws

SignificantTheoretical argumentOpen

Proof of guardrail ensures that the guardrail executed, but the guardrail can still err or be jailbroken. Because the guardrail must be open source, a malicious developer can attack it with jailbreaks while still presenting a valid proof. In the authors' evaluation, Llama Guard 3 reached an F1 score of 0.56 on the unsafe class of the ToxicChat dataset. The authors state that proof of guardrail should not be interpreted or advertised as proof of safety.

Sources: [1]

Search

Full search page