Declared safeguards were applied during inference
On this page
Mechanisms
- R2Safeguard attestationprimaryAttests that a measured safeguard program (guardrail, filter, monitor) mediated the attested responses; coverage of all traffic is not established.
- R3Model identity attestation⚠supportingLinks an attested evaluation to the model later served 3.
- Can attest that measured policy software, such as filters and logging, wrapped the model 11. The evaluation variant is Attestable Audits.
- R2Confidential multi-party verificationsupportingPlan-scoped monitoring runs an agreed classifier over private usage records 5.
- R2Zero-knowledge proofs of inferencesupportingAttestable proposes that a proof could show an agreed input classifier was applied; South et al. prove evaluation results.
- Could require deployment-time safeguards on approved devices 1.
Implementations
- Assessed use: proving an output came from committed weightsProposed use: showing an agreed input classifier was applied.
- Assessed use: screening challenged records to show declared inference compute is not trainingScreening checks that outputs are free of blacklisted uses, including with inspector agents.