Mechanism · Safeguard attestation

Evidence & limits

On this page

R2Demonstrated for attesting that a declared safeguard mediated a service's responses

R2, narrowly: one prototype with public code attests a guardrail end to end on cloud enclaves. No independent security evaluation or reliance by another party has been documented.

Assessed use: attesting that a declared safeguard mediated a service's responses

Rubric assessment

  • R1 met: designs that state the claim (a response was produced after a specific guardrail ran) and their trust assumptions are published 1, alongside designs for attesting inference properties 4 and plan-scoped monitoring 6.
  • R2 met through the proof-of-guardrail prototype. Its code is public 2, and its authors report end-to-end results on production cloud enclave hardware (AWS Nitro Enclaves) against a stated adversary, a developer who skips or modifies the guardrail 1. PAL*M adds attested session inference on Intel TDX with an H100 GPU, but its properties do not cover safeguards 4.
  • R3 not met: no party other than a developer is documented as relying on safeguard attestation, and the proof-of-guardrail code is described as a proof of concept that is not production-ready 2. Tinfoil reports that it is rolling out enclave-run safeguards in its chat service and can prove to an auditor or regulator that they are running, but as of mid-September 2026 the rollout was still under way 3.
  • R4 not met: no independent audit or red-team has been published, and the monitoring prototype has not been stress-tested by a counterparty 7.

Confidence is low because the demonstration is far from a frontier serving stack. It reaches both its guardrail model and the agent's backend model through external APIs, attests only the responses for which attestation is offered, and runs on CPU enclaves 1. Its README states that the enclave does not yet restrict the agent's command execution, which could be used to bypass the guardrail 2. Intel TDX attestations have also been forged by attackers with physical access to the memory bus 12 13.

Gaps to the next level
  • Reliance by a party other than the developer on safeguard attestations for a verification decision, or a production-grade system that is generally available.
  • An independent public evaluation (audit, red-team or peer-reviewed security analysis) of a safeguard-attestation system.
  • A demonstration in which the safeguard model itself runs inside the attested boundary on GPU hardware at a realistic serving scale.
  • A published way to show that all of a service's traffic, not only attested responses, passed through the attested safeguard path.

Assessed 2026-09-25 against rubric v1.1.

Evidence

  • Proof of guardrail. The authors implemented it for OpenClaw agents on AWS Nitro Enclaves, with Llama Guard 3 for content safety and a fact-checking tool 1. They report 34% added latency on average compared with running outside the enclave 1. The code is public; its README calls it a proof of concept that is not production-ready 2.
  • Tinfoil safeguards. Tinfoil reports that it is rolling out safeguard models, gpt-oss-safeguard with Kimi-K3 reviewing flagged cases, inside its secure enclaves for Tinfoil Chat 3. It states that the whole pipeline is open source and attested, so that it can prove to an auditor, a regulator or anyone else that the safeguards are running and what the policy is 3. Only the fact of a violation, with the account and conversation ID, leaves the enclave 3.
  • PAL*M. It is implemented on Intel TDX with an NVIDIA H100 4. The authors report that PAL*M's added time was 3.8–11.4% of total run time in session inference, and they modelled the protocol formally with the Tamarin prover 4.
  • Attestable Audits. A prototype ran a 4-bit 8-billion-parameter model in CPU-only AWS Nitro Enclaves at 1.84 tokens per second 5.
  • Scoped monitoring. A reference implementation runs in Tinfoil confidential virtual machines 7. Its authors state that user data and plan execution in the demo are not actually secure, and that the system has not been stress-tested by a counterparty 7.

Limitations

  • Effectiveness gap. An attested guardrail can still be jailbroken 1.
  • Selective attestation. Provers might attest only favourable executions 1 4.
  • External APIs. The prototype called its guardrail model and backend model through external APIs 1.
  • Hardware attacks. With physical access to the memory bus, researchers forged Intel TDX attestations and, by pairing them with relayed H100 attestations, made a workload outside TEE protection pass as GPU confidential computing 12. A second team forged TDX attestation reports with an active interposer 13. Other attacks forged AMD SEV-SNP attestations, one through a DDR4 interposer and one from software alone on platforms without AMD's fix 14 15.
  • Scale. Frontier model inference typically needs several GPUs 8. CPU inference, which the enclave prototype had to use, cost 21.7 times as much per token as GPU inference and ran about 100 times slower 5.

For confidential workflows that combine several parties' private inputs, see Confidential multi-party verification.

Known flaws

Blockers

Search

Full search page