Mechanism · Safeguard attestation
Technical detail
On this page
- Proof-of-guardrail protocol. (1) A wrapper program f bundles the public guardrail g, its configuration and the mediation of all agent inputs and outputs. (2) When f is loaded, the enclave records a measurement m, a hash that depends on the binary of f. (3) The private agent is loaded afterwards as a secret input, so it is not part of m. (4) For a user input x and response r, f returns a document signed with the platform's attestation key that contains m and d = Hash(x, r). (5) The verifier checks the certificate chain against the platform's published root, compares m with the expected measurement of the open-source f, and checks d 1.
- Prototype costs. On AWS Nitro Enclaves, the authors report 34% added latency on average over the same agent and guardrail run outside the enclave (24.8–38.0% across the four measured steps), 97.8 ms to generate an attestation and 5.1 ms to verify it. Holding the whole guardrail runtime in enclave memory needs an m5.xlarge instance, which costs 18.5 times as much per hour as a t3.micro 1.
- PAL*M. It defines single and session inference properties, r = M(M_tok(q)) and its multi-turn form over the chat history. It binds hashes of the query, tokenizer, model and response into an Intel TDX quote, together with an NVIDIA H100 attestation token and a verifier challenge 4. PAL*M's added time was 3.8–11.4% of total run time for session inference and 45.5–66.4% for single prompts, across Llama-3.1-8B, Gemma-3-4B and Phi-4-Mini, so one attested prompt took 1.8 to 3 times as long as on the same TDX machine without PAL*M's measurements 4.