The declared evaluation was run
The claim is that a reported test result came from running the agreed evaluation procedure on the specified model and test data.
Checking this lets an evaluator verify how the result was produced, while weights and tests can stay private 1 4. The evidence establishes execution under stated trust assumptions. Whether the tests adequately measure a risk or capability remains a separate question.
Attestable Audits binds the model, audit code, data and result in an enclave attestation 1. PAL*M measures the model, tokenizer and evaluation data and attests the operation that produced a metric 2.
Linking an evaluation to later service outputs needs evidence for served-model identity. In the PySyft pilot, proprietary model code could not all be inspected or allowlisted, so the participants accepted an additional trust assumption 4.
Evaluation integrity means that the declared model, test data and procedure produced the reported result. Attestable Audits describes that binding explicitly 1. PAL*M includes evaluation among its attested operations 2. Cove composes reviewed, attested workflow stages and their input and output certificates 3. These checks depend on the measured software and the hardware attestation roots. PySyft's double-blind pilot adds practical evidence for confidential evaluation, with participant-accepted opaque code and cloud verification assumptions 4. An accepted result still needs an interpretation of what the evaluation measures. A later deployment needs its own link to the evaluated model.
On this page
Why it matters
Checking an evaluation's execution links the reported result to a particular model, test data and procedure 1 2. That link can be checked while the weights and test data remain private 1 4.
Attestable Audits identifies two problems with ordinary benchmarks: results are not verifiable, and model weights and benchmark data may need to stay confidential 1. Its audit attestation binds the model hash, the hash of the audit code and data, and the result 1. PAL*M similarly defines evaluation evidence in terms of the operation, model, tokenizer, test data and metric 2.
Evaluation execution and subsequent service identity need separate evidence. Attestable Audits includes an inference protocol that checks the served model against the audited model and binds responses to the audit result 1. That second step addresses The declared model is the one being served.
Why it is hard
- Private weights and private tests can belong to parties that do not trust each other. Double-blind evaluation uses an attested enclave to run their agreed computation without sharing those inputs 4.
- Evidence must bind the relevant inputs and procedure to the result. Attesting the enclave's launch image is one step; Attestable Audits also measures the model and audit artifacts 1.
- Verifiers need trustworthy reference measurements and attestation roots. Cove's reference implementation trusts Intel TDX, the container runtime, pinned components and human review of the workflow 3.
- Opaque code can leave an accepted assumption inside that boundary. In the Gemini pilot, not all model code could be inspected or allowlisted. The participants also note that guest OS builds were not independently reproducible and Google's services signed and verified the attestation 4.