{
  "schema_version": "1.4.0",
  "rubric_version": "1.1",
  "license": "CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)",
  "record": {
    "id": "C-0011",
    "slug": "declared-evaluation-was-run",
    "title": "The declared evaluation was run",
    "aliases": [
      "evaluation integrity",
      "evaluation authenticity",
      "verifiable evaluations"
    ],
    "status": "draft",
    "last_reviewed": "2026-10-08",
    "review_interval_days": 90,
    "steward": null,
    "provenance": {
      "drafted_by": "ai",
      "reviewed_by": []
    },
    "risk_flags": [],
    "flags": [],
    "one_liner": "The reported test result comes from running the agreed evaluation procedure on the specified model and test data.",
    "summary": "The claim is that a reported test result came from running the agreed evaluation procedure on the specified model and test data. Checking this lets an evaluator verify how the result was produced, while weights and tests can stay private [[S-0009]] [[S-3320]]. The evidence establishes execution under stated trust assumptions. Whether the tests adequately measure a risk or capability remains a separate question.\n\n[[I-0007|Attestable Audits]] binds the model, audit code, data and result in an enclave attestation [[S-0009]]. [[I-0022|PAL\\*M]] measures the model, tokenizer and evaluation data and attests the operation that produced a metric [[S-0012]]. Linking an evaluation to later service outputs needs evidence for [[C-0005|served-model identity]]. In [[I-0017|the PySyft pilot]], proprietary model code could not all be inspected or allowlisted, so the participants accepted an additional trust assumption [[S-3320]].",
    "claim_class": "positive",
    "editors_synthesis": {
      "assessment": true,
      "markdown": "Evaluation integrity means that the declared model, test data and procedure produced the reported result. [[I-0007|Attestable Audits]] describes that binding explicitly [[S-0009]]. [[I-0022|PAL\\*M]] includes evaluation among its attested operations [[S-0012]]. [[I-0023|Cove]] composes reviewed, attested workflow stages and their input and output certificates [[S-1505]]. These checks depend on the measured software and the hardware attestation roots. [[I-0017|PySyft's double-blind pilot]] adds practical evidence for confidential evaluation, with participant-accepted opaque code and cloud verification assumptions [[S-3320]]. An accepted result still needs an interpretation of what the evaluation measures. A later deployment needs its own link to [[C-0005|the evaluated model]].",
      "text": "Evaluation integrity means that the declared model, test data and procedure produced the reported result. Attestable Audits describes that binding explicitly [S-0009]. PAL*M includes evaluation among its attested operations [S-0012]. Cove composes reviewed, attested workflow stages and their input and output certificates [S-1505]. These checks depend on the measured software and the hardware attestation roots. PySyft's double-blind pilot adds practical evidence for confidential evaluation, with participant-accepted opaque code and cloud verification assumptions [S-3320]. An accepted result still needs an interpretation of what the evaluation measures. A later deployment needs its own link to the evaluated model."
    },
    "sources": [
      {
        "source": "S-0009",
        "supports": "audit attestation binds model, audit code and data, and result; separate inference protocol",
        "locator": "§3; Algorithms 2–3"
      },
      {
        "source": "S-0012",
        "supports": "proof of evaluation binds model, tokenizer, evaluation data and metric; hardware and threat model",
        "locator": "§3.2; §4.3.3; Table 5"
      },
      {
        "source": "S-1505",
        "supports": "Cove certificates, reviewed node manifests and trust assumptions",
        "locator": "docs/internal/architecture.md; docs/internal/security_model.md"
      },
      {
        "source": "S-3320",
        "supports": "double-blind evaluation workflow and pilot; accepted proprietary code; guest OS and cloud verification assumptions",
        "locator": "§2.5; §3; §4"
      }
    ],
    "concepts": [
      "K-0004",
      "K-0006",
      "K-0024"
    ],
    "order": 11,
    "type": "claim",
    "goals": [
      {
        "id": "G-0003",
        "title": "Deploy only evaluated models",
        "url": "https://trustbutveri.fyi/goals/deploy-only-evaluated-models/",
        "relevance": "direct"
      }
    ],
    "url": "https://trustbutveri.fyi/claims/declared-evaluation-was-run/",
    "source_file": "content/claims/declared-evaluation-was-run.md",
    "flags_all": [],
    "body_markdown": "## Why it matters\n\nChecking an evaluation's execution links the reported result to a particular model, test data and procedure\n[[S-0009]] [[S-0012]]. That link can be checked while the weights and test data remain private\n[[S-0009]] [[S-3320]].\n\n[[I-0007|Attestable Audits]] identifies two problems with ordinary benchmarks: results are not verifiable, and model\nweights and benchmark data may need to stay confidential [[S-0009]]. Its audit attestation binds the model\nhash, the hash of the audit code and data, and the result [[S-0009]]. [[I-0022|PAL\\*M]] similarly defines evaluation\nevidence in terms of the operation, model, tokenizer, test data and metric [[S-0012]].\n\nEvaluation execution and subsequent service identity need separate evidence. [[I-0007|Attestable Audits]] includes\nan inference protocol that checks the served model against the audited model and binds responses to the\naudit result [[S-0009]]. That second step addresses [[C-0005]].\n\n## Why it is hard\n\n- Private weights and private tests can belong to parties that do not trust each other. Double-blind\n  evaluation uses an attested enclave to run their agreed computation without sharing those inputs\n  [[S-3320]].\n- Evidence must bind the relevant inputs and procedure to the result. Attesting the enclave's launch image\n  is one step; [[I-0007|Attestable Audits]] also measures the model and audit artifacts [[S-0009]].\n- Verifiers need trustworthy reference measurements and attestation roots. [[I-0023|Cove]]'s reference implementation\n  trusts Intel TDX, the container runtime, pinned components and human review of the workflow [[S-1505]].\n- Opaque code can leave an accepted assumption inside that boundary. In the Gemini pilot, not all model\n  code could be inspected or allowlisted. The participants also note that guest OS builds were not\n  independently reproducible and Google's services signed and verified the attestation [[S-3320]].",
    "body_text": "Why it matters Checking an evaluation's execution links the reported result to a particular model, test data and procedure [S-0009] [S-0012]. That link can be checked while the weights and test data remain private [S-0009] [S-3320]. Attestable Audits identifies two problems with ordinary benchmarks: results are not verifiable, and model weights and benchmark data may need to stay confidential [S-0009]. Its audit attestation binds the model hash, the hash of the audit code and data, and the result [S-0009]. PAL*M similarly defines evaluation evidence in terms of the operation, model, tokenizer, test data and metric [S-0012]. Evaluation execution and subsequent service identity need separate evidence. Attestable Audits includes an inference protocol that checks the served model against the audited model and binds responses to the audit result [S-0009]. That second step addresses The declared model is the one being served. Why it is hard - Private weights and private tests can belong to parties that do not trust each other. Double-blind evaluation uses an attested enclave to run their agreed computation without sharing those inputs [S-3320]. - Evidence must bind the relevant inputs and procedure to the result. Attesting the enclave's launch image is one step; Attestable Audits also measures the model and audit artifacts [S-0009]. - Verifiers need trustworthy reference measurements and attestation roots. Cove's reference implementation trusts Intel TDX, the container runtime, pinned components and human review of the workflow [S-1505]. - Opaque code can leave an accepted assumption inside that boundary. In the Gemini pilot, not all model code could be inspected or allowlisted. The participants also note that guest OS builds were not independently reproducible and Google's services signed and verified the attestation [S-3320].",
    "addressed_by": [
      {
        "id": "M-0025",
        "title": "Confidential multi-party verification",
        "url": "https://trustbutveri.fyi/mechanisms/confidential-multi-party-verification/",
        "role": "primary",
        "note": "Attestable Audits binds the model, audit code and data, and result. PySyft's pilot ran an agreed private evaluation with participant-accepted opaque-code and cloud trust assumptions (S-0009, S-3320)."
      },
      {
        "id": "M-0008",
        "title": "TEE remote attestation for AI workloads",
        "url": "https://trustbutveri.fyi/mechanisms/tee-remote-attestation/",
        "role": "supporting",
        "note": "Measures the enclave software and supplies the hardware root for evaluation attestations. Attestable Audits additionally binds model, audit and result (S-0009); enclave launch attestation alone does not establish the full evaluation claim."
      },
      {
        "id": "I-0007",
        "title": "Attestable Audits",
        "url": "https://trustbutveri.fyi/implementations/attestable-audits/",
        "role": "primary",
        "note": "The audit protocol binds the model hash, audit code and data, and result in a published attestation (S-0009)."
      },
      {
        "id": "I-0023",
        "title": "Cove",
        "url": "https://trustbutveri.fyi/implementations/cove/",
        "role": "primary",
        "note": "Reviewed workflow manifests and attested certificates bind evaluation stages to admitted inputs, results and predecessor evidence, under the documented host and review assumptions (S-1505)."
      },
      {
        "id": "I-0022",
        "title": "PAL*M",
        "url": "https://trustbutveri.fyi/implementations/palm/",
        "role": "primary",
        "note": "Evaluation attestations bind the operation, model, tokenizer, test dataset and reported metric (S-0012)."
      },
      {
        "id": "I-0017",
        "title": "PySyft double-blind evaluations",
        "url": "https://trustbutveri.fyi/implementations/pysyft-double-blind-evaluations/",
        "role": "primary",
        "note": "Both parties approve an attested evaluation on their submitted assets. The pilot accepted proprietary model code that could not all be inspected or allowlisted; it does not establish ongoing serving identity (S-3320)."
      }
    ],
    "referenced_by": [
      {
        "id": "G-0003",
        "title": "Deploy only evaluated models",
        "url": "https://trustbutveri.fyi/goals/deploy-only-evaluated-models/"
      },
      {
        "id": "M-0025",
        "title": "Confidential multi-party verification",
        "url": "https://trustbutveri.fyi/mechanisms/confidential-multi-party-verification/"
      },
      {
        "id": "M-0008",
        "title": "TEE remote attestation for AI workloads",
        "url": "https://trustbutveri.fyi/mechanisms/tee-remote-attestation/"
      },
      {
        "id": "I-0007",
        "title": "Attestable Audits",
        "url": "https://trustbutveri.fyi/implementations/attestable-audits/"
      },
      {
        "id": "I-0023",
        "title": "Cove",
        "url": "https://trustbutveri.fyi/implementations/cove/"
      },
      {
        "id": "I-0022",
        "title": "PAL*M",
        "url": "https://trustbutveri.fyi/implementations/palm/"
      },
      {
        "id": "I-0017",
        "title": "PySyft double-blind evaluations",
        "url": "https://trustbutveri.fyi/implementations/pysyft-double-blind-evaluations/"
      }
    ]
  }
}