Implementation · PySyft double-blind evaluations

Evidence & limits

On this page

R2Demonstrated for evaluating a private model on private prompts, neither party seeing the other's inputs

A participant report describes an end-to-end evaluation with private assets on commercial GPU hardware; PySyft's double-blind workflow has not been shown as a generally available service.

Assessed use: evaluating a private model on private prompts, neither party seeing the other's inputs

Rubric assessment

  • R1 met: the report states the mutual-confidentiality claim, the enclave trust assumptions and the submission and approval procedure 1.
  • R2 met: AVERI evaluated Gemini 2.5 Flash Lite with private MLCommons prompts on an H100 with Intel TDX and PySyft v0.10.x. The report gives the stack and workflow, and reports a separate evaluation with Singapore AISI that used private prompts 1.
  • R3 not met for this implementation: OpenMined documents a pilot with real private assets, but not an available production service or reliance on its result for a verification decision 1 2.
  • R4 not met: the participant report includes no independent public security evaluation of the workflow. Confidence is medium. The demonstration and its limits come from the participants' own report 1.
Gaps to the next level
  • A generally available production workflow, or documented reliance by another party on its result for a verification decision.
  • An independent public security evaluation that leaves no critical flaw open.

Assessed 2026-09-25 against rubric v1.1.

Evidence

  • In the 2026 pilot, AVERI evaluated Gemini 2.5 Flash Lite on private MLCommons AILuminate prompts in Google Cloud Confidential Space. The instance used one NVIDIA H100 with Intel TDX and PySyft v0.10.x. AVERI staff decrypted and scored the outputs 1.
  • The same report describes a separate evaluation with Singapore AISI, using a private prompt set focused on harmful content in Singapore's context 1.

Limitations

  • The pilot's authors could not inspect or allowlist all model code. AVERI accepted that condition 1.
  • The guest operating system builds were not independently reproducible, and Google's services signed and verified the attestation 1.
  • The TEE findings report physical-host attacks on the pilot's TDX and H100 hardware class 3 4. The pilot report does not evaluate this boundary 1.
  • The published evaluation used one H100. The authors name many-node confidential GPU clusters as a next step 1.

Known flaws

Blockers

Search

Full search page