Implementation · EZKL

Evidence & limits

On this page

R2Demonstrated for proving a language model's output follows from committed weights, against a cheating prover

Public, independently audited code proves small models against a cheating prover, but its published results stop far below the size of the language models that verification claims concern.

Assessed use: proving a language model's output follows from committed weights, against a cheating prover

Rubric assessment

  • R1 met: the claim is that a stated model produced an output from a stated input, with either kept private 1. South et al. set out its use for verifiable evaluations of models with private weights 2.
  • R2 met: the code is public 1. The stated adversary is a prover who tries to convince a verifier of an incorrect result, and an independent audit tested for exactly that 3. South et al. report end-to-end proofs for models of up to about a million parameters 2.
  • R3 not met for this use, which is checking that a served language model is the declared one. The library is public and versioned 1, and Trail of Bits reports that other projects use its attestation and verifier contracts in production 3. No source says what those projects prove, and no party is documented relying on EZKL to verify a language model's outputs. The largest language model South et al. prove with it has 250,000 parameters 2, far below the models that claims about served AI concern. This matches the assessment of Zero-knowledge proofs of inference.
  • R4 not met, because R3 is not. Trail of Bits' 2025 audit would otherwise count: at its fix review no high-severity finding remained unresolved 3.
Gaps to the next level
  • A production-grade release that proves language models at the scale verification claims concern, or reliance by another party on such proofs for a verification decision.

Assessed 2026-10-08 against rubric v1.1.

Evidence

  • South et al., whose authors include two from EZKL, report proofs for eight models, from a 30-parameter linear regression to a 1.07-million-parameter VAE decoder 2. A 250,000-parameter nanoGPT took 2,781 s to prove and 2.69 s to verify, with a 219 GB proving key 2.
  • Trail of Bits reviewed EZKL and its verifier contracts for Zkonduit in January 2025, with 11 engineer-weeks of effort 3. It reported 34 findings, 8 of high severity, including circuit soundness bugs and ways to bypass data attestation. At its March 2025 fix review every high-severity finding was resolved 3.
  • Trail of Bits reports that other projects use EZKL's attestation and verifier contracts in production 3.

Limitations

  • South et al. write that proof time and resource requirements "grow dramatically with large models", and that the proving key's size limits model size 2.
  • The proof covers the quantized circuit. Trail of Bits showed that quantization can activate a backdoor that is dormant in the full-precision model 3.
  • The audit did not fully review several higher-level operations, such as einsum and scatter_nd, or all of the command-line tools 3.
  • Limits common to proofs of inference, such as proofs that do not bind the computation spent, are covered in Zero-knowledge proofs of inference.

Known flaws

Blockers

  • Proving cost grows steeply with model size: a 250,000-parameter nanoGPT took 2,781 s to prove and needed a 219 GB proving key, which South et al. name as the main limit on model size.

Search

Full search page