Mechanism · Zero-knowledge proofs of inference
Evidence & limits
On this page
R2Demonstrated for proving a language model's output follows from committed weights, against a cheating prover
R2 through zkLLM, which has public, artifact-evaluated code and peer-reviewed results at 13 billion parameters; no implementation is yet production-grade for language models.
Assessed use: proving a language model's output follows from committed weights, against a cheating prover
Rubric assessment
- R1 met: ZKML, zkLLM and NanoZK describe the claim (an output equals the committed model applied to an input) and their assumptions 1 3 4.
- R2 met through zkLLM. Its code is public and received CCS 2024 artifact-evaluation badges 2. Its published end-to-end results use a 13-billion-parameter model on a data-centre GPU, against a stated adversary: a cheating polynomial-time prover 1.
- R3 not met for this use. zkLLM's README says the code is not ready for industrial applications 2. Attestable's results come without public code, paper or reproducible artifacts 8. The ezkl library is public, and Trail of Bits reports that other projects use its verifier contracts in production 6. The language model South et al. prove with ezkl is a 250,000-parameter nanoGPT, and the largest model in their results is a VAE decoder of about 1.07 million parameters 5. Both are far below the scale this use concerns. Lagrange calls DeepProve "production-ready" 14, but its repository has no releases, and its reported benchmarks are GPT-2 and Gemma 3 at 512 tokens 15. No source documents a party other than a developer relying on any of these systems for a verification decision about a language model.
- R4 not met. Only ezkl has an independent audit: it left no high-severity finding unresolved 6. As of September 2026 no independent audit or red-team of zkLLM has been published. One independent analysis, tested on a small transformer, shows that valid proofs of LLM inference do not bind the computation spent, so a much smaller model can pass as the declared one 12.
- An implementation that is production-grade and available for language models, or relied on by a party other than its developer for a verification decision.
- An independent public security evaluation (audit, red-team or peer-reviewed analysis) of a proof system that handles language models.
Assessed 2026-10-08 against rubric v1.1.
Mechanism properties
| Threat model | Adversarial prover |
|---|---|
| Adversarial evaluation | Independent red-team |
| Hardware needed | None |
| Prover cooperation | Required |
| Confidentiality | Partial |
Evidence
- ZKML. It proved a distilled 81.3-million-parameter GPT-2 in 3,651.67 seconds on a 128-vCPU, 1 TB machine. This was the largest model its authors could prove with under 1 TB of RAM 3.
- South et al. They used ezkl to prove evaluations of small models. A 250,000-parameter nanoGPT took 2,781 seconds to prove and needed a 219 GB proving key 5.
- zkLLM. It proved OPT models up to 13B and LLaMa-2 models of 7B and 13B, each proof covering one 2,048-token forward pass on one A100 GPU. LLaMa-2-13B took 803 seconds to prove and 3.95 seconds to verify, with a 188 kB proof 1. Its code is public 2.
- zkGPT. It proved GPT-2 inference in 21.8 seconds with 32 threads on a 16-core CPU server, with a 101 KB non-interactive proof that verifies in 0.35 seconds. Its code is public 13.
- NanoZK. It reports proofs of 3.2 to 3.7 KB per attention or MLP sub-circuit, about 83 KB in total for 12 layers. It measured full transformer blocks on CPU only up to width 128. Assuming a GPU speedup, it projects that a 12-layer GPT-2 would take about 14 minutes to prove sequentially 4.
- Attestable. It reports proving a 31-billion-parameter Gemma model at 53 tokens per second for one 16K-token sequence on one H100 GPU, with 4.35 to 7.92 MiB proofs and 157 to 648 ms CPU verification. No code or paper accompanies these results 8.
- Lagrange DeepProve. Lagrange reports a zero-knowledge proof of full GPT-2 inference 14. Its public repository, under Lagrange's own licence, reports proving 512 tokens of GPT-2 in 7.6 minutes on a 24-core CPU server with 504 GB of memory, with a 10.7 MiB proof that verifies in 1.3 seconds 15.
- Independent audit of ezkl. Trail of Bits reviewed the ezkl library in January 2025, at its developer's request. It reported three high-severity circuit soundness issues and four ways to bypass data attestation in ezkl's smart contracts. Its fix review found every high-severity finding resolved 6.
- Hollow-LLM attack. Researchers at the University of Southern California ran a ghost-weight attack through zkGPT's proof procedure on a 6-layer, 512-dimensional transformer declared as up to 12 layers and 1,024 dimensions. Outputs were unchanged and serving cost stayed at the small model's level 12.
A verification system design for AI agreements lists ZKPs as a "tentative plan B" that, if they mature, could "remove the need for secure computing hardware setups" 11.
Limitations
Cost.
- zkLLM needs about 13 minutes (803 seconds) per 2,048-token forward pass at 13B scale on one A100 1.
- The verification system design calls the overhead "heavy" 11.
- The survey names "limited circuit expressiveness, high proving cost, and deployment complexity" as the main implementation bottlenecks 7.
Expressiveness. ZKML does not support branching or variable-length loops, so language models need fixed-length inputs 3. Floating-point emulation remains open 11. Attestable reports a 16K-token context limit 8.
Implementation soundness. In ezkl, Trail of Bits found circuits with missing constraints that "would allow a malicious prover to convince a verifier of incorrect calculations"; these were fixed 6. The zkLLM README says its code "has NOT undergone security auditing and is NOT ready for industrial applications". It also says prover and verifier run side by side, and that a deployment would need Fiat–Shamir to be non-interactive 2.
Quantisation. Trail of Bits also built a ResNet-18 backdoor that is dormant in the full-precision model and active after ezkl's quantisation. Whether it persists through proving was left for further investigation 6.
Coverage. A proof covers only the outputs proven 9. For accounting of other work, see Proofs of useful work for capacity accounting. A proof also does not show how much computation produced an output. In the Hollow-LLM attack, weights with the declared architecture let a much smaller model do the work, and the proofs remain valid 12.
Known flaws
Blockers
Showing that proven inference was the only work done needs a compute-accounting mechanism such as proof-of-work accounting, which is only proposed 9.