Mechanism · Zero-knowledge proofs of inference
Technical detail
On this page
Systems encode a network as an arithmetic circuit, or as sumcheck and lookup relations, over a finite field, with tensors scaled to fixed point 1 3.
- ZKML (EuroSys 2024) compiles models to halo2 circuits with KZG commitments (trusted setup) or IPA commitments (transparent), and optimises circuit layout. A distilled GPT-2 (81.3M parameters) took 3,651.67 s to prove with KZG, 18.70 s to verify and a 28,128-byte proof on a 128-vCPU, 1 TB machine 3.
- ezkl also targets halo2 with KZG. South et al. proved a 250,000-parameter nanoGPT in 2,781 s, verified it in 2.69 s and needed a 219 GB proving key 5. A Trail of Bits audit of ezkl (January 2025, commit bdcba5c) reported 34 findings, 8 of high severity; three high-severity findings were circuit soundness issues, and all high-severity findings were resolved at the March 2025 fix review 6.
- zkLLM (CCS 2024) uses sumcheck-based arguments, a parallel lookup argument (tlookup) for non-arithmetic tensor operations, an attention protocol (zkAttn) and Hyrax commitments over BLS12-381, with values scaled by 2^16. On one A100 40 GB GPU with 2,048-token inputs, LLaMa-2-13B needed 986 s for the one-time weight commitment, 803 s to prove, 188 kB of proof, 3.95 s to verify and 23.1 GB of memory. Perplexity on C4 moved from 6.520 to 6.528 1.
- zkGPT (USENIX Security 2025) combines the GKR protocol, Lasso lookups and Hyrax commitments, and is made non-interactive with Fiat–Shamir. On a 16-core Xeon server with 200 GB of memory, it proved GPT-2 inference in 21.8 s with 32 threads, with a 101 KB proof verified in 0.35 s. Its code is archived on Zenodo 13.
- NanoZK (ICICS 2026) proves each transformer layer separately with Halo2/IPA over the Pallas curve, uses 16-bit lookup tables for softmax, GELU and normalisation, and chains layer boundaries with SHA-256 commitments. It reports 3.2 to 3.7 KB per sub-circuit (attention or MLP) proof, about 83 KB in total for 12 layers. It measured full-block proofs on CPU up to width 128 and, assuming a GPU speedup, projects about 68 s per block at GPT-2 width, or about 14 minutes to prove a 12-layer GPT-2 sequentially 4.
- Attestable reports a prover whose security rests only on hash functions, at 100-bit security, with 8-bit integer matrix multiplications. On one H100 it reports 4.35 to 7.92 MiB proofs, 157 to 648 ms CPU verification, and 53 tokens per second for one 16K-token sequence of a 31B Gemma model 8.