Mechanism · Zero-knowledge proofs of training constraints

Evidence & limits

On this page

R2Demonstrated for proving a training run followed a committed specification and data

Peer-reviewed end-to-end results exist for small models against a stated adversary, and public code proves single fine-tuning steps of 13-billion-parameter language models. The frontier-scale design is unbuilt.

Assessed use: proving a training run followed a committed specification and data

Rubric assessment

  • R1 met: Kaizen 1 and ZkAudit 2 define proofs of correct training on committed data. Peigné et al. describe a frontier design with stated claims, trust anchors and open problems 4.
  • R2 met through reproducible published end-to-end results and a public working implementation. Kaizen and ZkAudit are peer-reviewed, specify the protocol, setup and parameters, and state a cheating prover as the adversary. Kaizen measures proving per training iteration of a 10-million-parameter VGG-11, with recursive aggregation implemented 1. ZkAudit proves single SGD steps of MobileNet v2 and recommender models on AWS g4dn.8xlarge instances, and estimates the cost of proving full training runs 2. Kaizen links no code 1, and ZkAudit links only an anonymised review repository 2; the rubric does not require code on this route. VeriLoRA, peer-reviewed and with public code, proves one LoRA fine-tuning step on a single sample for LLaMA and OPT models of 3 to 13 billion parameters 3. All of these results are far below frontier training.
  • R3 not met: as of September 2026 no deployment or reliance by a third party has been published. The frontier design is unimplemented, and its authors present its costs as estimates, with target values "not yet measured" 4.
  • R4 not met: no independent evaluation has been published.
Gaps to the next level
  • Any use by a party other than the developer, or a production-grade, available implementation.
  • An independent public security evaluation of a proof-of-training system.
  • For the frontier use: an implementation of challenge-based step proofs at realistic model and cluster scale.

Assessed 2026-09-25 against rubric v1.1.

Mechanism properties

Threat modelAdversarial prover
Adversarial evaluationAnalysis
Hardware neededNone
Prover cooperationRequired
ConfidentialityPartial

Evidence

  • Kaizen (CCS 2024). It proves training of a 10-million-parameter VGG-11 on CIFAR-10 at batch size 16. The prover takes 15 minutes per iteration; the proof is 1.63 MB and verifies in 130 milliseconds, independent of the number of iterations 1. Its authors report "24× faster prover time" than generic recursive proof systems 1.
  • ZkAudit (ICML 2024). It proved single SGD steps for MobileNet v2 image classifiers and a recommender model on AWS g4dn.8xlarge instances. Proving one step on a single image took 47.5 to 328.3 seconds for MobileNet v2 (1.0, 224), depending on the fixed-point scale factor. The authors estimated, rather than generated, proofs of full training runs, at costs of hundreds to thousands of dollars 2.
  • VeriLoRA (NDSS 2026). It proves one LoRA fine-tuning step, covering the forward pass, backward pass and parameter update, on a single sample, for LLaMA and OPT models of 3 to 13 billion parameters on one A100 GPU. Proving a step takes minutes and verifying it takes seconds, and the code is public 3.
  • Other systems. The survey lists further verifiable-training systems in its Table IV 5.
  • The frontier design. Peigné et al. estimate 2 to 10% training-side overhead for a Llama 3.1 405B-scale run, and "a deployable proof of concept within approximately 36 months" 4. The paper reports no prototype or measurements of its own, and marks its target values as "not yet measured" 4.

The zkLLM authors wrote in 2024 that zero-knowledge proofs of LLM training "may pose insurmountable challenges" 6.

Limitations

Cost. A 10-million-parameter model needs minutes of proving per step 1. ZkAudit's authors leave scaling to larger models, such as language models, to future work 2. VeriLoRA reaches 13-billion-parameter language models, but for single-sample steps that update only low-rank adapters 3.

Open problems in the frontier design. Peigné et al. list 13, including 4:

  • zero-knowledge proofs of backpropagation;
  • deterministic attention backward passes with under 5% overhead;
  • an open-hardware network tap at line rate;
  • a way to tell silent data corruption apart from adversarial deviation;
  • coverage of mixture-of-experts, reinforcement-learning post-training and multi-site training.

They also report that current deterministic tensor-parallel all-reduce configurations lose 64 to 89% of bandwidth 4.

Attacks. As of September 2026 no attack on these proof systems has been published.

Known flaws

Blockers

Search

Full search page