Evidence & limits
On this page
R2Demonstrated for proving an output came from committed weights, against a prover who cheats
Public, artifact-evaluated code proves 13-billion-parameter models in peer-reviewed tests, but its authors say it is unaudited and not ready for production.
Assessed use: proving an output came from committed weights, against a prover who cheats
Rubric assessment
- R1 met: the paper states the claim, the threat model and the security theorems 1.
- R2 met: the code is public, tagged and archived on Zenodo. It received the CCS 2024 badges "Artifacts Available" and "Artifacts Evaluated--Functional" 2. The published end-to-end results use 13-billion-parameter models on a data-centre GPU 1. The stated adversary is a cheating polynomial-time prover 1.
- R3 not met: the README says the code "is NOT ready for industrial applications" and is no longer maintained 2. As of September 2026 no third party is known to rely on it.
- R4 not met: the README says the code "has NOT undergone security auditing" 2, and as of September 2026 no independent security evaluation of it has been published.
- A production-grade implementation, with prover and verifier separated and non-interactive proofs, or reliance by a third party for a verification decision.
- An independent public security evaluation (audit, red-team or third-party peer-reviewed analysis).
Assessed 2026-09-25 against rubric v1.1.
Evidence
- Setup. The paper reports results on one NVIDIA A100 GPU with 40 GB of memory, 12 CPU cores and 124.5 GB of system memory 1. The models were OPT (125M to 13B) and LLaMa-2 (7B and 13B), run on 2,048-token samples from C4 1.
- LLaMa-2-13B. Proving took 803 seconds and produced a 188 kB proof that verified in 3.95 seconds, using 23.1 GB of memory; the one-time weight commitment took 986 seconds 1.
- Accuracy. Perplexity changed little: for LLaMa-2-13B it moved from 6.520 to 6.528 1.
- Comparison. The authors compare zkLLM with an earlier system, zkML, on the same hardware 1. zkML ran out of memory beyond the size of GPT-2 (1.5 billion parameters), so its times for larger models are the authors' estimates 1.
- Artifact evaluation. The artifact received the CCS 2024 badges "Artifacts Available" and "Artifacts Evaluated--Functional" 2.
- Use by others. A verification system design for AI agreements uses zkLLM's figure of 803 seconds per 2,048-token forward pass on an A100 to judge ZKP overheads 3.
Limitations
Code maturity. The README states the code "has NOT undergone security auditing and is NOT ready for industrial applications" 2. It also names these gaps:
- prover and verifier run side by side;
- intermediate files are not meant as verifier inputs;
- an industrial deployment would need to separate the two parties and apply Fiat–Shamir 2.
Maintenance. The repository was archived in July 2025 2. The author states the project is no longer actively maintained 2.
Assumptions. The paper assumes a publicly known model structure 1. Its zero-knowledge guarantee is stated for a semi-honest verifier, which "accurately reports the outcome of the proof verification" but tries to learn the hidden parameters 1.
Scope. Proofs cover inference only 1. The authors write that extending zero-knowledge proofs to training LLMs "may pose insurmountable challenges" 1.
Attacks. As of September 2026 no attack on the soundness of zkLLM's proofs has been published. Its security rests on the paper's soundness and zero-knowledge theorems 1. An independent analysis, whose setting follows deployments such as zkLLM, shows that valid proofs do not bind the computation spent, so a much smaller model can pass as the declared one 4. Its authors demonstrated this with another system, zkGPT, on a small transformer 4. See Zero-knowledge proofs of inference.
Known flaws
Blockers
Proving takes about 13 minutes (803 seconds) of A100 time per 2,048-token forward pass at 13B parameters, plus a one-time weight commitment of 16 to 21 minutes.
The repository was archived on 10 July 2025 and the author states there is no plan for upgrades or maintenance.
No security audit of the code has been carried out.