Readiness levels
Every mechanism and implementation has a readiness level from R0 to R4. The level says how close a technique is to being usable for its verification purpose, not how mature the underlying technology is.
Each record lists the verification use its level covers and confidence in that assessment; ⚠ marks an open critical flaw.
R0Idea
Mentioned or sketched, without a written design, claim and assumptions.
No record is at this level.
R1Proposed
A design or theory is publicly described, including the claim it would verify and its assumptions.
Roughly technology readiness levels 1–3.
Mechanisms (10)
- Chip location verification
for bounding how far a chip is from trusted landmark servers when checked
Medium confidence - Chip registries and manufacturing records
for a checkable record of which chips were made and who declared owning them
Medium confidence - Hardware performance throttling and licensing
for performance limits a verifier can rely on, against an operator trying to bypass them
Medium confidence - Hardware-enabled guarantees (flexHEG) and guarantee processors
for checking and enforcing training-compute limits on chips, against adversaries up to states
Medium confidence - Memory wiping and proofs of secure erasure
for showing that no data from earlier work persists in memory the wipe reaches
Medium confidence - Network taps and certifiers
for committing a complete record of cluster traffic, so declared inference can be checked
Medium confidence - Proofs of useful work for capacity accounting
for bounding the spare capacity of declared hardware that could run training
Low confidence - Remote detection of data centres
for finding undeclared data centres above an agreed compute threshold
Medium confidence - Side-channel suppression for isolated facilities
for bounding physical covert channels out of a verified enclosure
Medium confidence - Whole-workload recomputation (reproducible packets)
for recomputing whole workloads to show a cluster runs only declared inference
Medium confidence
Implementations (7)
- AI 2040 inference-only verification stack
for showing that retrofitted data centres run only inference
Medium confidence - Attestable zero-knowledge inference prover
for proving an output came from committed weights
Low confidence - Data-centre memory challenging
for confirming data presence and bounding free memory across data-centre servers
Medium confidence - Low-trust AI compute verification system overview
for screening challenged records to show declared inference compute is not training
Medium confidence - Lucid sovereignty (location) certificates
for certifying the region where an attested workload ran at a given time
Medium confidence - RAND secure inference data center (SIDC) design
for the operator's own weight security, with no outside verification described
Low confidence - SASH confidential network logger
for telling inference from training on a mutually inspected cluster
Medium confidence
R2Demonstrated
A public working implementation, or reproducible published end-to-end results, obtained under conditions representative of the verification use in at least one key respect: realistic model or cluster scale, realistic hardware, or a stated adversary. Lab and research demonstrations reach this level; production use is R3.
Roughly technology readiness levels 4–6.
Mechanisms (11)
- Bandwidth limits and compartmentalization
for monitoring inter-node traffic with operator-run software on four GPUs
Low confidence - Bounding unexplained information in outputs
for bounding how much hidden information can leave in checked inference outputs
Low confidence - Confidential multi-party verification
for audits or evaluations of a private model that reveal neither party's inputs
Medium confidence - On-chip telemetry from timing, memory and performance counters⚠
for workload evidence from GPU counters and timing, assuming authentic measurements
Medium confidence - Safeguard attestation
for attesting that a declared safeguard mediated a service's responses
Low confidence - Tamper evidence for verifier devices
for detecting probing of proposed verifier hardware, using server and electronics prototypes as evidence
Medium confidence - Timed challenge-response and memory-occupation challenges
for detecting whether a GPU is doing other work
Medium confidence - Training-transcript verification (proof-of-learning)⚠
for checking from its transcript that a training run followed declared rules
Medium confidence - Workload classification from telemetry and side channels
for telling training from inference and other work using genuine telemetry, including disguised workloads
Medium confidence - Zero-knowledge proofs of inference
for proving a language model's output follows from committed weights, against a cheating prover
Medium confidence - Zero-knowledge proofs of training constraints
for proving a training run followed a committed specification and data
Medium confidence
Implementations (13)
- Attestable Audits
for showing users that the model answering them is the audited one
Low confidence - Batch-invariant inference kernels (Thinking Machines)
for exact recomputation of served outputs by a verifier, with a cooperating provider
Low confidence - Cove
for composing owner-approved confidential workflow stages on Intel TDX
Medium confidence - DiFR (Divergence From Reference)
for checking that outputs match the declared model, precision and sampling settings
Medium confidence - EZKL
for proving a language model's output follows from committed weights, against a cheating prover
Medium confidence - GPU contention probes
for detecting another workload running on the same GPU
Medium confidence - Kaizen
for proving gradient-descent training of small image models on committed data
Medium confidence - PAL*M
for attesting declared model operations on a confidential CPU–GPU prototype
Medium confidence - PySyft double-blind evaluations
for evaluating a private model on private prompts, neither party seeing the other's inputs
Medium confidence - SAGE
for attesting code execution on a GPU that lacks hardware trusted-execution support
Medium confidence - VeriLoRA
for proving single-sample LoRA fine-tuning steps for 3–13-billion-parameter language models
Medium confidence - VRAM-residency challenge
for detecting whether verifier-supplied data remains in a single GPU's memory
Medium confidence - zkLLM
for proving an output came from committed weights, against a prover who cheats
Medium confidence
R3In production
R2 holds, and it is production-grade and available, or a party other than its developer relies on it for a verification decision. Critical flaws may still be open: production use does not show that it holds up against a prover who tries to cheat.
Roughly technology readiness levels 7–8.
Mechanisms (4)
- Deterministic and bit-exact inference
for reproducing open-model inference from receipts in Gensyn's information-market service
Low confidence - Model identity attestation⚠
for showing users that a service runs the declared model weights
Medium confidence - Sampled inference recomputation
for checking untrusted workers' activations against the declared model, prompt and precision
Low confidence - TEE remote attestation for AI workloads⚠
for showing which software ran to a party that distrusts the operator holding the hardware
Medium confidence
Implementations (5)
- Apple Private Cloud Compute
for showing users which software serves their AI requests, not which model
Medium confidence - Pearl proof-of-useful-work blockchain
for checking matrix-multiplication work proofs for blockchain consensus
Low confidence - Tinfoil model identity (Modelwrap)⚠
for showing clients that the served weights match a committed hash
Medium confidence - TOPLOC
for checking that untrusted providers used the claimed model, prompt and precision
Low confidence - Verde and RepOps (Gensyn)
for reproducing declared-model inference from receipts in Gensyn's information-market service
Low confidence
R4Deployment-ready
R3 holds, and at least one independent public evaluation (an audit, a red-team or a peer-reviewed security analysis) left no critical flaw open.
Roughly technology readiness level 9.
No record is at this level.
How levels are assigned
- A record gets the highest level whose criteria all hold.
- The level is assessed for the verification use stated beside it. A commercial component does not put a verification workflow built on it in production. TEE attestation is in production (R3) because services rely on it to show users which software runs, but that does not make TEE-based verification between rival states deployment-ready (R4).
- "Reproducible" means the method, setup and parameters are published in enough detail for an independent team to repeat the work. Public code is not required, but its absence is noted.
- An open critical flaw does not by itself lower R1, R2 or R3, which describe development, demonstration and use. A break that invalidates the evidence a level rests on does lower it.
- "Production-grade" (R3) covers a system its developer runs for real users, even at a 0.x version. A release its developer labels alpha, beta or preview does not count unless someone else relies on it for a verification decision.
- Production use that has stopped still counts while the system remains available. The rationale dates the last documented use, and confidence is low if none is documented in the last 12 months.
- A mechanism is at least as ready as its most mature implementation for that use, and its rationale names that implementation.
- Ratings can go down.
- Adversarial evaluation is also judged for the verification use. Attacks on related products in other settings are context, not evidence.
What each rating records
Each rating states the verification use it is assessed for, and comes with a rationale that walks through the criteria, the sources it rests on, the gaps to the next level, a confidence, who assessed it and when, and a status (current, under review or disputed). Levels are editorial judgments under rubric v1.1, kept apart from the facts, and open to correction.
Confidence is how sure the editors are of the level: 13 low confidence, 37 medium confidence and 0 high confidence. The Status page lists the low-confidence ratings. The Methodology page covers flaws, citations and neutrality.