Mechanisms · category

Cryptographic & computational

Protocols that check computation itself: recomputation, zero-knowledge proofs, proofs of learning, proofs of work, challenge-response.

NameTypeReadinessVerifiesThreat model
Attestable Audits
A research prototype that runs AI safety benchmarks inside a trusted execution environment and publishes attestations binding the model, the audit and the results.
ImplementationR2DemonstratedThe declared model is the one being servedSemi-trusted prover
Bounding unexplained information in outputs
Limits the hidden information a facility's outputs can carry by measuring how much of those outputs the declared computation fails to predict.
MechanismR2DemonstratedModel weights or data have not left the facilityAdversarial prover
Confidential multi-party verification
Lets mutually distrusting parties run an agreed check over private models or records inside attested enclaves or zero-knowledge proofs, revealing only the result.
MechanismR2DemonstratedThe declared model is the one being servedSemi-trusted prover
Deterministic and bit-exact inference
Making model inference reproducible bit for bit, so that a verifier's re-run must match the provider's output exactly rather than approximately.
MechanismR2DemonstratedThe declared model is the one being servedAdversarial prover
DiFR (Divergence From Reference)
DiFR checks that an inference provider ran its declared model by comparing output tokens or activations with a trusted re-run using the same random seed.
ImplementationR2DemonstratedThe declared model is the one being servedAdversarial prover
Model identity attestation
Establishes that responses come from a specific, committed set of model weights, using enclave measurements or recomputation of sampled outputs.
MechanismR2DemonstratedThe declared model is the one being servedSemi-trusted prover
Pearl proof-of-useful-work blockchain
A blockchain whose mining is designed to be a by-product of GPU matrix multiplications in AI workloads, with public node and miner code.
ImplementationR2DemonstratedAdversarial prover
Proof-of-learning and training-transcript verification
A trainer logs checkpoints, data order and settings, so a verifier can re-run sampled training segments and check that the claimed training happened.
MechanismR2DemonstratedA training run stayed within declared limitsAdversarial prover
Safeguard attestation
Hardware-signed evidence that an AI service ran its declared safeguards, such as a guardrail classifier or monitor, when producing a given response.
MechanismR2DemonstratedDeclared safeguards were applied during inferenceSemi-trusted prover
Sampled inference recomputation
A verifier re-runs a random sample of an AI provider's logged queries on a trusted copy of the declared model and checks the outputs match.
MechanismR2DemonstratedThe declared model is the one being servedAdversarial prover
TEE remote attestation for AI workloads
Trusted execution environments (TEEs) in CPUs and GPUs sign reports of loaded software, so a remote party can check which code ran an AI workload.
MechanismR2DemonstratedThe declared model is the one being servedSemi-trusted prover
Tinfoil model identity (Modelwrap)
Tinfoil's method for proving which model weights its enclave-hosted inference service runs, by binding a dm-verity hash of the weights into remote attestation.
ImplementationR2DemonstratedThe declared model is the one being servedSemi-trusted prover
TOPLOC
TOPLOC is a hashing scheme from Prime Intellect that lets a verifier check whether an inference provider ran the model, prompt and precision it claims.
ImplementationR2DemonstratedThe declared model is the one being servedAdversarial prover
Zero-knowledge proofs of inference
A prover produces a cryptographic proof that an output came from running a committed model on a given input, without revealing the weights.
MechanismR2DemonstratedThe declared model is the one being servedAdversarial prover
Zero-knowledge proofs of training constraints
Cryptographic proofs that a training run followed a committed dataset, procedure and rules, checkable without revealing the model or the data.
MechanismR2DemonstratedA training run stayed within declared limitsAdversarial prover
zkLLM
zkLLM is a GPU-accelerated zero-knowledge proof system that proves a large language model's output came from committed weights without revealing those weights.
ImplementationR2DemonstratedThe declared model is the one being servedAdversarial prover
AI 2040 inference-only verification stack
A proposed retrofit that isolates data-centre inference units, taps their front-end traffic and recomputes random samples to check that only declared inference runs.
ImplementationR1ProposedThis compute runs inference, not trainingAdversarial prover
Attestable zero-knowledge inference prover
Attestable's zero-knowledge prover, which the company reports proves large language model outputs came from committed weights at tens of tokens per second.
ImplementationR1ProposedThe declared model is the one being servedAdversarial prover
Chip registries and manufacturing records
Recording each AI chip's identity and owner from the fab onwards, and cryptographically fixing manufacturing records, so that chips can be accounted for later.
MechanismR1ProposedCompute stock is at most a declared amountSemi-trusted prover
Low-trust AI compute verification system overview
A retrofittable reference design in which network taps commit to all facility traffic, and air-gapped, independently sourced checkers later re-run randomly challenged records.
ImplementationR1ProposedThis compute runs inference, not trainingAdversarial prover
Memory wiping and proofs of secure erasure
Overwriting all of a device's memory in a way a verifier can check, so that nothing from earlier, undeclared work survives the wipe.
MechanismR1ProposedAdversarial prover
Network taps and certifiers
Devices on a cluster's network links that copy and hash all traffic, so a verifier can later check sampled records against declared work.
MechanismR1ProposedThis compute runs inference, not trainingAdversarial prover
Proofs of useful work and resource exhaustion
Cryptographic evidence that hardware performed a given amount of agreed computation, proposed as a way to show no spare capacity remained for other work.
MechanismR1ProposedThere is no undeclared relevant computeAdversarial prover
Reproducible computation packets
Organizing all AI workloads in a facility into discrete, reproducible units, so that a verifier can recompute a random sample and check each one.
MechanismR1ProposedThis compute runs inference, not trainingAdversarial prover
SASH confidential network logger
An open-source prototype that routes a facility's inference traffic through a logger and re-runs requests on a separate cluster to check it serves inference.
ImplementationR1ProposedThis compute runs inference, not trainingSemi-trusted prover
Timed challenge-response and memory-occupation challenges
A verifier sends unpredictable questions that a device can answer in time only if it holds specified data, or dedicates specified resources, locally.
MechanismR1ProposedDeclared hardware is idle or shut downAdversarial prover

Includes records that list this as a secondary category.