A fully reproducible inference stack needs substantial software and tooling, and per-packet network reproducibility may need considerable software, firmware and possibly hardware work.
Passive optical taps work at 400G, but the 800G and 1600G line rates now arriving in data centres are undemonstrated.
Checking that taps are correctly installed and stay in place at scale is not a solved problem, and hardening the recomputation server inside the prover's facility needs significant research.
There is no plan yet for quickly scaling side-channel defences on a frontier cluster; only early theoretical pieces exist.
Memory wiping may use existing algorithms, but hardware testing is at an early stage.
Robust red-teaming of recomputation schemes has not started, and most algorithm development remains academic.
A public working implementation or reproducible end-to-end results for the integrated stack (isolated inference unit, taps, packetization and recomputation) at realistic scale or against a stated adversary.
Next readiness level
A hardened recomputation server and a method for checking that taps are correctly installed and remain in place.
Next readiness level
Red-teaming of recomputation and of the completeness measures (side channels, memory wiping).
Next readiness level
Assessment and limits · Assessment sources
Related organizations: Amodo Design, AI Futures Project
Published binaries cannot be rebuilt from source and carry no symbols, so checking what the attested software does needs reverse engineering.
Outside developers can use PCC only with an entitlement from Apple, which is limited to small developers.
An independent public evaluation of the attestation and transparency-log chain, including what attestation covers at runtime, that leaves no critical flaw open.
Next readiness level
Evidence that the Google Cloud deployment's attestation resists attackers with physical access to TDX and NVIDIA hardware.
Next readiness level
Reproducible builds, so that published binaries can be checked against published source.
Next readiness level
Assessment and limits · Assessment sources
The prototype needs porting to GPU confidential computing to handle larger models; the authors expect an overhead as small as 5 times there.
As of September 2026 no code has been released for the prototype.
Reliance by a party other than the developers for a verification decision, or a production-grade, available implementation, for example on GPU confidential computing at realistic model scale.
Next readiness level
An independent public security evaluation (audit, red-team or peer-reviewed analysis).
Next readiness level
Assessment and limits · Assessment sources
Related organizations: University of Cambridge
No paper, protocol specification or code is public, so the reported results cannot be reproduced.
Attestable reports a context window limited to 16K tokens.
Covering computation that is not proven relies on proof-of-work accounting, which Attestable has only proposed.
A public working implementation, or reproducible published end-to-end results, such as a paper with a protocol specification and benchmarks others can rerun.
Next readiness level
Any independent security analysis of the proof system.
Next readiness level
Assessment and limits · Assessment sources
Related organizations: Attestable
No cap that a verifier can check has been implemented or red-teamed.
The verifier must know that all traffic leaving a pod crosses the capped, monitored links.
Shaping devices and routing assignments must be trusted by both parties; Amodo has not yet fully analysed resilience to a compromised DPU.
Advances in low-communication training could shrink the margin that the cap enforces.
A production-grade cap or bandwidth monitor for this use, or reliance on one by a party other than its developer for a verification decision.
Next readiness level
Enforcement outside the prover's control, such as a shaper at each pod uplink, built and tested on data-centre hardware (Lucid reports a proof of concept in development).
Next readiness level
Measured, not only modelled, training inefficiency under a cap at realistic scale, including low-communication methods.
Next readiness level
Red-teaming of cap bypass (for example through parallel scale-up switches) and of trust in the shaping devices.
Next readiness level
Assessment and limits · Assessment sources
Related organizations: Amodo Design, Lucid Computing, AI Futures Project, Machine Intelligence Research Institute
Batch invariance costs throughput: on Qwen3-8B the improved deterministic build took 42 s against 26 s for vLLM's default, and SGLang reports an average slowdown of 34.35% on its FlashInfer and FlashAttention 3 backends.
Outputs are identical only while the model, inference implementation and device stay fixed, so provider and verifier must run the same stack.
A production-grade verification stack that uses these kernels for exact-match checks, or a party other than the developers relying on such checks for a verification decision.
Next readiness level
Assessment and limits · Assessment sources
Related organizations: Thinking Machines Lab
The prover's compute must be isolated so that all traffic passes through the verifier's interlock; any unmonitored path voids the bound.
Physical side channels need separate suppression, and one design treats a low residual bandwidth, rather than zero, as the realistic target.
Tolerance for numerical nondeterminism sets the size of the residual channel; bit-exact replay would remove it but needs full hardware and software metadata.
Recomputation over confidential weights and inputs needs a protected setting: prover recomputation in a verifier-controlled enclosure, verifier recomputation in a prover-controlled enclosure, or zero-knowledge proofs.
No prototype of the facility-level architecture exists to red-team.
A prototype of the facility-level architecture (interlock, commitments, challenge-based prediction) with published results.
Next readiness level
A bound that holds against an adversary who controls the prompt distribution, for example with entropy-calibrated tolerances, evaluated independently.
Next readiness level
Reliance by a party other than the developer on an unexplained-information bound for a verification decision, or a production-grade deployment.
Next readiness level
Assessment and limits · Assessment sources
No public code or reproducible end-to-end location results are available for the reported H100 prototype.
Per-chip keys must be provisioned and protected against extraction; hardware-integrated, tamper-resistant versions still need R&D.
The time limit forces a trade-off: a limit at the speed of light in fibre can be beaten by faster links, while one at the vacuum speed of light makes honest chips fail often.
A trusted landmark network must be built and secured, and who should operate it, under what oversight, is unsettled.
A public implementation, or reproducible end-to-end results, on data-centre accelerators with a real landmark network.
Next readiness level
Published measurements of false-positive and false-negative rates under realistic internet routing.
Next readiness level
An evaluation against a stated adversary covering delay manipulation, faster network paths, landmark compromise and key extraction.
Next readiness level
Assessment and limits · Assessment sources
Related organizations: Lucid Computing, Institute for AI Policy and Strategy, Center for a New American Security, NVIDIA
No AI chip registry operates, and covering re-exports would need cooperation from re-exporters and foreign governments that may not be feasible everywhere.
Linking records to physical chips needs hard-to-spoof unique IDs and inspections.
Chips produced before a registry starts must be reconstructed from supplier records.
A public pilot registry or published commitment to manufacturing records at realistic scale.
Next readiness level
End-to-end results for sampled chain-of-custody checks, including how hard-to-spoof IDs are read and matched.
Next readiness level
An adversarial evaluation of record falsification, forged serial numbers and unrecorded chips.
Next readiness level
Assessment and limits · Assessment sources
Related organizations: RAND, AI Futures Project, Institute for AI Policy and Strategy
Frontier inference typically needs the resources of several GPUs, and published confidential evaluations ran on a single server, so the double-blind pilot's authors name many-node confidential H100 or B200 clusters as the next milestone.
Zero-knowledge audits have been shown on image classifiers and a recommender model, not language models at frontier scale, and a counterfactual audit of the recommender cost $8,456.
Trust rests on a small number of hardware vendors, and a per-CPU Intel attestation key has been extracted by physical attack.
Parties must negotiate the plan or workflow and handle false positives and appeals, which the Auditor-in-a-Box authors list as open problems.
Released evidence must be designed to limit collateral leakage, which requires declaring protected properties in advance and calibrating on labelled executions.
A production-grade, available confidential evaluation or audit workflow, or reliance by a party other than its developer on such a result for a verification decision.
Next readiness level
An independent public evaluation (audit, red-team or peer-reviewed security analysis) of a confidential multi-party verification system, such as Confidential Space with PySyft or Cove, that leaves no critical flaw open.
Next readiness level
Confidential evaluation on multi-node enclave clusters at the scale of the largest frontier models, or zero-knowledge audits at that scale.
Next readiness level
Assessment and limits · Assessment sources
Related organizations: University of Cambridge, Machine Intelligence Research Institute, Tinfoil, OpenMined
Docker policy cannot prove that arbitrary guest workloads cannot generate quotes when quote channels are globally exposed.
The production workflow depends on owners reviewing manifests, allow rules and provisioning code.
A production-grade, available confidential evaluation workflow, or another party's documented reliance on its results.
Next readiness level
Assessment and limits · Assessment sources
No network-level probe has been shown to distinguish one server's memory contents from another's.
Most of the design needs a probe in close proximity to the challenged memory.
Filling a pod's volatile memory takes tens of minutes, and on-board SSDs take hours.
Remote memory access must be excluded during challenges.
A network-level challenge that distinguishes memory contents between servers, with public code or measurements described in enough detail to repeat.
Next readiness level
Measured reaction times from different points in the network hierarchy.
Next readiness level
A threat model and red-teaming of evasion under repeated challenges.
Next readiness level
Assessment and limits · Assessment sources
Related organizations: Machine Intelligence Research Institute
Batch-invariant kernels cost throughput: in Thinking Machines' Qwen3-8B test, an improved deterministic build took 42 s against 26 s for vLLM's default, and SGLang reports an average 34.35% slowdown on its FlashInfer and FlashAttention 3 backends.
Coverage is incomplete: the bit-exact emulator targets dense blocks on NVIDIA GPUs and excludes mixture-of-experts inference and training, and vLLM's batch-invariant mode is in beta, with open work on AMD hardware and speculative decoding.
Amodo's status page for the AI 2040 verification plan rates a reproducible inference stack for that plan as 'not started'.
Exact replay requires the prover to disclose weights, software versions, parallelism and batch sizes to whoever recomputes.
An independent public evaluation (audit, red-team or peer-reviewed security analysis) of bit-exact verification that leaves no critical flaw open.
Next readiness level
Assessment and limits · Assessment sources
The verifier needs the model weights, so outsiders cannot use the method to verify providers of closed-weights models.
The verifier must know and match the provider's sampling procedure, and in one prototype a sampling mismatch in a newer vLLM version produced large spurious logit differences.
No independent red-team of DiFR's consistency check has been published, Amodo rates recomputation red-teaming 'not started', and the one independent attack study targets an exfiltration detector built on the same statistic.
Reliance by a party other than the developers on DiFR for a verification decision, or a production-grade release.
Next readiness level
An independent public security evaluation (audit, red-team or peer-reviewed analysis) against adaptive adversaries.
Next readiness level
Assessment and limits · Assessment sources
Related organizations: Amodo Design
Proving cost grows steeply with model size: a 250,000-parameter nanoGPT took 2,781 s to prove and needed a 219 GB proving key, which South et al. name as the main limit on model size.
A production-grade release that proves language models at the scale verification claims concern, or reliance by another party on such proofs for a verification decision.
Next readiness level
Assessment and limits · Assessment sources
Related organizations: Zkonduit
Continuous probes add power draw, occupy GPU memory and reduce inference throughput.
No thresholds or statistical tests define when a timing shift counts as a detection.
Production-grade probe tooling, or use by a party other than the authors for a verification decision.
Next readiness level
Detection thresholds with measured false-positive and false-negative rates.
Next readiness level
Results on multi-GPU servers and against an operator who tries to hide a workload.
Next readiness level
Assessment and limits · Assessment sources
The throttles need new microarchitecture in future chips, and chipmakers would have to adopt it.
Secure licensing and trigger infrastructure is missing, such as a guarantee processor that issues or checks licenses.
Licenses denominated in work need secure meters for the licensed quantities.
No throttle has been evaluated on real hardware or against red-team attempts at bypass.
A public hardware or FPGA prototype, or end-to-end results described in enough detail to repeat (released simulator changes and scripts would also serve).
Next readiness level
End-to-end evaluation on training and inference workloads on current architectures.
Next readiness level
A trigger or licensing path that is cryptographically validated and shown to resist interception.
Next readiness level
Attestation by which a remote verifier can confirm the throttle state.
Next readiness level
Assessment and limits · Assessment sources
Related organizations: RAND, Center for a New American Security
Integrated flexHEG needs substantial help from the accelerator manufacturer, and the authors estimate 3.7–7.9 years, from when the manufacturer starts work, for such hardware to displace other accelerators in frontier development.
State-level attackers who hold the hardware can likely compromise the best current secure enclosures.
Rival states would need to trust the design and manufacture of guarantee processors and enclosures, for example through open design, redundant processors from each side or oversight of production.
Restricting future rule updates would need a formal language for rules, which the authors judge most likely infeasible for early flexHEG versions.
Governing all relevant chips depends on knowing where they are, through chip registries and detection of undeclared facilities.
A public prototype of a guarantee processor or Interlock on a real accelerator data path, with published end-to-end results.
Next readiness level
A secure enclosure evaluated against invasive physical attacks, with published cost-to-circumvent estimates.
Next readiness level
A specified ruleset language and a multi-party update protocol implemented and analysed.
Next readiness level
Chipmaker engagement, needed for integrated designs.
Next readiness level
Assessment and limits · Assessment sources
Related organizations: RAND, Center for a New American Security
Proving a VGG-11 gradient-descent iteration with 10 million parameters and batch size 16 took about 15 minutes.
A production-grade, available proof-of-training implementation, or another party's documented reliance on its proofs.
Next readiness level
Assessment and limits · Assessment sources
Empirical feasibility of passive optical splitting at 53–112 GBaud under realistic conditions is an open question.
Exact replay needs complete hardware and software metadata, and the tolerable slowdown from emulation is an open question.
Tamper-evident, rapidly mass-manufacturable and retrofittable enclosures for side-channel defence are an open research question, and physical security against covert communication in every monitored data centre is challenging.
A mass-manufacturable, good-enough side-channel defence, particularly power-line filtering, has not been constructed or red-teamed.
Distinguishing one server's DRAM contents from another's by challenge-response timing, and a general challenge-response protocol for diverse data types, are open.
The threat model is under-developed and needs input from cybersecurity and AI threat-modelling experts.
A public working implementation or reproducible end-to-end results for the capture-then-challenge pipeline, at realistic line rates or against a stated adversary.
Next readiness level
Demonstrated bit-exact replay of production inference inside a secure auditing environment built from independently sourced components.
Next readiness level
A developed threat model and red-teaming of the side-channel, egress and inspector-agent components.
Next readiness level
Assessment and limits · Assessment sources
Related organizations: Machine Intelligence Research Institute
The specification is an unfinished draft with no public implementation or evaluation.
It needs a globally distributed, trusted anchor fleet and an endorser to run the anchor directory.
A public reference implementation, or reproducible results with a real anchor fleet and realistic network paths.
Next readiness level
Published location accuracy and false-rejection rates.
Next readiness level
An independent security review of the protocol and its trust assumptions.
Next readiness level
Assessment and limits · Assessment sources
Related organizations: Lucid Computing
Wipes take time: tens of minutes for a pod's volatile memory and hours for SSDs, displacing work.
Timed challenges must exclude remote memory and other helpers.
All memory stores in a system must be inventoried and wiped at the same time.
A public end-to-end wipe of an accelerator server, covering fill and timed challenges by a separate verifier, for host memory, HBM and storage.
Next readiness level
Coverage of memory the wipe cannot reach: drive-controller DRAM, firmware stores, NICs, DPUs and switches, and the algorithm's own working memory.
Next readiness level
Adversarial evaluation in a data-centre setting, including remote-memory (RDMA) help during challenges.
Next readiness level
Assessment and limits · Assessment sources
Related organizations: Amodo Design, AI Futures Project, Machine Intelligence Research Institute
Attestation that resists physical attackers, for the enclave variant.
Numerical nondeterminism limits how tightly recomputation can pin down the model and sampling.
The recomputation variant needs the verifier to hold the declared weights.
Independent security evaluation of a deployed model-identity scheme that leaves no critical flaw open.
Next readiness level
Resistance of the enclave variant to physical attackers (see TEE remote attestation for AI workloads).
Next readiness level
Third-party verification for private models beyond consistency across requests.
Next readiness level
Tooling for audit-time checking of transparency records.
Next readiness level
Assessment and limits · Assessment sources
Related organizations: Tinfoil, University of Cambridge, Machine Intelligence Research Institute
No complete verification tap has been demonstrated at production frontend link rates, and on the tested CPU no hash algorithm reached line rate with minimum-size frames.
Nondeterministic inference leaves covert capacity in outputs that hashing cannot remove.
Taps and gateway devices need tamper-evident housing and physical monitoring so that traffic cannot bypass them.
Radio, power-line and thermal channels are not addressed by network-level designs.
Red-teaming by specialists is called for but has not been reported.
Hashing and tapping demonstrated at production frontend link rates (400 Gbps class), including minimum-size frames.
Next readiness level
A built active warden or Secure Gateway Device, with measured residual covert bandwidth.
Next readiness level
Red-teaming of tap bypass, covert channels and physical security.
Next readiness level
Assessment and limits · Assessment sources
Related organizations: Amodo Design, Singapore AI Safety Hub (SASH), AI Futures Project, Machine Intelligence Research Institute, Hardware AI Governance Lab
Shipping accelerators need a tamper-resistant, authenticated telemetry path.
NVIDIA's full confidential-computing mode disables the hardware performance counters its profiling tools use, so telemetry that needs them conflicts with it.
Continuous challenge puzzles cost power and throughput on production workloads.
Evaluation has not gone beyond single nodes, framework-level evasion and one vendor's hardware.
Use by a party other than the developers for a verification decision.
Next readiness level
Telemetry read paths that the operator cannot forge, such as signed counters from a root of trust or a guarantee processor.
Next readiness level
Calibrated false-positive and false-negative rates, with a detection-theoretic threshold framework.
Next readiness level
Independent red-teaming, including custom-kernel and multi-node evasion.
Next readiness level
Assessment and limits · Assessment sources
Related organizations: Machine Intelligence Research Institute
The paper plans prototype code release after peer review.
Overhead in memory-mapped dataset preprocessing and attribute-distribution experiments accounted for 37.64–42.50% of total instrumented runtime.
A production-grade, available property-attestation implementation, or documented reliance by another party on its attestations.
Next readiness level
Assessment and limits · Assessment sources
Related organizations: University of Waterloo
Built for consensus rather than capacity bounding; verifying that declared hardware has no spare capacity would also need a credible compute estimate.
Performance figures are provider-reported, and the benchmark reports no baseline of the certified model without mining.
Bit-exact verification depends on reproducing GPU arithmetic deterministically.
An independent public security evaluation of the protocol or its implementation.
Next readiness level
For verification use: a demonstration that proof rates can bound the spare capacity of declared hardware.
Next readiness level
Assessment and limits · Assessment sources
Related organizations: Pearl Research Labs
Bounding spare capacity needs a credible estimate of the compute available to the actor, including third-party access 6.
Proofs of work cannot find facilities that were never declared 6.
As of September 2026 no implementation, demonstration or independent evaluation of proofs of work for capacity bounding has been published.
A public implementation or reproducible end-to-end result that uses proofs of work to bound the spare capacity of declared hardware against a stated adversary.
Next readiness level
A method for the verifier to obtain a credible estimate of the prover's available compute.
Next readiness level
Assessment and limits · Assessment sources
Related organizations: Attestable, Pearl Research Labs, Machine Intelligence Research Institute
The pilot could not inspect or allowlist all model code, and the guest operating system builds were not independently reproducible.
The pilot ran on one H100; the authors name many-node confidential GPU clusters as the next scale target.
A generally available production workflow, or documented reliance by another party on its result for a verification decision.
Next readiness level
An independent public security evaluation that leaves no critical flaw open.
Next readiness level
Assessment and limits · Assessment sources
Related organizations: OpenMined
No prototype exists; RAND recommends prototyping key security features and integration now.
The report describes internal integrity checks, audit logging and accreditation, but no way for a party outside the operator to verify the facility's properties.
Human review of every prompt and response makes each request take three to five minutes, with the review steps as the rate-limiting factor.
Detailed design information is withheld from the public report and is to be evaluated privately with stakeholders, which limits independent public scrutiny.
A public working prototype, or reproducible published results, for key features such as the diode-gated realm topology and cross-realm protocols.
Next readiness level
A published way for a party other than the operator to verify the facility's claims, for example weight confidentiality or which model is served.
Next readiness level
Assessment and limits · Assessment sources
Related organizations: RAND
Wide-area, automated detection of data centres is not yet practical and needs large training datasets.
No measured detection or false-alarm rates for finding undeclared facilities have been published.
Recent high-resolution imagery is costly, is limited by weather and needs trained analysts.
Published end-to-end results on finding previously unknown large facilities over a wide area, with measured miss and false-alarm rates.
Next readiness level
An evaluation against a stated concealment adversary, for example disguised or underground facilities.
Next readiness level
Assessment and limits · Assessment sources
Related organizations: Planet Labs, AI Futures Project, Epoch AI
No published design shows that all of a provider's traffic passes through the attested safeguard path; current evidence covers individual attested responses.
Frontier model inference typically needs several GPUs, GPU confidential computing is less mature than CPU support, and CPU inference, which an enclave prototype had to use, ran about 100 times slower than GPU inference.
Trust rests on a small number of hardware vendors, and a per-CPU Intel attestation key has been extracted by physical attack.
Safeguard evidence must be bound to the model actually served, which depends on model-identity attestation.
No independent red-team or audit of a safeguard-attestation system has been published, and the available prototypes are described by their authors as proofs of concept that have not been stress-tested by a counterparty.
Reliance by a party other than the developer on safeguard attestations for a verification decision, or a production-grade system that is generally available.
Next readiness level
An independent public evaluation (audit, red-team or peer-reviewed security analysis) of a safeguard-attestation system.
Next readiness level
A demonstration in which the safeguard model itself runs inside the attested boundary on GPU hardware at a realistic serving scale.
Next readiness level
A published way to show that all of a service's traffic, not only attested responses, passed through the attested safeguard path.
Next readiness level
Assessment and limits · Assessment sources
Related organizations: University of Cambridge, Machine Intelligence Research Institute, Tinfoil
The verifier must know the exact hardware configuration of the GPU.
The verifier runs in an SGX enclave on the same host as the GPU, so the scheme inherits trust in that enclave.
A production-grade release, or use by a party other than the authors for a verification decision.
Next readiness level
Evaluation on current AI accelerators and with AI inference or training workloads.
Next readiness level
An independent security evaluation of the timing margin against proxy and optimisation attacks.
Next readiness level
Assessment and limits · Assessment sources
In tap-based retrofit designs, recording all inference traffic needs network taps and recomputation servers that can ingest it, in the worst case one recomputation-server network interface per inference front-end interface.
In retrofit designs, the recomputation server must sit inside the prover's data centre, possibly under the prover's physical control, and still be protected from a compromised provider, which Amodo rates 'not on track'.
No independent red-team of a recomputation consistency check has been published (the one independent attack study targets the weight-exfiltration bound), and Amodo rates recomputation red-teaming 'not started'.
Tolerance-based checks need calibration on trusted hardware and exact knowledge of the provider's sampling procedure, and in one prototype a sampling-implementation mismatch produced large spurious differences.
The verifier needs the model weights, so checking a closed-weights model requires a trusted, confidential recomputation environment, which the retrofit designs place inside the prover's facility.
An independent audit, red-team or peer-reviewed security analysis of a recomputation consistency check against an adaptive adversary.
Next readiness level
Assessment and limits · Assessment sources
Related organizations: Prime Intellect, Amodo Design, AI Futures Project, Machine Intelligence Research Institute
The planned FPGA logger has not yet been built.
The design has not been scaled to production traffic volumes.
Exact-match recomputation requires reproducible inference.
A hardware logger (the planned FPGA version) on a real data-centre link, with realistic model size or traffic volume.
Next readiness level
A stated adversary and threat model, with testing against it.
Next readiness level
Random sampling and hashed or certified traffic records, as described in the blog post, implemented in the public code.
Next readiness level
Assessment and limits · Assessment sources
Related organizations: Singapore AI Safety Hub (SASH), Future of Life Institute
No prototype or red-team exists; the design is a first-pass viability study.
Volume costs of TEMPEST-grade power-line filters are uncertain, because existing products are mostly made to order.
A prototype enclosure for at least one AI rack or scalable unit, with measured attenuation for each channel class.
Next readiness level
A red-team exercise against the prototype by a stated adversary.
Next readiness level
Validated costs for filters, jamming and optical conversion at production scale.
Next readiness level
Assessment and limits · Assessment sources
Related organizations: Machine Intelligence Research Institute
No tamper-evident enclosure has been designed for AI verifier hardware at retrofit scale.
Battery-backed designs add bulk, limit operating temperature (+10 °C to +35 °C for the IBM 4765) and complicate transport.
Active monitoring needs power, and visual inspection of large enclosures faces access limits.
No evaluation has been published in the AI verification setting.
An enclosure or sensing design built for AI verifier devices (taps, gateways, recomputation servers) and deployable at data-centre scale.
Next readiness level
An independent public evaluation (red team or certification) of such an enclosure in the AI verification setting.
Next readiness level
Inspection protocols suited to host-controlled AI facilities.
Next readiness level
Assessment and limits · Assessment sources
Vendor threat models exclude sophisticated physical attacks, but in international verification the prover holds the hardware.
Negative claims such as "no undeclared training" need chip-wide accounting of all workloads, which attestation does not provide.
Multi-GPU and multi-node coverage is incomplete, because Hopper leaves NVLink traffic unencrypted and NVIDIA's April 2026 release notes list no multi-node confidential mode.
Rival parties have not agreed on trust roots and key provenance they would accept.
CPU-only enclaves are costly for large models, because in the Attestable Audits prototype CPU inference cost 21.7 times as much per token as GPU inference and the enclave roughly doubled the CPU cost.
Attestation that survives an attacker who physically holds the hardware, for example through memory integrity and freshness protection or tamper-responsive enclosures, confirmed by independent red-teaming.
Next readiness level
GPU attestation cryptographically bound to the specific confidential VM it serves.
Next readiness level
Coverage of whole-chip and multi-node activity, beyond a single deployment.
Next readiness level
Roots of trust and key provenance that rival parties accept, beyond one vendor's certificate authority.
Next readiness level
Measurement of runtime configuration as well as launch state.
Next readiness level
Assessment and limits · Assessment sources
Related organizations: NVIDIA, Tinfoil, University of Cambridge, Machine Intelligence Research Institute, Future of Life Institute
No network-level memory challenge across data-centre servers has been demonstrated.
Challenges that fill memory displace workloads; filling a pod's volatile memory takes tens of minutes and SSDs take hours.
Outside help, such as remote memory, must be excluded during challenges.
Production-grade challenge tooling, or use by a party other than the developers for a verification decision.
Next readiness level
A network-level challenge that bounds free memory across accelerator servers, with public code or measurements described in enough detail to repeat.
Next readiness level
Quantified false-positive and false-negative rates under adversarial conditions.
Next readiness level
Evaluation against known attack classes on timed attestation, such as compression and relocation.
Next readiness level
Assessment and limits · Assessment sources
Related organizations: Amodo Design, Machine Intelligence Research Institute
The underlying TEE attestation does not resist attackers with physical access to the host.
No independent evaluation of the model-identity chain has been published.
An independent security evaluation of Modelwrap and the attestation chain that leaves no critical flaw open.
Next readiness level
A supported tool for audit-time verification from transparency records.
Next readiness level
A way for third parties to learn something about private models beyond consistency across requests.
Next readiness level
Assessment and limits · Assessment sources
Related organizations: Tinfoil
No independent security evaluation has been published, and Amodo Design rates red-teaming of recomputation schemes as 'not started'.
The verifier must run the model itself, which suits the paper's setting of providers serving open-weights models.
An independent public security evaluation, such as an audit, red-team or peer-reviewed analysis, that tests adaptive attacks like the spoofing and speculative-decoding cases the TOPLOC authors list.
Next readiness level
Assessment and limits · Assessment sources
Related organizations: Prime Intellect
The verifier must receive the training data, weights and code 1 6.
Transcripts are large: weight checkpoints may each require terabytes 5.
The verifier must reproduce training segments, which may be infeasible if the prover uses specialised or proprietary hardware 6.
The noise tolerance needed for honest reproduction is what structurally correct spoofs exploit 4.
Tying transcripts to real chips needs on-chip weight-snapshot logging, chip inspections and a trusted chip-owner directory 5.
Use by a party other than the developer, or a production-grade implementation, at realistic training scale.
Next readiness level
Verification rules with formal robustness arguments, as Fang et al. argue are needed, or an independent red-team of the post-2023 defences.
Next readiness level
Assessment and limits · Assessment sources
Reproducibility costs throughput: RepOps added 98% to Llama-8B inference time on an A100 in the paper, and Gensyn reports a threefold cut in REE's reproducible-mode overhead without absolute figures.
The providers who re-run a job and the referee need the model and data, and the guarantee holds only if at least one provider is honest.
An independent public security evaluation of the Verde dispute protocol and of RepOps reproducibility across hardware, which the paper asserts but does not test.
Next readiness level
Assessment and limits · Assessment sources
Related organizations: Gensyn
Each single-sample step took 121.93–249.38 seconds to prove, with commitment generation taking another 156–554 seconds in the reported experiments.
A production-grade, available LoRA proving implementation, or another party's documented reliance on its proofs.
Next readiness level
Assessment and limits · Assessment sources
The challenge data occupies a large part of GPU memory for as long as the test runs.
No test across servers has been reported, and the MIRI overview lists network-level probing of memory contents as undemonstrated.
Production-grade tooling, or use by a party other than the authors for a verification decision.
Next readiness level
Detection thresholds with measured false-positive and false-negative rates.
Next readiness level
A test across servers, where the verifier is not on the same host as the GPU.
Next readiness level
Assessment and limits · Assessment sources
Workloads are not reproducible by default, and achieving reproducibility may cost performance.
Network packets are not individually reproducible by default; making them so may need considerable software, firmware and hardware work. Amodo rates this 'not on track'.
All traffic must reach the recomputation server via network taps, and the server's integrity is critical.
Recomputing training steps needs checkpoints: writing one at every step would cost more than 100% overhead, so Amodo's design needs a spare data-parallel replica that tracks the weights instead.
A public implementation, or reproducible end-to-end results, of packet-based recomputation beyond single inference requests, under realistic model scale, hardware or a stated adversary.
Next readiness level
Assessment and limits · Assessment sources
Related organizations: Amodo Design, AI Futures Project
Software telemetry is trustworthy only if on-chip counters are read over a path the operator cannot tamper with.
No independent red-team or third-party reliance has been reported.
Results do not yet cover multi-node clusters, other vendors or multi-tenant serving.
Reliance by a verifier other than the developers, or a production-grade, available system.
Next readiness level
An independent red-team or peer-reviewed security analysis.
Next readiness level
Results at multi-node cluster scale and across hardware vendors.
Next readiness level
A tamper-resistant, authenticated telemetry path (see On-chip telemetry from timing, memory and performance counters), or physical sensing validated across devices.
Next readiness level
Assessment and limits · Assessment sources
Related organizations: Machine Intelligence Research Institute, Intelligence Security Laboratories, Amodo Design
Proving takes about 13 minutes (803 seconds) per 2,048-token forward pass of a 13B model on one A100 1, and a verification system design calls the overhead heavy 11.
ZKML and zkLLM prove fixed-point arithmetic 1 3, and floating-point emulation in ZKPs is described as an open problem 11.
zkLLM's code is unaudited, interactive and archived 2; the one audited ZK inference library, ezkl, had high-severity circuit soundness bugs before its fixes 6.
Showing that proven inference was the only work done needs a compute-accounting mechanism such as proof-of-work accounting, which is only proposed 9.
An implementation that is production-grade and available for language models, or relied on by a party other than its developer for a verification decision.
Next readiness level
An independent public security evaluation (audit, red-team or peer-reviewed analysis) of a proof system that handles language models.
Next readiness level
Assessment and limits · Assessment sources
Related organizations: Attestable, University of Waterloo
Proving costs minutes per training step even for small models: 15 minutes per VGG-11 iteration 1 and 47.5 to 328.3 seconds per single-image MobileNet v2 SGD step 2.
The frontier design is unbuilt and lists 13 open problems, including zero-knowledge proofs of backpropagation and deterministic attention backward passes with low overhead 4.
The frontier design needs deterministic training; current deterministic tensor-parallel all-reduce is reported to lose 64 to 89% of bandwidth 4.
The frontier design needs an open-hardware network tap at line rate, listed as an open problem 4.
Mixture-of-experts, reinforcement-learning post-training and multi-site training are not yet covered 4.
Any use by a party other than the developer, or a production-grade, available implementation.
Next readiness level
An independent public security evaluation of a proof-of-training system.
Next readiness level
For the frontier use: an implementation of challenge-based step proofs at realistic model and cluster scale.
Next readiness level
Assessment and limits · Assessment sources
Proving takes about 13 minutes (803 seconds) of A100 time per 2,048-token forward pass at 13B parameters, plus a one-time weight commitment of 16 to 21 minutes.
The repository was archived on 10 July 2025 and the author states there is no plan for upgrades or maintenance.
No security audit of the code has been carried out.
A production-grade implementation, with prover and verifier separated and non-interactive proofs, or reliance by a third party for a verification decision.
Next readiness level
An independent public security evaluation (audit, red-team or third-party peer-reviewed analysis).
Next readiness level
Assessment and limits · Assessment sources
Related organizations: University of Waterloo