# AI verification proposal

A proposal built with the Proposal Explorer of the AI Verification Tech Map (https://trustbutveri.fyi/), from its records of 2026-10-09. Interactive version: https://trustbutveri.fyi/explorer/?mechanisms=M-0002,M-0009,M-0017

How to read it: a claim is something one party wants to verify about another's AI hardware or software. A mechanism is a general technique for verifying claims; it is "aimed at" a claim when that is its direct purpose, and "supporting" when it contributes without being aimed at it. A claim is addressed when a mechanism in the proposal is aimed at it and is not excluded by the filters; addressed does not mean verified, so check that mechanism's development status, security evidence and findings. Definitions: https://trustbutveri.fyi/about/methodology/ (roles, properties and findings) and https://trustbutveri.fyi/about/readiness/ (development status).

## Filters

Filters apply to mechanisms only and describe the setting the proposal is for.

None set. Every mechanism on the map was available.

## Overview

One row per mechanism, read from its record. Open failures: critical / significant / minor. The last three columns are the editors' reading of what the verifier sees. Findings are grouped as known failures, scope limitations and open questions. Only known failures count as failures. Counts are an inventory of published findings, not a risk score.

| Mechanism | Development | Security evidence | Prover | Attack testing | Hardware | Open failures | Weights | Inputs and outputs | Training data |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| Deterministic and bit-exact inference | Operational use | Published security analysis | Adversarial | Analysis | None | 0 / 0 / 0 | depends | depends | not involved |
| Hardware-enabled guarantees (flexHEG) and guarantee processors | Proposed | Published security analysis | Adversarial | Analysis | New chip design | 0 / 3 / 0 | hidden | hidden | hidden |
| Tamper evidence for verifier devices | Research demonstration | Published security analysis | Adversarial | Analysis | Retrofit device | 0 / 2 / 0 | not involved | not involved | not involved |

## Claims

No claims chosen.

## Mechanisms

### Deterministic and bit-exact inference

Making model inference reproducible bit for bit, so that a verifier's re-run must match the provider's output exactly rather than approximately. ([Deterministic and bit-exact inference](https://trustbutveri.fyi/mechanisms/deterministic-inference/))

- Assessment: mechanism family.
- Development: Operational use (legacy code R3), assessed for reproducing open-model inference from receipts in Gensyn's information-market service.
- Security evidence: Published security analysis. Independent evaluation: unassessed. Formal proof: unassessed. Deployment assurance: unassessed.
- Claims in this proposal: none of them.
- Threat model: adversarial prover. Hardware: none. Prover cooperation: required. Attack testing: analysis. Category: Cryptographic & computational.
- What the verifier sees: model weights depends; inputs and outputs depends; training data not involved. Exact replay needs the weights, configuration and replayed requests inside the recomputation environment. What the verifier sees depends on whether that environment keeps them confidential.

### Hardware-enabled guarantees (flexHEG) and guarantee processors

A proposed add-on for AI chips: an auditable guarantee processor, sealed in a tamper-protected enclosure, that would check and enforce agreed rules on chip use. ([Hardware-enabled guarantees (flexHEG) and guarantee processors](https://trustbutveri.fyi/mechanisms/flexheg-guarantee-processors/))

- Assessment: mechanism family.
- Development: Proposed (legacy code R1), assessed for checking and enforcing training-compute limits on chips, against adversaries up to states.
- Security evidence: Published security analysis. Independent evaluation: unassessed. Formal proof: unassessed. Deployment assurance: unassessed.
- Claims in this proposal: none of them.
- Threat model: adversarial prover. Hardware: new chip design. Prover cooperation: required. Attack testing: analysis. Category: On-chip & hardware-enabled.
- What the verifier sees: model weights hidden; inputs and outputs hidden; training data hidden. The guarantee processor sees the chip's traffic inside a sealed enclosure and reports only whether rules were kept.

### Tamper evidence for verifier devices

Enclosures, seals and sensors that make physical interference with verification hardware visible, or that destroy the hardware's secrets when tampering occurs. ([Tamper evidence for verifier devices](https://trustbutveri.fyi/mechanisms/tamper-evidence-for-verifier-devices/))

- Assessment: mechanism family.
- Development: Research demonstration (legacy code R2), assessed for detecting probing of proposed verifier hardware, using server and electronics prototypes as evidence.
- Security evidence: Published security analysis. Independent evaluation: unassessed. Formal proof: unassessed. Deployment assurance: unassessed.
- Claims in this proposal: none of them.
- Threat model: adversarial prover. Hardware: retrofit device. Prover cooperation: partial. Attack testing: analysis. Category: Off-chip devices & sensors.
- What the verifier sees: model weights not involved; inputs and outputs not involved; training data not involved. Protects verifier devices; it does not handle model data.


## Properties

**Operational use**

- Deterministic and bit-exact inference: Operational use (legacy code R3), assessed for reproducing open-model inference from receipts in Gensyn's information-market service

**Built for an adversarial prover**

- Deterministic and bit-exact inference
- Hardware-enabled guarantees (flexHEG) and guarantee processors
- Tamper evidence for verifier devices

**No new hardware needed**

- Deterministic and bit-exact inference


## Attack testing

Attack testing records published testing for this use. It does not by itself show independent review, a formal proof or that a deployed system is secure.

**Testing history**

- Deterministic and bit-exact inference: Analysis
- Hardware-enabled guarantees (flexHEG) and guarantee processors: Analysis
- Tamper evidence for verifier devices: Analysis


## Limits

**Open significant failures**

- State attackers can likely defeat current secure enclosures (known failure, theoretical argument, in Hardware-enabled guarantees (flexHEG) and guarantee processors; https://trustbutveri.fyi/mechanisms/flexheg-guarantee-processors/evidence/flaws/1/) [13][15]. The flexHEG authors write that "nation-state attackers can likely compromise the best current secure enclosures", and that the marginal cost of circumvention per device is hard to estimate. RAND similarly judges that anti-tamper measures "would not be insurmountable for a determined and well-resourced adversary", although they raise costs and can reveal tampering.
- Firmware-only retrofits rely on Secure Boot, which fault injection can bypass (known failure, theoretical argument, in Hardware-enabled guarantees (flexHEG) and guarantee processors; https://trustbutveri.fyi/mechanisms/flexheg-guarantee-processors/evidence/flaws/2/) [13]. Part II notes that the most common attack on Secure Boot replaces the firmware and applies a voltage glitch while the signature is being checked. It also notes that sophisticated actors may use microprobing or laser voltage probing to read key registers.
- FLOP accounting can be laundered through external data (known failure, theoretical argument, in Hardware-enabled guarantees (flexHEG) and guarantee processors; https://trustbutveri.fyi/mechanisms/flexheg-guarantee-processors/evidence/flaws/4/) [13]. Results of earlier or parallel workloads could be hidden in the "external data" fed to a device, which would falsify the total FLOP count unless the inputs are explained or time delays are imposed.
- Seals are often defeated with simple methods (known failure, demonstrated attack, in Tamper evidence for verifier devices; https://trustbutveri.fyi/mechanisms/tamper-evidence-for-verifier-devices/evidence/flaws/1/) [24][25]. Mechanism-class evidence. Published defeats of general security seals. They warn about proposed verifier-device seals, but do not demonstrate defeat of an AI verification enclosure or sensor. In 1996 a Los Alamos vulnerability assessment defeated all 94 security seals it examined, with 132 defeats in total, using rapid, inexpensive, low-tech methods. It found that seal cost did not predict security. In 2001 Johnston reported that high-tech seals are often easier to defeat than low-tech ones.
- Attack classes outside published models (known failure, open question, in Tamper evidence for verifier devices; https://trustbutveri.fyi/mechanisms/tamper-evidence-for-verifier-devices/evidence/flaws/3/) [18][19][26]. Mechanism-class evidence. The radio compensation result is emulated using measured channel data under a known-reference attacker model. It is not a physical bypass demonstration against an AI verifier enclosure. The authors of the batteryless cover say they cannot assess chemical-solvent attacks, which exceed their expertise, and deem cover removal impractical. Anti-Tamper Radio's reference can drift as the environment or measurement system ages; the authors suggest gradually renewing the reference. A 2025 follow-up by some of the same authors shows, by emulation on measured channel data, that an attacker who knows the reference channel and the needle's effect on it could inject a signal that cancels the change caused by a needle insertion. It proposes a reconfigurable intelligent surface that randomizes the channel as a countermeasure.

**Scope limitations**

- Some kernels remain genuinely nondeterministic (scope limitation, open question, in Deterministic and bit-exact inference; https://trustbutveri.fyi/mechanisms/deterministic-inference/evidence/flaws/1/) [1]. The bit-exact work separates kernels that are deterministic but not batch-invariant from truly nondeterministic ones that use atomic functions. Some integer de-quantization kernels use atomic additions and remain nondeterministic, so exact replay needs backends that avoid them.
- Cross-hardware replay relies on reverse-engineered, closed behaviour (scope limitation, open question, in Deterministic and bit-exact inference; https://trustbutveri.fyi/mechanisms/deterministic-inference/evidence/flaws/2/) [1][3]. Emulating one GPU's rounding on another requires reverse-engineering tensor-core arithmetic and modelling proprietary kernel choices. Hawkeye covers a subset of NVIDIA architectures and states that attention and other higher-level operations need further reverse engineering. For the bit-exact emulator, a proprietary Hopper kernel family is an open edge case.
- Many important rules cannot be checked on-chip (scope limitation, theoretical argument, in Hardware-enabled guarantees (flexHEG) and guarantee processors; https://trustbutveri.fyi/mechanisms/flexheg-guarantee-processors/evidence/flaws/3/) [12][14]. Malicious intent "is not a technical property observable on-chip", and misuse depends on what is done with a computation's results. A guarantee processor cannot easily tell whether a network is the whole system or one expert in a mixture-of-experts system. Part III judges that a fully local ruleset "may not be entirely feasible" for the same reason.
- Coverage stops at flexHEG-equipped chips (scope limitation, open question, in Hardware-enabled guarantees (flexHEG) and guarantee processors; https://trustbutveri.fyi/mechanisms/flexheg-guarantee-processors/evidence/flaws/6/) [12][14]. Motivated actors will always be able to use some compute that is not flexHEG-equipped. Recalling existing consumer GPUs would likely be impractical, and reaching perfect coverage, or conclusively proving that no secret government data centres exist, would be "practically quite difficult".

  Related mechanism: Chip registries and manufacturing records (R1, not in the proposal). Accounts for which chips exist and who holds them.

  Related mechanism: Remote detection of data centres (R1, not in the proposal). Looks for undeclared facilities that hold other chips.
- Security depends on inspection protocols (scope limitation, theoretical argument, in Tamper evidence for verifier devices; https://trustbutveri.fyi/mechanisms/tamper-evidence-for-verifier-devices/evidence/flaws/2/) [23][25]. Mechanism-class evidence. An inspection and protocol requirement drawn from safeguards and enclosure studies, not a reported break of a deployed AI verifier. Johnston argues that a seal is no better than the protocols for using it, and that inspectors are usually given little useful information on how to detect tampering. The Sandia survey notes that larger enclosures are hard to inspect fully and that sensor data must be authenticated.

**Open questions**

- Supply-chain diversion and hidden backdoors (open question, open question, in Hardware-enabled guarantees (flexHEG) and guarantee processors; https://trustbutveri.fyi/mechanisms/flexheg-guarantee-processors/evidence/flaws/5/) [13][14]. Components could be diverted before a guarantee processor is added, and backdoors could be introduced during design or manufacturing. Open-source designs and physical scans of randomly selected chips are proposed as countermeasures. Part III proposes international oversight of production and extensive testing of a random sample of finished devices.

  Related mechanism: Chip registries and manufacturing records (R1, not in the proposal). Records each chip's identity and owner from the fab onwards, which bears on diversion before a guarantee processor is fitted. It does not address hidden backdoors.

**Not yet demonstrated**

- Hardware-enabled guarantees (flexHEG) and guarantee processors: Proposed (legacy code R1), assessed for checking and enforcing training-compute limits on chips, against adversaries up to states

**Need new chip designs**

- Hardware-enabled guarantees (flexHEG) and guarantee processors


## Possible additions

Mechanisms on the map, not in the proposal, that the records connect to an unaddressed or partly addressed claim, an open failure or a dependency. Pointers, not recommendations: each brings its own readiness level and findings, and none is claimed to close a failure.

- **Chip registries and manufacturing records** (Proposed (legacy code R1), assessed for a checkable record of which chips were made and who declared owning them)
  - Hardware-enabled guarantees (flexHEG) and guarantee processors waits on it: Governing all relevant chips depends on knowing where they are, through chip registries and detection of undeclared facilities.
- **TEE remote attestation for AI workloads** (Operational use (legacy code R3), assessed for showing which software ran to a party that distrusts the operator holding the hardware)
  - Hardware-enabled guarantees (flexHEG) and guarantee processors depends on it.


## Dependencies

**Missing prerequisites**

- TEE remote attestation for AI workloads (Operational use (legacy code R3), assessed for showing which software ran to a party that distrusts the operator holding the hardware), needed by Hardware-enabled guarantees (flexHEG) and guarantee processors
- Chip registries and manufacturing records (Proposed (legacy code R1), assessed for a checkable record of which chips were made and who declared owning them), needed by Hardware-enabled guarantees (flexHEG) and guarantee processors

**Blockers**

- Deterministic and bit-exact inference: Batch-invariant kernels cost throughput: in Thinking Machines' Qwen3-8B test, an improved deterministic build took 42 s against 26 s for vLLM's default, and SGLang reports an average 34.35% slowdown on its FlashInfer and FlashAttention 3 backends. (performance & compatibility) [2][4]
- Deterministic and bit-exact inference: Coverage is incomplete: the bit-exact emulator targets dense blocks on NVIDIA GPUs and excludes mixture-of-experts inference and training, and vLLM's batch-invariant mode is in beta, with open work on AMD hardware and speculative decoding. (performance & compatibility) [1][5][27]
- Deterministic and bit-exact inference: Amodo's status page for the AI 2040 verification plan rates a reproducible inference stack for that plan as 'not started'. (performance & compatibility) [28]
- Deterministic and bit-exact inference: Exact replay requires the prover to disclose weights, software versions, parallelism and batch sizes to whoever recomputes. (privacy & leakage) [1][11]
- Hardware-enabled guarantees (flexHEG) and guarantee processors: Integrated flexHEG needs substantial help from the accelerator manufacturer, and the authors estimate 3.7–7.9 years, from when the manufacturer starts work, for such hardware to displace other accelerators in frontier development. (access & governance) [13]
- Hardware-enabled guarantees (flexHEG) and guarantee processors: State-level attackers who hold the hardware can likely compromise the best current secure enclosures. (hardware trust; waits on Tamper evidence for verifier devices) [13][15]
- Hardware-enabled guarantees (flexHEG) and guarantee processors: Rival states would need to trust the design and manufacture of guarantee processors and enclosures, for example through open design, redundant processors from each side or oversight of production. (hardware trust) [12][14]
- Hardware-enabled guarantees (flexHEG) and guarantee processors: Restricting future rule updates would need a formal language for rules, which the authors judge most likely infeasible for early flexHEG versions. (protocol soundness) [12]
- Hardware-enabled guarantees (flexHEG) and guarantee processors: Governing all relevant chips depends on knowing where they are, through chip registries and detection of undeclared facilities. (coverage & hidden compute; waits on Chip registries and manufacturing records) [14]
- Tamper evidence for verifier devices: No tamper-evident enclosure has been designed for AI verifier hardware at retrofit scale. (hardware trust) [11]
- Tamper evidence for verifier devices: Battery-backed designs add bulk, limit operating temperature (+10 °C to +35 °C for the IBM 4765) and complicate transport. (performance & compatibility) [19]
- Tamper evidence for verifier devices: Active monitoring needs power, and visual inspection of large enclosures faces access limits. (access & governance) [23]
- Tamper evidence for verifier devices: No evaluation has been published in the AI verification setting. (adversarial validation) [11]


## What the verifier sees

- Model weights: shown by none; depends on the design for Deterministic and bit-exact inference; hidden by Hardware-enabled guarantees (flexHEG) and guarantee processors; not involved in Tamper evidence for verifier devices; unspecified for none.
- Inputs and outputs: shown by none; depends on the design for Deterministic and bit-exact inference; hidden by Hardware-enabled guarantees (flexHEG) and guarantee processors; not involved in Tamper evidence for verifier devices; unspecified for none.
- Training data: shown by none; depends on the design for none; hidden by Hardware-enabled guarantees (flexHEG) and guarantee processors; not involved in Deterministic and bit-exact inference and Tamper evidence for verifier devices; unspecified for none.

## Implementations

- Deterministic and bit-exact inference: [Batch-invariant inference kernels (Thinking Machines)](https://trustbutveri.fyi/implementations/batch-invariant-inference-kernels/) (R2, open-source project); [Verde and RepOps (Gensyn)](https://trustbutveri.fyi/implementations/gensyn-verde-repops/) (R3, product); [Low-trust AI compute verification system overview](https://trustbutveri.fyi/implementations/low-trust-compute-verification-system-overview/) (R1, proposed architecture)
- Hardware-enabled guarantees (flexHEG) and guarantee processors: none on the map
- Tamper evidence for verifier devices: [AI 2040 inference-only verification stack](https://trustbutveri.fyi/implementations/ai-2040-inference-only-verification-plan/) (R1, proposed architecture)

## Sources

1. Bit-Exact AI Inference Verification Without Performance Tradeoffs, N. Cankaya (2026). https://arxiv.org/abs/2606.00279
2. Defeating Nondeterminism in LLM Inference, H. He & Thinking Machines Lab (2025). https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/
3. Hawkeye: Reproducing GPU-Level Non-Determinism, E. Badash et al. (2026). https://proceedings.mlsys.org/paper_files/paper/2026/hash/e217c271a57c365a246b0ad39e668ba8-Abstract-Conference.html
4. Towards Deterministic Inference in SGLang and Reproducible RL Training, The SGLang Team (2025). https://www.lmsys.org/blog/2025-09-22-sglang-deterministic/
5. Batch Invariance (vLLM documentation), vLLM project (2026). https://github.com/vllm-project/vllm/blob/main/docs/features/batch_invariance.md
6. gensyn-ai/ree: Gensyn Reproducible Execution Environment (GitHub repository), Gensyn (2026). https://github.com/gensyn-ai/ree
7. EigenCloud Brings Verifiable AI to Mass Market with EigenAI and EigenCompute Launches, EigenCloud (2025). https://www.eigenlabs.org/blog/eigencloud-brings-verifiable-ai-to-mass-market-with-eigenai-and-eigencompute-launches/
8. Building Delphi: Pricing, Settlement, and Agentic Trading, D. Jedamski (2026). https://www.gensyn.ai/blog/building-delphi-pricing-settlement-and-agentic-trading
9. Reproducible Execution Environment (REE) (Gensyn documentation), Gensyn (2026). https://docs.gensyn.ai/tech
10. What is Delphi? (Delphi documentation), Gensyn (2026). https://docs.delphi.fyi/
11. A System Overview for Near-Term, Low-Trust AI Compute Verification, N. Cankaya (2026). https://intelligence.org/wp-content/uploads/2026/06/A-system-overview-for-near-term-low-trust-AI-compute-verification.pdf
12. Flexible Hardware-Enabled Guarantees for AI Compute, J. Petrie et al. (2025). https://arxiv.org/abs/2506.15093
13. Technical Options for Flexible Hardware-Enabled Guarantees, J. Petrie & O. Aarne (2025). https://arxiv.org/abs/2506.03409
14. International Security Applications of Flexible Hardware-Enabled Guarantees, O. Aarne & J. Petrie (2025). https://arxiv.org/abs/2506.15100
15. Hardware-Enabled Governance Mechanisms: Developing Technical Solutions to Exempt Items Otherwise Classified Under Export Control Classification Numbers 3A090 and 4A090, G. Kulp et al. (2024). https://www.rand.org/pubs/working_papers/WRA3056-1.html
16. Secure, Governable Chips: Using On-Chip Mechanisms to Manage National Security Risks from AI & Advanced Computing, O. Aarne et al. (2024). https://www.cnas.org/publications/reports/secure-governable-chips
17. Hardware-Enabled Mechanisms for Verifying Responsible AI Development, A. O'Gara et al. (2025). https://arxiv.org/abs/2505.03742
18. Anti-Tamper Radio: System-Level Tamper Detection for Computing Systems, P. Staat et al. (2022). https://ieeexplore.ieee.org/document/9833631/
19. Secure Physical Enclosures from Covers with Tamper-Resistance, V. Immler et al. (2019). https://tches.iacr.org/index.php/TCHES/article/view/7334
20. ImpedanceVerif: On-Chip Impedance Sensing for System-Level Tampering Detection, T. Mosavirik et al. (2023). https://eprint.iacr.org/2022/946
21. IBM 4765 Cryptographic Coprocessor Security Module: Security Policy, IBM Corporation (2012). https://csrc.nist.gov/csrc/media/projects/cryptographic-module-validation-program/documents/security-policies/140sp1505.pdf
22. PHYSEC SEAL: Change detection for maximum safety, PHYSEC GmbH (2026). https://www.physec.de/en/solutions/physec-seal/
23. Tamper-Indicating Enclosures, A Current Survey, H. A. Smartt & Z. N. Gastelum (2015). https://www.osti.gov/servlets/purl/1256541
24. Physical Security and Tamper-Indicating Devices, R. G. Johnston & A. R. E. Garcia (1996). https://www.osti.gov/servlets/purl/459707
25. Tamper Detection for Safeguards and Treaty Monitoring: Fantasies, Realities, and Potentials, R. G. Johnston (2001). https://www.nonproliferation.org/wp-content/uploads/npr/81john.pdf
26. Anti-Tamper Radio Meets Reconfigurable Intelligent Surface for System-Level Tamper Detection, M. S. Tabar et al. (2025). https://arxiv.org/abs/2503.14279
27. [Feature]: Batch Invariant Feature and Performance Optimization (vLLM issue #27433), vLLM project contributors (2025). https://github.com/vllm-project/vllm/issues/27433
28. AI 2040 Plan A — Verification SITREP, Amodo Design (2026). https://amododesign.com/ai-verification/plan-a-sitrep/
