# AI verification proposal

A proposal built with the Proposal Explorer of the AI Verification Tech Map (https://trustbutveri.fyi/), from its records of 2026-10-09. Interactive version: https://trustbutveri.fyi/explorer/?mechanisms=M-0001,M-0010&implementations=M-0001:I-0002

How to read it: a claim is something one party wants to verify about another's AI hardware or software. A mechanism is a general technique for verifying claims; it is "aimed at" a claim when that is its direct purpose, and "supporting" when it contributes without being aimed at it. A claim is addressed when a mechanism in the proposal is aimed at it and is not excluded by the filters; addressed does not mean verified, so check that mechanism's development status, security evidence and findings. Definitions: https://trustbutveri.fyi/about/methodology/ (roles, properties and findings) and https://trustbutveri.fyi/about/readiness/ (development status).

## Filters

Filters apply to mechanisms only and describe the setting the proposal is for.

None set. Every mechanism on the map was available.

## Overview

One row per mechanism, read from its record. Open failures: critical / significant / minor. The last three columns are the editors' reading of what the verifier sees. Findings are grouped as known failures, scope limitations and open questions. Only known failures count as failures. Counts are an inventory of published findings, not a risk score.

| Mechanism | Development | Security evidence | Prover | Attack testing | Hardware | Open failures | Weights | Inputs and outputs | Training data |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| Sampled inference recomputation / DiFR (Divergence From Reference) | Research demonstration | Published security analysis | Adversarial | Analysis | None | 0 / 1 / 1 | unspecified | unspecified | unspecified |
| On-chip telemetry from timing, memory and performance counters | Research demonstration | Published attack testing | Semi-trusted | Red-teamed | Existing features | 0 / 2 / 0 | depends | depends | depends |

## Claims

No claims chosen.

## Mechanisms

### Sampled inference recomputation

DiFR checks that an inference provider ran its declared model by comparing output tokens or activations with a trusted re-run using the same random seed. ([Sampled inference recomputation](https://trustbutveri.fyi/mechanisms/sampled-inference-recomputation/))

- Assessment: selected implementation [DiFR (Divergence From Reference)](https://trustbutveri.fyi/implementations/difr/).
- Development: Research demonstration (legacy code R2), assessed for checking that outputs match the declared model, precision and sampling settings.
- Security evidence: Published security analysis. Independent evaluation: unassessed. Formal proof: unassessed. Deployment assurance: unassessed.
- Claims in this proposal: none of them.
- Threat model: adversarial prover. Hardware: none. Prover cooperation: required. Attack testing: analysis. Category: Cryptographic & computational.
- What the verifier sees: model weights unspecified; inputs and outputs unspecified; training data unspecified. This Explorer has no asset-specific exposure assessment for this implementation. Check its source and deployment assumptions.

### On-chip telemetry from timing, memory and performance counters

Uses on-chip measurements, such as task timings, whether data sits in chip memory, and performance counters, as evidence of what AI chips are running. ([On-chip telemetry from timing, memory and performance counters](https://trustbutveri.fyi/mechanisms/on-chip-telemetry/))

- Assessment: mechanism family.
- Development: Research demonstration (legacy code R2), assessed for workload evidence from GPU counters and timing, assuming authentic measurements.
- Security evidence: Published attack testing. Independent evaluation: unassessed. Formal proof: unassessed. Deployment assurance: unassessed.
- Claims in this proposal: none of them.
- Threat model: semi-trusted prover. Hardware: existing features. Prover cooperation: partial. Attack testing: red-teamed. Category: On-chip & hardware-enabled.
- What the verifier sees: model weights depends; inputs and outputs depends; training data depends. Counters do not read weights or data, but richer counters can leak secrets through side channels.


## Properties

**Built for an adversarial prover**

- Sampled inference recomputation

**No new hardware needed**

- Sampled inference recomputation
- On-chip telemetry from timing, memory and performance counters


## Attack testing

Attack testing records published testing for this use. It does not by itself show independent review, a formal proof or that a deployed system is secure.

**Testing history**

- Sampled inference recomputation / DiFR (Divergence From Reference): Analysis
- On-chip telemetry from timing, memory and performance counters: Red-teamed


## Limits

**Open significant failures**

- Statistical tolerance leaves a covert channel (known failure, demonstrated attack, in DiFR (Divergence From Reference); https://trustbutveri.fyi/implementations/difr/evidence/flaws/1/) [2][7][9]. Statistical schemes can put an upper bound on an adversary's covert bandwidth, but cannot close it. In the companion exfiltration study, the detector cut exfiltratable information to under 0.5% under benign prompt traffic. It did not cut it to zero. Its authors called the channel's size under adversarial prompts an open empirical question. An independent study attacked that detector, which uses the same Gumbel-margin statistic. An adversary who controls the prompts roughly doubled the bits leaked per token. Across six models, this cut the slowdown from 146–254 times under benign prompts to 60–118 times. The attack widens the exfiltration bound. It does not target DiFR's check that outputs match the declared configuration.
- Adversarially disguised fine-tuning partly evades classification (known failure, demonstrated attack, in On-chip telemetry from timing, memory and performance counters; https://trustbutveri.fyi/mechanisms/on-chip-telemetry/evidence/flaws/2/) [14]. Across 20 evasion strategy families in five monitor-evader rounds, the classifier's accuracy against the most challenging disguised workloads held out from training was 43–87%. White-box LoRA fine-tuning was the only evasion family detected less than half the time. The evaluation covered single nodes, PyTorch-level evasion and NVIDIA hardware.
- Counters leak information about protected workloads (known failure, theoretical argument, in On-chip telemetry from timing, memory and performance counters; https://trustbutveri.fyi/mechanisms/on-chip-telemetry/evidence/flaws/4/) [15][16]. Performance counters have been used as a side channel against TEEs, for example in CounterSEVeillance. NVIDIA disables performance counters in full confidential-computing mode, stating that they could provide an avenue for side-channel attacks. Richer counters for verification therefore pull against confidentiality.

**Family finding context**

- Context for DiFR (Divergence From Reference). Findings from the mechanism family appear here as context. They apply to an implementation only when its own record lists them, under the conditions stated there. Tolerance for numerical noise leaves a covert channel (known failure, demonstrated attack, in Sampled inference recomputation; https://trustbutveri.fyi/mechanisms/sampled-inference-recomputation/evidence/flaws/1/) [2][7][9]. Schemes that accept approximate matches can put an upper bound on an adversary's covert bandwidth, but they cannot close the channel. The weight-exfiltration detector cut exfiltratable information to under 0.5%, not to zero, on a 30-billion-parameter mixture-of-experts model under benign prompt traffic. Its authors called the channel's size under adversarial prompts an open empirical question. An independent study showed that an adversary who controls the prompts roughly doubles the bits leaked per token. Across six models, that cut the slowdown from 146–254 times under benign prompts to 60–118 times. The attack widens the exfiltration bound. It does not target the check that outputs match the declared model.

  Related mechanism: Deterministic and bit-exact inference (R3, not in the proposal). Bit-exact inference would remove the numerical tolerance if exact replay can be deployed with the required weights and configuration.
- Context for DiFR (Divergence From Reference). Findings from the mechanism family appear here as context. They apply to an implementation only when its own record lists them, under the conditions stated there. Only recorded traffic is checked (scope limitation, theoretical argument, in Sampled inference recomputation; https://trustbutveri.fyi/mechanisms/sampled-inference-recomputation/evidence/flaws/2/) [2][10]. Recomputation checks that recorded, declared workloads are correct. It cannot show that the record is complete. The published schemes do not cover hidden workloads run on the same compute, or substituted work. Rinberg et al. say their exfiltration-detection scheme cannot stand alone.

  Related mechanism: Network taps and certifiers (R1, not in the proposal). Taps copy and hash all traffic on the monitored links, which bears on whether the traffic record is complete. They do not show what else ran on the same chips.
- Context for DiFR (Divergence From Reference). Findings from the mechanism family appear here as context. They apply to an implementation only when its own record lists them, under the conditions stated there. Some inference optimizations are not covered (known failure, theoretical argument, in Sampled inference recomputation; https://trustbutveri.fyi/mechanisms/sampled-inference-recomputation/evidence/flaws/3/) [1][11]. TOPLOC's authors state that it cannot detect speculative decoding in which a cheaper model does the decoding. They did not test whether it distinguishes types of key-value (KV) cache compression. DiFR was evaluated only on sampling from a single model. Its authors sketch an extension to one speculative-decoding algorithm but do not test it.
- Context for DiFR (Divergence From Reference). Findings from the mechanism family appear here as context. They apply to an implementation only when its own record lists them, under the conditions stated there. Mixed hardware widens the honest baseline (known failure, open question, in Sampled inference recomputation; https://trustbutveri.fyi/mechanisms/sampled-inference-recomputation/evidence/flaws/4/) [1]. When honest reference runs span different GPU types, the spread of benign scores grows. In DiFR's tests on Qwen3-30B-A3B, pooling A100 and H200 runs left Token-DiFR unable to separate the two smallest tested changes, a temperature of 1.1 instead of 1.0 and a simulated top-2 sampling bug, at the target false-positive rate, while cross-entropy separated them. Matched provider and verifier environments, or pooling that weights rare large deviations, restored detection.

**Open minor failures**

- Mixed hardware widens the honest baseline (known failure, open question, in DiFR (Divergence From Reference); https://trustbutveri.fyi/implementations/difr/evidence/flaws/2/) [1]. For Qwen3-30B-A3B, pooling honest runs across A100 and H200 GPUs and parallelism setups broadened the honest score distribution. Token-DiFR then failed to separate the two smallest tested changes, a temperature of 1.1 instead of 1.0 and a simulated top-2 sampling bug, at the target false-positive rate, while cross-entropy did. The authors report that matched provider and verifier environments, or pooling that weights rare large deviations, restore detection.

**Scope limitations**

- Software-read telemetry can be forged by the operator (scope limitation, theoretical argument, in On-chip telemetry from timing, memory and performance counters; https://trustbutveri.fyi/mechanisms/on-chip-telemetry/evidence/flaws/1/) [12][14]. NVML-based classification assumes trustworthy telemetry. Without a tamper-resistant read path, an authenticated telemetry channel and secure boot of the monitoring software, an operator who controls the full software stack could forge counter values. Monfared et al. start from the same premise: current GPUs expose little trusted telemetry and can be modified or virtualized.

  Related mechanism: Hardware-enabled guarantees (flexHEG) and guarantee processors (R1, not in the proposal). A guarantee processor on the chip would give the tamper-resistant, authenticated telemetry path the flaw says is missing.
- Timing challenges do not identify the individual chip (scope limitation, theoretical argument, in On-chip telemetry from timing, memory and performance counters; https://trustbutveri.fyi/mechanisms/on-chip-telemetry/evidence/flaws/3/) [12]. GEMM and VDF challenges can be answered by identical GPUs elsewhere, and floating-point fingerprints distinguish GPU models, not individual devices. GPU virtualization adds timing leakage that prevents attributing compute use.

**Open questions**

- Speculative decoding and multi-model sampling not evaluated (open question, open question, in DiFR (Divergence From Reference); https://trustbutveri.fyi/implementations/difr/evidence/flaws/3/) [1]. The algorithms and experiments cover sampling from a single LLM. Speculative decoding was not evaluated. The authors sketch an extension to one speculative-decoding algorithm, without experiments. They note that other variants would need modified verification and extra metadata.
- No quantified error rates or formal thresholds for timing primitives (open question, open question, in On-chip telemetry from timing, memory and performance counters; https://trustbutveri.fyi/mechanisms/on-chip-telemetry/evidence/flaws/5/) [12]. Monfared et al. state that false-positive and false-negative rates are not quantified and leave hardware-specific formal thresholds to future work.


## Possible additions

Mechanisms on the map, not in the proposal, that the records connect to an unaddressed or partly addressed claim, an open failure or a dependency. Pointers, not recommendations: each brings its own readiness level and findings, and none is claimed to close a failure.

- **Hardware-enabled guarantees (flexHEG) and guarantee processors** (Proposed (legacy code R1), assessed for checking and enforcing training-compute limits on chips, against adversaries up to states)
  - On-chip telemetry from timing, memory and performance counters waits on it: Shipping accelerators need a tamper-resistant, authenticated telemetry path.
- **TEE remote attestation for AI workloads** (Operational use (legacy code R3), assessed for showing which software ran to a party that distrusts the operator holding the hardware)
  - On-chip telemetry from timing, memory and performance counters depends on it.


## Dependencies

**Missing prerequisites**

- TEE remote attestation for AI workloads (Operational use (legacy code R3), assessed for showing which software ran to a party that distrusts the operator holding the hardware), needed by On-chip telemetry from timing, memory and performance counters

**Blockers**

- Sampled inference recomputation: The verifier needs the model weights, so outsiders cannot use the method to verify providers of closed-weights models. (privacy & leakage) [1]
- Sampled inference recomputation: The verifier must know and match the provider's sampling procedure, and in one prototype a sampling mismatch in a newer vLLM version produced large spurious logit differences. (performance & compatibility) [1][4]
- Sampled inference recomputation: No independent red-team of DiFR's consistency check has been published, Amodo rates recomputation red-teaming 'not started', and the one independent attack study targets an exfiltration detector built on the same statistic. (adversarial validation) [6][7]
- On-chip telemetry from timing, memory and performance counters: Shipping accelerators need a tamper-resistant, authenticated telemetry path. (hardware trust; waits on Hardware-enabled guarantees (flexHEG) and guarantee processors) [13][14]
- On-chip telemetry from timing, memory and performance counters: NVIDIA's full confidential-computing mode disables the hardware performance counters its profiling tools use, so telemetry that needs them conflicts with it. (privacy & leakage) [15][16]
- On-chip telemetry from timing, memory and performance counters: Continuous challenge puzzles cost power and throughput on production workloads. (performance & compatibility) [12]
- On-chip telemetry from timing, memory and performance counters: Evaluation has not gone beyond single nodes, framework-level evasion and one vendor's hardware. (adversarial validation) [14]


## What the verifier sees

- Model weights: shown by none; depends on the design for On-chip telemetry from timing, memory and performance counters; hidden by none; not involved in none; unspecified for Sampled inference recomputation.
- Inputs and outputs: shown by none; depends on the design for On-chip telemetry from timing, memory and performance counters; hidden by none; not involved in none; unspecified for Sampled inference recomputation.
- Training data: shown by none; depends on the design for On-chip telemetry from timing, memory and performance counters; hidden by none; not involved in none; unspecified for Sampled inference recomputation.

## Implementations

- Sampled inference recomputation: [AI 2040 inference-only verification stack](https://trustbutveri.fyi/implementations/ai-2040-inference-only-verification-plan/) (R1, proposed architecture); [DiFR (Divergence From Reference)](https://trustbutveri.fyi/implementations/difr/) (R2, research prototype); [Low-trust AI compute verification system overview](https://trustbutveri.fyi/implementations/low-trust-compute-verification-system-overview/) (R1, proposed architecture); [SASH confidential network logger](https://trustbutveri.fyi/implementations/sash-confidential-network-logger/) (R1, research prototype); [TOPLOC](https://trustbutveri.fyi/implementations/toploc/) (R3, open-source project)
- On-chip telemetry from timing, memory and performance counters: none on the map

## Sources

1. DiFR: Inference Verification Despite Nondeterminism, A. Karvonen et al. (2025). https://arxiv.org/abs/2511.20621
2. Verifying LLM Inference to Detect Model Weight Exfiltration, R. Rinberg et al. (2025). https://arxiv.org/abs/2511.02620
3. adamkarvonen/difr (GitHub repository), A. Karvonen (2025). https://github.com/adamkarvonen/difr
4. Scaling Recomputation Inference Verification, Amodo Design (2026). https://amododesign.com/notes/2026-09-02-scaling-recomputation-inference-verification/
5. Amodo-Design/Inference-Recomputation-Prototype (GitHub repository), Amodo Design (2026). https://github.com/Amodo-Design/Inference-Recomputation-Prototype
6. AI 2040 Plan A — Verification SITREP, Amodo Design (2026). https://amododesign.com/ai-verification/plan-a-sitrep/
7. Adversarial Entropy Inflation Against Gumbel-Based Inference Verification, N. Kezins (2026). https://arxiv.org/abs/2608.23375
8. An Inference Verification Prototype — Stage 1, Amodo Design (2026). https://amododesign.com/notes/2026-06-29-inference-verification-prototype/
9. Bit-Exact AI Inference Verification Without Performance Tradeoffs, N. Cankaya (2026). https://arxiv.org/abs/2606.00279
10. Example Schemes for Verifying High-Stakes AI Agreements, Amodo Design (2026). https://amododesign.com/notes/2026-06-23-verification-algorithms/
11. TOPLOC: A Locality Sensitive Hashing Scheme for Trustless Verifiable Inference, J. M. Ong et al. (2025). https://proceedings.mlr.press/v267/ong25a.html
12. Timing and Memory Telemetry on GPUs for AI Governance, S. K. Monfared et al. (2026). https://arxiv.org/abs/2602.09369
13. Guaranteeable Memory: An HBM-Based Chiplet for Verifiable AI Workloads, J. Petrie (2025). https://openreview.net/forum?id=uc79kOv0MV
14. Detecting Hidden ML Training With Zero-Overhead Telemetry, R. Rahman & S. Tajdari (2026). https://arxiv.org/abs/2606.19262
15. On TEEs for Privacy-Preserving Monitoring in AI Governance, Gloria Z (2026). https://techgov.intelligence.org/blog/on-tees-for-privacy-preserving-monitoring-in-ai-governance
16. NVIDIA Secure AI with Blackwell and Hopper GPUs (White Paper), NVIDIA (2025). https://docs.nvidia.com/nvidia-secure-ai-with-blackwell-and-hopper-gpus-whitepaper.pdf
