Mechanism · On-chip telemetry from timing, memory and performance counters
Evidence & limits
On this page
R2Demonstrated for workload evidence from GPU counters and timing, assuming authentic measurements
Detailed results exist on current GPUs against a stated adversary, but the data can be forged in software and no one outside the developers has evaluated the methods.
Assessed use: workload evidence from GPU counters and timing, assuming authentic measurements
Rubric assessment
- R1 met: primitives, claims and threat models are published 1 2 3.
- R2 met through reproducible published results. Rahman and Tajdari report end-to-end results on nine NVIDIA GPU models, including H100, H200 and B200, over five monitor-evader rounds against a stated adversary. They state that code for their workloads is in the project repository, but give no repository URL 3; the rubric does not require code for this route. Monfared et al. report measurements on T4 and H100 GPUs with language models of up to 72B parameters 1. No Implementation record realises this mechanism, and the most mature work is the NVML classifier.
- R3 not met: no production-grade verification system built on these signals is available, and no source reports a party other than the developers relying on them for a verification decision. NVIDIA's Fleet Intelligence is a production telemetry service, but NVIDIA describes it as a tool for operators to monitor their own fleets 9. Trustworthy collection also needs hardware protections that Rahman and Tajdari list as prerequisites 3.
- R4 not met: as of September 2026 no independent evaluation has been published.
- Use by a party other than the developers for a verification decision.
- Telemetry read paths that the operator cannot forge, such as signed counters from a root of trust or a guarantee processor.
- Calibrated false-positive and false-negative rates, with a detection-theoretic threshold framework.
- Independent red-teaming, including custom-kernel and multi-node evasion.
Assessed 2026-10-08 against rubric v1.1.
Evidence
- Rahman and Tajdari. They tested nine NVIDIA GPU models from four architecture generations, including A100, H100, H200 and B200, with up to eight GPUs per node. The classifier reached 98.2% binary accuracy at identifying training across their whole corpus. After five monitor-evader rounds covering 20 evasion strategy families, it reached 43–87% accuracy on the most challenging disguised workloads held out from training. The authors state that code for the workloads is in the project repository, but give no URL 3.
- Monfared et al. They measured T4 and H100 GPUs running language models from 1.1B to 72B parameters. With 60 GB of challenge data on an H100, the residency test separated on-device from off-device data by more than 350 ms. Continuous puzzles added noticeable power overhead and cut throughput. The residency test's power and throughput costs were negligible, but by design it occupies a large part of the GPU's memory. The paper mentions no code release 1.
- Guaranteeable Memory. It is a workshop proposal with no reported prototype 2.
- NVIDIA Fleet Intelligence. NVIDIA reports that this service, generally available since May 2026, uses a read-only host agent to stream GPU power, temperature, performance, health and configuration telemetry to an NVIDIA cloud service. It has released the agent as open source. NVIDIA reports that the agent also collects GPU attestation evidence signed with on-device certificates, but does not say that the telemetry values themselves are signed 9.
Limitations
- Forgery. If the operator controls the software that reads the counters, the counters prove little 3.
- Disguise. Adversarial disguise works in part: white-box LoRA fine-tuning was detected less than half the time 3.
- Outsourcing and virtualization. Timing challenges can be answered by identical GPUs, and virtualization confounds attribution 1.
- Checking matrix results. Floating-point rounding complicates the check 1.
- Error rates. The timing primitives have no calibrated error rates 1.
- Evaluation scope. The classifier was evaluated only on single nodes, with evasion at the PyTorch level and sampling at about 1 Hz 3.
- Power sampling. Yang and colleagues found that on A100 and H100 GPUs the built-in power reading, which nvidia-smi obtains through NVML, samples only 25% of runtime. The GPU can draw very different power in the other 75% without the reading showing it 8. They also found the reading's error to be within about ±5% in most cases, against the ±5 W that NVIDIA claims 8.
- Leakage. Richer counters risk leaking secrets 6.
Known flaws
Blockers
Shipping accelerators need a tamper-resistant, authenticated telemetry path.
NVIDIA's full confidential-computing mode disables the hardware performance counters its profiling tools use, so telemetry that needs them conflicts with it.
Continuous challenge puzzles cost power and throughput on production workloads.
Evaluation has not gone beyond single nodes, framework-level evasion and one vendor's hardware.