Mechanism · On-chip telemetry from timing, memory and performance counters
Technical detail
On this page
- Counter-based classification. Rahman and Tajdari sample nine always-available NVML counters at 1 Hz: GPU and memory utilization, memory used, power, temperature, SM and memory clocks, and PCIe TX/RX. They extract 166 features over 5–60 s windows, including memory slope and epoch periodicity from an FFT of power 3.
- Memory-hard proof of work. In Monfared et al.'s challenge suite, these puzzles expose parallel effort and HBM use 1.
- Verifiable delay functions. Based on sequential modular squaring, they expose sequential compute pressure 1.
- GEMM puzzles. They target tensor-core throughput; the authors note that over 90% of LLM floating-point operations are GEMMs. Results can be checked with Freivalds' algorithm, subject to floating-point rounding discrepancies 1.
- VRAM residency test. It runs bandwidth-bound Argon2id over challenge data. With 60 GB of challenge data on an H100, the response time for data held in HBM and for data in pinned host memory reached over PCIe differed by more than 350 ms 1.
- Guaranteeable Memory. A guarantee chiplet beneath the HBM stacks would observe memory traffic directly. The author argues that the HBM standard makes it compatible with multiple leading accelerators 2.
- Metering targets. Candidate targets for licensing include floating-point and integer arithmetic, memory, NVLink and PCIe transfer volume, energy and clock cycles 5.