Mechanism · Hardware performance throttling and licensing

Technical detail

On this page
  • Knobs evaluated by Ma et al. L2 capacity through way masking, L2 latency through configurable request buffering, L2 bandwidth through credit-based rate limiting, and shared-memory port access through bank arbitration. None is exposed to a software-visible interface 1.
  • Setup. The simulator was AccelSim, configured as an NVIDIA A100, with utilization metrics validated on a real GPU with Nsight Compute. The workloads were CUTLASS StreamK GEMM kernels, for prefill (M=4096) and decode (M=128), and FlashAttention, with shapes from DeepSeek-V3, Llama-3-70B and Mixtral-8x7B, plus six non-LLM CUDA sample workloads 1.
  • Results. At one-eighth of resource availability, performance fell by up to 80% (L2 latency for decode; L2 associativity for prefill). After a throttle was applied, performance settled within about 5K cycles for shared-memory ports, 5–7K cycles for L2 response rate and about 80K cycles for L2 associativity. Each mechanism needs fewer than about 10K flip-flops 1.
  • Selectivity. Memory-side knobs affected LLM kernels selectively. Disabling compute cores degraded nearly every workload together 1.
  • Offline licensing (RAND). A secure message authenticated on the GPU might authorize, for example, 10^18 arithmetic operations, after which the GPU would fall back to 1% of full performance 2.
  • Embedded off-switch (Petrie). Each security block uses public-key signatures, with nonces against replay, to check for recent authorization. Petrie estimates roughly 40,000 transistors per block, and about 0.5% of the die for 10,000 blocks 7. The proof-of-concept README reports signature checks of about 5 million cycles for ECDSA and about 300,000 cycles for HSS/LMS 8.

Search

Full search page