Implementation · Attestable Audits

CPU-only enclaves force small, quantized models and high cost

On this page

← All known flaws

MinorOpen questionOpen

Memory limits required 4-bit quantization, and the quantized model scored 51.4% on zero-shot MMLU. CPU inference cost 21.7 times as much per token as GPU inference, and the enclave roughly doubled the CPU cost 1. The authors wrote that H100 confidential computing had no multi-GPU support 1. NVIDIA's white paper of August 2025 describes a protected-PCIe mode that passes all eight GPUs of a Hopper HGX node to one confidential VM, with NVLink traffic unencrypted 2.

Sources: [1] · [2]

Search

Full search page