Implementation · Attestable Audits
CPU-only enclaves force small, quantized models and high cost
On this page
MinorOpen questionOpen
Memory limits required 4-bit quantization, and the quantized model scored 51.4% on zero-shot MMLU. CPU inference cost 21.7 times as much per token as GPU inference, and the enclave roughly doubled the CPU cost 1. The authors wrote that H100 confidential computing had no multi-GPU support 1. NVIDIA's white paper of August 2025 describes a protected-PCIe mode that passes all eight GPUs of a Hopper HGX node to one confidential VM, with NVLink traffic unencrypted 2.