Mechanism · Memory wiping and proofs of secure erasure

Technical detail

On this page
  • Origin. Perito and Tsudik introduced proofs of secure erasure (PoSE) for embedded devices with bounded memory and a small ROM 2.
  • Bursuc et al.'s model. The adversary is a distant, memory-unbounded helper A0 and a local device A1 bounded to memory M. The protocol has an initialization phase (fill memory), r timed challenge-response rounds, and a verification phase that accepts only if the answers are correct and each round-trip time is at most Δ 3.
  • Variants. An unconditional variant fills memory with random bits sent by the verifier. A graph-based variant sends only a seed: the prover computes labels of a depth-robust graph with a hash function and stores the output labels, and a construction with in-place labelling needs only about the output size plus O(w) memory. A lightweight variant relaxes depth-robustness to a small constant for speed 3.
  • Prototype. Bursuc et al. ran their prototype on a standard desktop computer and simulated the erasure of 32 KB, which took 0.25 s; compiled for a 32-bit architecture, the program is 3.4 KB 3.
  • Amodo's implementation. It follows Bursuc et al.: node labels are L(n) = H(n ∥ L(p1) ∥ … ∥ L(pk)), only "robust" labels are stored and challengeable, and the scheme is secure if q < γ, where q is the number of hashes a cheater can compute within the round-trip time and γ is the minimum number of hashes needed to recompute a robust label 5.
  • Amodo's parameters and throughput. An assumed RTT of 1 ms; γ = 65,536 for host RAM and 32,768 for GPU HBM; Dual-AES-PRF on CPU and BLAKE3 on GPU. The measured throughputs, 126.3 MiB/s for 120 GB of RAM and 244.5 MiB/s for 140 GB of HBM, scale to 2,565 s for the memory sizes in a GB200 tray 5.
  • Disk path. A later optimization parallelized label generation across GPUs and reached 6 min 41 s per TB with 6 GPUs on one fast drive; it needs 25 GiB of working memory, which is left unattested if used on HBM 6. The NVL72 storage estimate assumes 30.72 TB of drives per tray and an estimated 840 MiB/s per B200 GPU, four GPUs per tray 6. Code for this path (CUDA graph labeller, multi-GPU disk-wipe benchmark with sample verification, wipe-time calculator) is public under the MIT licence 7.

Search

Full search page