Implementation · GPU contention probes · Draft
Technical detail
On this page
- Hash search. A custom CUDA implementation of Argon2id runs with 1 pass, 1 lane and 1 MiB of working memory per instance. The difficulty targets about one valid hash per 2^24 trials 1.
- Delay function. The reported runs use 256 parallel instances of repeated modular squaring, each of 2^20 steps 1.
- Matrix multiplication. The square matrices have dimension 32,768, with FP16 multiplication and FP32 accumulation. Results are checked with Freivalds' algorithm over 5 rounds 1.
- Co-running models. The contention experiments used TinyLlama-1.1B at about 2 GB, Qwen2.5-7B at about 9 GB and Llama-2 in FP16 at about 12 GB 1.