Implementation · GPU contention probes · Draft

Technical detail

On this page
  • Hash search. A custom CUDA implementation of Argon2id runs with 1 pass, 1 lane and 1 MiB of working memory per instance. The difficulty targets about one valid hash per 2^24 trials 1.
  • Delay function. The reported runs use 256 parallel instances of repeated modular squaring, each of 2^20 steps 1.
  • Matrix multiplication. The square matrices have dimension 32,768, with FP16 multiplication and FP32 accumulation. Results are checked with Freivalds' algorithm over 5 rounds 1.
  • Co-running models. The contention experiments used TinyLlama-1.1B at about 2 GB, Qwen2.5-7B at about 9 GB and Llama-2 in FP16 at about 12 GB 1.

Search

Full search page