Data-centre memory challenging
Data-centre memory challenging is a design in Naci Cankaya's low-trust verification overview, published by the Machine Intelligence Research Institute.
A probe near the memory sends challenges and times the answers. The design has two uses. One confirms that information is present where the prover says it is.
The other confirms that no free memory or storage remains, by first filling capacity with incompressible noise and then sampling it at random. Response time is the main evidence, because each tier of memory answers at a different speed. The author treats the design as an optional addition to network taps.
It has not been built. The author knows of no demonstration that a network-level probe can tell one server's memory from another's. Fast remote memory access must be ruled out by latency or by physical disconnection, and filling a pod's memory takes tens of minutes.
The design, its two uses and its assumptions are published, but nothing has been built or measured.
Assessed use: confirming data presence and bounding free memory across data-centre servers
On this page
What it is
Data-centre memory challenging is a design in Naci Cankaya's system overview for low-trust AI compute verification, a working draft from the Machine Intelligence Research Institute's Technical Governance Team 1. It applies timed challenge-response to the memory and storage of AI servers. The author considers it "an optional evidence collection mechanism in addition to network taps" 1. The whole system is covered in Low-trust AI compute verification system overview.
How it works
The design has two uses: "A) verifying presence of information B) verifying the absence of free memory/storage" 1.
- Presence. The verifier challenges the prover about data that the prover claims to hold. Response time is the main evidence, because an answer fetched from another device arrives measurably later 1.
- Absence. Incompressible noise is loaded into the device until its capacity is full, and random samples are then challenged 1.
The overview groups memory into response-time domains. A domain is "any set of locations whose mutual latency differences fall below the measurement resolution at the location of the pinging device" 1.
The latency figures it gives set the margins 1:
- A local DRAM read takes about 70–200 ns.
- A fast data-centre SSD has an average read latency of a few microseconds.
- A remote memory access round trip over InfiniBand or RoCE takes about 1–2 µs.
The author notes that most of these mechanisms need a probe close to the challenged memory 1. He expects that one probe per NVLink scale-up system would cost far less than the hardware it monitors 1.