Mechanism · Bandwidth limits and compartmentalization

Evidence & limits

On this page

R2Demonstrated for monitoring inter-node traffic with operator-run software on four GPUs

R2, narrowly: a public software prototype on data-centre GPUs monitors inter-node traffic against a limit and detects low-communication training, but no cap that a verifier can check has been built.

Assessed use: monitoring inter-node traffic with operator-run software on four GPUs

Rubric assessment

  • R1 met: the AI 2040 plan proposes isolated inference units created by removing back-end networking, on the grounds that high-bandwidth links are mostly needed only for training 1. Lucid Computing gives a detailed design with a cap, a pod definition, adversary strategies, favourable assumptions and residual risks 2. The MIRI overview describes perimeters of monitored links sized to pods for inference 4.
  • R2 met: MIRI's Technical Governance Team published an end-to-end prototype, with code, that monitors inter-node traffic against a threshold on two nodes with four A100 GPUs 5. It tested violating training, compliant inference and a stated adversary, DiLoCo training 5. The prototype alerts rather than throttles, and the authors call its software-only design "trivially spoofable" 5. Amodo reports building a DPU-based per-node rate limiter on 400G links, for weight security under a trusted controller 7. Lucid states that its pod-level cap is "still at the design stage and not yet implemented or red-teamed" 2. The mechanism's implementations, RAND secure inference data center (SIDC) design and AI 2040 inference-only verification stack, are proposed architectures at R1.
  • R3 not met: no production-grade cap or monitor for this use is publicly available, and no reliance on one by a party other than its developer is documented.

Confidence is low: the R2 evidence is one small, software-only prototype described in a blog post, and it monitors a limit without enforcing one.

Gaps to the next level
  • A production-grade cap or bandwidth monitor for this use, or reliance on one by a party other than its developer for a verification decision.
  • Enforcement outside the prover's control, such as a shaper at each pod uplink, built and tested on data-centre hardware (Lucid reports a proof of concept in development).
  • Measured, not only modelled, training inefficiency under a cap at realistic scale, including low-communication methods.
  • Red-teaming of cap bypass (for example through parallel scale-up switches) and of trust in the shaping devices.

Assessed 2026-10-08 against rubric v1.1.

Mechanism properties

Threat modelAdversarial prover
Adversarial evaluationAnalysis
Hardware neededRetrofit device
Prover cooperationRequired
ConfidentialityPreserving

Evidence

  • Designs. The AI 2040 plan, Lucid's design brief and the MIRI overview are designs and analysis 1 2 4. Lucid describes its design as "still at the design stage and not yet implemented or red-teamed" 2. It reports that its engineers are building a proof of concept for red-teaming at a partner cluster 2.
  • Monitoring prototype. MIRI's Technical Governance Team monitored inter-node traffic on a two-node Azure cluster with four A100 GPUs joined by one 100 Gb Ethernet link, and published its code 5. Fully sharded fine-tuning of Llama 3.1 8B and a mixture-of-experts model averaged 2–3 GB/s between the nodes, while LLM inference averaged about 10 kB/s 5. DiLoCo cut total cross-node traffic from about 10 TB to about 100 GB, but its regular spikes of hundreds of MB/s remained distinguishable from inference 5. The prototype raises alerts and does not throttle, and the authors call their software-only version "trivially spoofable", because the prover controls the whole measurement pipeline 5.
  • Node-level limiter. Amodo reports 390 Gbps unencrypted and 193 Gbps encrypted throughput between two DPUs on a 400G link 7. That demonstration targets weight security under a trusted controller, and Amodo has "not yet fully analysed resilience to a single compromised DPU" 7.
  • Operator egress limits. Anthropic reports that its ASL-3 security measures include egress bandwidth controls, which limit the rate of outbound traffic from environments that hold model weights 6.

Limitations

  • Low-communication training. DiLoCo matched fully synchronous training on 8 workers while communicating 500 times less 8. Rahman writes that this family theoretically allows large-scale training at under 100 Mbps 9. Lucid's bounds include DiLoCo- and SWARM-family methods, but it notes that extreme activation compression or modular paradigms could erode the margin 2.
  • Inference that crosses pod boundaries. Expert-parallel mixture-of-experts inference uses all-to-all communication between devices holding different experts, which can be a bottleneck 11. A cap calibrated for token-only ingress and egress may disrupt that serving architecture.
  • Low-traffic training. LoRA fine-tuning and reinforcement learning involve less inter-node traffic than ordinary training, and the MIRI authors leave whether a limit would stop them to future work 5. They judge that training for longer at lower throughput is not a major concern, because a reasonable limit would slow training by a prohibitive factor, such as 100x 5.
  • Routing control. If the operator can group pods freely behind routers, the bound falls to about 90–220x uncompressed and as low as about 25x with compression 2.
  • Hidden storage. Undeclared per-pod storage weakens the bound 2.
  • Scale-up bypass. In GB200 topologies, GPUs reach other nodes through NVSwitches without a NIC on the path, and compromising one or two parallel switches would bypass a limit 7.
  • Durability. Lucid positions the design as "a short-to-medium-term deterrent and Phase-1 milestone" 2.

Known flaws

Blockers

Search

Full search page