Cap frontier training

Draft

A cap on frontier training sets an amount of compute that no AI training run may exceed.

Shavit describes a hardware-monitoring framework whose main aim is to give governments high confidence that no actor uses large quantities of specialised chips for a training run that breaks agreed rules 1. A draft international agreement by Scher and colleagues prohibits training runs above 10^24 floating-point operations (FLOP) 2. Verifying a cap means checking the training runs a party declares, and showing that no large run happens on compute it has not declared.

On this page

Proposals

  • Shavit analyses how governments could enforce rules on large training runs, and verify each other's compliance with them, by monitoring the computing hardware used for training 1. The design has three parts 1. Chips occasionally save snapshots of the weights in their memory. The trainer keeps enough information to prove how those weights were produced. The chip supply chain is monitored, so that no actor can avoid discovery by amassing a large quantity of untracked chips.
  • Scher and colleagues propose an international agreement that restricts the scale of AI training 2. The limits are FLOP thresholds, verified by tracking AI chips and verifying how they are used 2. Training runs above 10^24 FLOP are prohibited, and runs above 10^22 FLOP must be approved and monitored 2.
  • Wasil and colleagues examine ten verification methods against two kinds of violation: unauthorised training, such as runs above a FLOP threshold, and unauthorised data centres 3.

Thresholds in law

The EU AI Act uses a compute threshold to trigger duties. It presumes that a general-purpose AI model trained with more than 10^25 FLOP has high-impact capabilities, and its provider must notify the European Commission 5.

Claims

What would have to be verified to check this goal. The two levels are the editors' judgment. How goals link to claims

Direct

  • The cap is a limit on each training run, so every declared run has to be shown to stay under it. 1 2
  • A run on compute that was never declared escapes every check on declared runs. Wasil and colleagues list unauthorised data centres, such as ones above an agreed capacity limit, as a violation alongside unauthorised training, and say methods are needed to detect them. 3

Supporting

  • Compute declared as inference could be used to train. In the agreement proposed by Scher and colleagues, chip monitoring starts by telling inference on existing models from training of new ones. 2
  • Shavit's design monitors the chip supply chain so that no actor can avoid discovery by amassing a large quantity of untracked chips, which means accounting for the chips each party holds. 1
  • The agreement proposed by Scher and colleagues prohibits concentrations of chips above 16 H100-equivalents outside monitored facilities, so enforcing its thresholds includes knowing where chips are. 2

Start a proposal from these claims →

Outside this map

What the goal also needs that no record on this map covers.

  • The choice of thresholdThis map treats where to set a threshold, and when to change it, as a policy decision. Sastry and colleagues describe training compute as only a high-level proxy for a model's capabilities, and note that thresholds have to change over time as algorithms and hardware progress. 4

Sources

  1. BY. Shavit (2023). What does it take to catch a Chinchilla? Verifying Rules on Large-Scale Neural Network Training via Compute Monitoring. arXiv. Source recordSupports: aim of the monitoring framework; weight snapshots, training records and supply-chain monitoring · abstract
  2. BA. Scher et al. (2025). An International Agreement to Prevent the Premature Creation of Artificial Superintelligence. Machine Intelligence Research Institute. Source recordSupports: training limits as FLOP thresholds, verified by chip tracking and chip-use verification; 10^24 and 10^22 FLOP; concentrations above 16 H100-equivalents outside monitored facilities; inference versus training · abstract; §4
  3. BA. R. Wasil et al. (2024). Verification methods for international AI agreements. arXiv. Source recordSupports: unauthorised training and unauthorised data centres as the two kinds of violation; ten verification methods; need for methods that detect unauthorised data centres · abstract; executive summary; What to verify
  4. BG. Sastry et al. (2024). Computing Power and the Governance of Artificial Intelligence. arXiv. Source recordSupports: training compute as only a high-level proxy for capability; thresholds change with algorithmic and hardware progress · §2.C
  5. AEuropean Parliament & Council of the European Union (2024). Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Official Journal of the European Union, OJ L, 2024/1689. Source recordSupports: presumption of high-impact capabilities above 10^25 FLOP; duty to notify the Commission · Arts 51, 52

Search

Full search page