{
  "schema_version": "1.4.0",
  "rubric_version": "1.1",
  "license": "CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)",
  "record": {
    "id": "G-0001",
    "slug": "cap-frontier-training",
    "title": "Cap frontier training",
    "aliases": [
      "training compute cap",
      "compute threshold",
      "FLOP threshold"
    ],
    "status": "draft",
    "last_reviewed": "2026-10-03",
    "review_interval_days": 90,
    "steward": null,
    "provenance": {
      "drafted_by": "ai",
      "reviewed_by": [
        "claude-review"
      ]
    },
    "risk_flags": [],
    "flags": [],
    "one_liner": "Keep every AI training run below an agreed amount of compute.",
    "summary": "A cap on frontier training sets an amount of compute that no AI training run may exceed. Shavit describes a hardware-monitoring framework whose main aim is to give governments high confidence that no actor uses large quantities of specialised chips for a training run that breaks agreed rules [[S-0029]]. A draft international agreement by Scher and colleagues prohibits training runs above 10^24 floating-point operations (FLOP) [[S-0063]]. Verifying a cap means checking the training runs a party declares, and showing that no large run happens on compute it has not declared.",
    "claims": [
      {
        "id": "C-0007",
        "title": "A training run stayed within declared limits",
        "url": "https://trustbutveri.fyi/claims/training-within-declared-limits/",
        "assessment": true,
        "relevance": "direct",
        "note": "The cap is a limit on each training run, so every declared run has to be shown to stay under it.",
        "sources": [
          "S-0029",
          "S-0063"
        ],
        "editorial": false
      },
      {
        "id": "C-0010",
        "title": "There is no undeclared relevant compute",
        "url": "https://trustbutveri.fyi/claims/no-undeclared-compute/",
        "assessment": true,
        "relevance": "direct",
        "note": "A run on compute that was never declared escapes every check on declared runs. Wasil and colleagues list unauthorised data centres, such as ones above an agreed capacity limit, as a violation alongside unauthorised training, and say methods are needed to detect them.",
        "sources": [
          "S-0062"
        ],
        "editorial": false
      },
      {
        "id": "C-0004",
        "title": "This compute runs inference, not training",
        "url": "https://trustbutveri.fyi/claims/inference-not-training/",
        "assessment": true,
        "relevance": "supporting",
        "note": "Compute declared as inference could be used to train. In the agreement proposed by Scher and colleagues, chip monitoring starts by telling inference on existing models from training of new ones.",
        "sources": [
          "S-0063"
        ],
        "editorial": false
      },
      {
        "id": "C-0001",
        "title": "Compute stock is at most a declared amount",
        "url": "https://trustbutveri.fyi/claims/compute-stock-is-bounded/",
        "assessment": true,
        "relevance": "supporting",
        "note": "Shavit's design monitors the chip supply chain so that no actor can avoid discovery by amassing a large quantity of untracked chips, which means accounting for the chips each party holds.",
        "sources": [
          "S-0029"
        ],
        "editorial": false
      },
      {
        "id": "C-0002",
        "title": "Chips are where they are declared to be",
        "url": "https://trustbutveri.fyi/claims/chips-are-where-declared/",
        "assessment": true,
        "relevance": "supporting",
        "note": "The agreement proposed by Scher and colleagues prohibits concentrations of chips above 16 H100-equivalents outside monitored facilities, so enforcing its thresholds includes knowing where chips are.",
        "sources": [
          "S-0063"
        ],
        "editorial": false
      }
    ],
    "outside": [
      {
        "label": "the choice of threshold",
        "text": "This map treats where to set a threshold, and when to change it, as a policy decision. Sastry and colleagues describe training compute as only a high-level proxy for a model's capabilities, and note that thresholds have to change over time as algorithms and hardware progress.",
        "sources": [
          "S-0053"
        ],
        "editorial": false
      }
    ],
    "sources": [
      {
        "source": "S-0029",
        "supports": "aim of the monitoring framework; weight snapshots, training records and supply-chain monitoring",
        "locator": "abstract"
      },
      {
        "source": "S-0063",
        "supports": "training limits as FLOP thresholds, verified by chip tracking and chip-use verification; 10^24 and 10^22 FLOP; concentrations above 16 H100-equivalents outside monitored facilities; inference versus training",
        "locator": "abstract; §4"
      },
      {
        "source": "S-0062",
        "supports": "unauthorised training and unauthorised data centres as the two kinds of violation; ten verification methods; need for methods that detect unauthorised data centres",
        "locator": "abstract; executive summary; What to verify"
      },
      {
        "source": "S-0053",
        "supports": "training compute as only a high-level proxy for capability; thresholds change with algorithmic and hardware progress",
        "locator": "§2.C"
      },
      {
        "source": "S-3542",
        "supports": "presumption of high-impact capabilities above 10^25 FLOP; duty to notify the Commission",
        "locator": "Arts 51, 52"
      }
    ],
    "order": 1,
    "type": "goal",
    "url": "https://trustbutveri.fyi/goals/cap-frontier-training/",
    "source_file": "content/goals/cap-frontier-training.md",
    "flags_all": [],
    "body_markdown": "## Proposals\n\n- **Shavit** analyses how governments could enforce rules on large training runs, and verify each other's compliance with them, by monitoring the computing hardware used for training [[S-0029]]. The design has three parts [[S-0029]]. Chips occasionally save snapshots of the weights in their memory. The trainer keeps enough information to prove how those weights were produced. The chip supply chain is monitored, so that no actor can avoid discovery by amassing a large quantity of untracked chips.\n- **Scher and colleagues** propose an international agreement that restricts the scale of AI training [[S-0063]]. The limits are FLOP thresholds, verified by tracking AI chips and verifying how they are used [[S-0063]]. Training runs above 10^24 FLOP are prohibited, and runs above 10^22 FLOP must be approved and monitored [[S-0063]].\n- **Wasil and colleagues** examine ten verification methods against two kinds of violation: unauthorised training, such as runs above a FLOP threshold, and unauthorised data centres [[S-0062]].\n\n## Thresholds in law\n\nThe EU AI Act uses a compute threshold to trigger duties. It presumes that a general-purpose AI model trained with more than 10^25 FLOP has high-impact capabilities, and its provider must notify the European Commission [[S-3542]].",
    "body_text": "Proposals - Shavit analyses how governments could enforce rules on large training runs, and verify each other's compliance with them, by monitoring the computing hardware used for training [S-0029]. The design has three parts [S-0029]. Chips occasionally save snapshots of the weights in their memory. The trainer keeps enough information to prove how those weights were produced. The chip supply chain is monitored, so that no actor can avoid discovery by amassing a large quantity of untracked chips. - Scher and colleagues propose an international agreement that restricts the scale of AI training [S-0063]. The limits are FLOP thresholds, verified by tracking AI chips and verifying how they are used [S-0063]. Training runs above 10^24 FLOP are prohibited, and runs above 10^22 FLOP must be approved and monitored [S-0063]. - Wasil and colleagues examine ten verification methods against two kinds of violation: unauthorised training, such as runs above a FLOP threshold, and unauthorised data centres [S-0062]. Thresholds in law The EU AI Act uses a compute threshold to trigger duties. It presumes that a general-purpose AI model trained with more than 10^25 FLOP has high-impact capabilities, and its provider must notify the European Commission [S-3542].",
    "referenced_by": []
  }
}