Glossary

FLOP accounting

Estimating or verifying how many floating-point operations a training run or other workload used, often to compare against a threshold in a rule.

FLOP accounting estimates or verifies the total number of floating-point operations (FLOP) that a training run or other workload performs, usually to compare it with a threshold set by a rule 1 2.

Shavit lists total training compute among the rules a verifier might enforce, noting that it has proven to be an indicator of model capabilities 1. US Executive Order 14110 required reporting for models trained with more than 10^26 operations 2 until its revocation in January 2025 8, and a proposed international agreement sets a prohibited threshold of 10^24 FLOP and a monitored threshold of 10^22 FLOP 3. Proposed ways to count or cap FLOP include:

  • Hardware time. Shavit converts a FLOP threshold into accelerator-days by assuming that every accelerator runs at its full rate with perfect parallelization, a conservative assumption that gives the fewest accelerator-days a run of that size could occupy 1.
  • Energy. Energy use can be converted into an approximate FLOP count 4.
  • Telemetry. RAND lists estimating a workload's model FLOPs utilization (MFU) and physical signature, such as power, as a research problem 5, the aim of on-chip telemetry.
  • On-chip budgets. Offline licensing ties chip use to a renewable licence carrying a compute budget 6, as in hardware performance throttling and licensing.

Training compute is only a high-level proxy for capability, and algorithmic progress means thresholds may need to change 2; one system overview argues that coarse totals such as FLOP counts are not enough and seeks evidence about individual workloads 7.

Related

Used in

Sources

  1. BY. Shavit (2023). What does it take to catch a Chinchilla? Verifying Rules on Large-Scale Neural Network Training via Compute Monitoring. arXiv. Source recordSupports: total training compute as a rule and an indicator of capabilities; a threshold of H FLOPs converted to chip-days using each chip's FLOPs per day at full, perfectly parallel use · §2.1; §3.2, Table 1
  2. BG. Sastry et al. (2024). Computing Power and the Governance of Artificial Intelligence. arXiv. Source recordSupports: EO 14110 reporting threshold of 10^26 operations; compute as a high-level proxy for capabilities; thresholds may need updating · thresholds; limitations
  3. BA. Scher et al. (2025). An International Agreement to Prevent the Premature Creation of Artificial Superintelligence. Machine Intelligence Research Institute. Source recordSupports: Strict Threshold 10^24 FLOP and Monitored Threshold 10^22 FLOP · §4
  4. BA. R. Wasil et al. (2024). Verification methods for international AI agreements. arXiv. Source recordSupports: energy estimates converted into an approximation of FLOPs · Energy monitoring
  5. BM. Baker et al. (2025). Verifying International Agreements on AI: Six Layers of Verification for Rules on Large-Scale AI Development and Deployment. RAND Corporation. Source recordSupports: estimating MFU and physical signature (e.g. power) as a research problem · Table 2, Appendix A.6
  6. BG. Kulp et al. (2024). Hardware-Enabled Governance Mechanisms: Developing Technical Solutions to Exempt Items Otherwise Classified Under Export Control Classification Numbers 3A090 and 4A090. RAND Corporation. Source recordSupports: offline licensing with a renewable licence carrying a compute budget · p. viii
  7. BN. Cankaya (2026). A System Overview for Near-Term, Low-Trust AI Compute Verification. Machine Intelligence Research Institute. Source recordSupports: coarse metrics like total FLOPs insufficient; per-workload evidence sought · verification goals
  8. AExecutive Office of the President (2025). Executive Order 14148: Initial Rescissions of Harmful Executive Orders and Actions. Federal Register, 90 FR 8237 (document 2025-01901, published 2025-01-28). Source recordSupports: revocation of EO 14110 on 20 January 2025 · Sec. 2(ggg)