FLOP accounting
Estimating or verifying how many floating-point operations a training run or other workload used, often to compare against a threshold in a rule.
FLOP accounting estimates or verifies the total number of floating-point operations (FLOP) that a training run or other workload performs, usually to compare it with a threshold set by a rule 1 2.
Shavit lists total training compute among the rules a verifier might enforce, noting that it has proven to be an indicator of model capabilities 1. US Executive Order 14110 required reporting for models trained with more than 10^26 operations 2 until its revocation in January 2025 8, and a proposed international agreement sets a prohibited threshold of 10^24 FLOP and a monitored threshold of 10^22 FLOP 3. Proposed ways to count or cap FLOP include:
- Hardware time. Shavit converts a FLOP threshold into accelerator-days by assuming that every accelerator runs at its full rate with perfect parallelization, a conservative assumption that gives the fewest accelerator-days a run of that size could occupy 1.
- Energy. Energy use can be converted into an approximate FLOP count 4.
- Telemetry. RAND lists estimating a workload's model FLOPs utilization (MFU) and physical signature, such as power, as a research problem 5, the aim of on-chip telemetry.
- On-chip budgets. Offline licensing ties chip use to a renewable licence carrying a compute budget 6, as in hardware performance throttling and licensing.
Training compute is only a high-level proxy for capability, and algorithmic progress means thresholds may need to change 2; one system overview argues that coarse totals such as FLOP counts are not enough and seeks evidence about individual workloads 7.
Related
Used in
- R1Hardware-enabled guarantees (flexHEG) and guarantee processors
- R1Hardware performance throttling and licensing
- R2On-chip telemetry from timing, memory and performance counters⚠
- R2Proof-of-learning and training-transcript verification⚠
- R1Proofs of useful work and resource exhaustion
- R2Zero-knowledge proofs of training constraints
- Compute stock is at most a declared amount
- There is no undeclared relevant compute
- A training run stayed within declared limits
Sources
- BY. Shavit (2023). What does it take to catch a Chinchilla? Verifying Rules on Large-Scale Neural Network Training via Compute Monitoring. arXiv. Source recordSupports: total training compute as a rule and an indicator of capabilities; a threshold of H FLOPs converted to chip-days using each chip's FLOPs per day at full, perfectly parallel use · §2.1; §3.2, Table 1
- BG. Sastry et al. (2024). Computing Power and the Governance of Artificial Intelligence. arXiv. Source recordSupports: EO 14110 reporting threshold of 10^26 operations; compute as a high-level proxy for capabilities; thresholds may need updating · thresholds; limitations
- BA. Scher et al. (2025). An International Agreement to Prevent the Premature Creation of Artificial Superintelligence. Machine Intelligence Research Institute. Source recordSupports: Strict Threshold 10^24 FLOP and Monitored Threshold 10^22 FLOP · §4
- BA. R. Wasil et al. (2024). Verification methods for international AI agreements. arXiv. Source recordSupports: energy estimates converted into an approximation of FLOPs · Energy monitoring
- BM. Baker et al. (2025). Verifying International Agreements on AI: Six Layers of Verification for Rules on Large-Scale AI Development and Deployment. RAND Corporation. Source recordSupports: estimating MFU and physical signature (e.g. power) as a research problem · Table 2, Appendix A.6
- BG. Kulp et al. (2024). Hardware-Enabled Governance Mechanisms: Developing Technical Solutions to Exempt Items Otherwise Classified Under Export Control Classification Numbers 3A090 and 4A090. RAND Corporation. Source recordSupports: offline licensing with a renewable licence carrying a compute budget · p. viii
- BN. Cankaya (2026). A System Overview for Near-Term, Low-Trust AI Compute Verification. Machine Intelligence Research Institute. Source recordSupports: coarse metrics like total FLOPs insufficient; per-workload evidence sought · verification goals
- AExecutive Office of the President (2025). Executive Order 14148: Initial Rescissions of Harmful Executive Orders and Actions. Federal Register, 90 FR 8237 (document 2025-01901, published 2025-01-28). Source recordSupports: revocation of EO 14110 on 20 January 2025 · Sec. 2(ggg)