Mechanism · Sampled inference recomputation
Some inference optimizations are not covered
On this page
SignificantTheoretical argumentOpen
TOPLOC's authors state that it cannot detect speculative decoding in which a cheaper model does the decoding. They did not test whether it distinguishes types of key-value (KV) cache compression 7. DiFR was evaluated only on sampling from a single model. Its authors sketch an extension to one speculative-decoding algorithm but do not test it 2.