Mechanism · Sampled inference recomputation

Some inference optimizations are not covered

On this page

← All known flaws

SignificantTheoretical argumentOpen

TOPLOC's authors state that it cannot detect speculative decoding in which a cheaper model does the decoding. They did not test whether it distinguishes types of key-value (KV) cache compression 7. DiFR was evaluated only on sampling from a single model. Its authors sketch an extension to one speculative-decoding algorithm but do not test it 2.

Sources: [7] · [2]

Search

Full search page