Mechanism · On-chip telemetry from timing, memory and performance counters

Adversarially disguised fine-tuning partly evades classification

On this page

← All known flaws

SignificantDemonstrated attackOpen

Across 20 evasion strategy families in five monitor-evader rounds, the classifier's accuracy against the most challenging disguised workloads held out from training was 43–87%. White-box LoRA fine-tuning was the only evasion family detected less than half the time. The evaluation covered single nodes, PyTorch-level evasion and NVIDIA hardware 3.

Sources: [3]

Search

Full search page