Implementation · Low-trust AI compute verification system overview
Inspector agents may be manipulable
On this page
SignificantOpen questionOpen
Automated compliance screening with LLM-based inspector agents must resist prompt-injection attacks. Adversarially trained systems might hide malicious actions with steganography, which makes backdoor detection an open problem.
Sources: [1]