Implementation · Low-trust AI compute verification system overview

Inspector agents may be manipulable

On this page

← All known flaws

SignificantOpen questionOpen

Automated compliance screening with LLM-based inspector agents must resist prompt-injection attacks. Adversarially trained systems might hide malicious actions with steganography, which makes backdoor detection an open problem.

Sources: [1]

Search

Full search page