AI Grimoire

Adversarial ML / agents

Agents

An agent that reads untrusted content and then acts on it has no reliable boundary between instruction and data. This is less a solved threat than a taxonomy of ways the boundary fails.

1 entry.

Entries

06.05.1
Prompt Injection Taxonomystandard
Direct, indirect and tool-mediated instruction hijack surfaces.