What AgentWarden cannot do
Last updated 18 September 2026
Security software earns trust by being honest about its edges. These are ours.
It cannot stop everything.
A hijacked session is refused or diverted before a sensitive action runs. But a process already mid-flight, or an agent with no hook support, is detected rather than stopped. Your agent’s own sandbox is still the strongest protection, and AgentWarden shows you whether it is on.
Depth depends on the agent.
Claude Code and Codex are read in full detail, because their session logs are read directly and every step is judged against what you asked. About thirty other agents are recognised running: the commands they start, their skills and their settings are checked, but they are not yet read line by line.
Patterns can be evaded.
Rules are patterns, and a new technique can slip past one. The decoys, and the “read hidden instructions, then acted” link, exist to catch what individual rules miss. The daily research loop turns newly published attacks into signed rule updates, usually the same day.
It is not enterprise endpoint security.
AgentWarden is strong personal-grade protection for the computer it runs on, with no kernel-level enforcement and no independent audit yet. We would rather you knew that now.