AgentGuards

How Hooks Work

AgentGuards hooks sit between you and your AI coding agent. Before Claude responds to a message, and before it runs any shell command, the hooks check whether the action is safe. Nothing to call manually — it all happens in the background.

1

UserPromptSubmit

Every message you type is scanned before Claude ever sees it. If the message contains a jailbreak attempt, prompt injection, PII, or other policy violations, it is blocked and Claude never responds.

Gets blocked

  • Attempts to override Claude's instructions
  • Requests to exfiltrate secrets or credentials
  • Messages containing PII (emails, passwords, card numbers)
  • Jailbreak patterns

Passes through

  • Normal coding requests
  • Questions, explanations, code reviews
  • Anything that doesn't trip a guardrail

If a message is blocked you'll see an AgentGuards explanation. Rephrase and try again.

2

PreToolUse

When Claude wants to run a shell command, AgentGuards evaluates it first and returns one of three decisions.

DecisionWhat happens
AllowCommand runs immediately, no prompt
DenyCommand is hard-blocked — it never runs
AskClaude Code pauses and asks you to approve it

Always allowed (safe baseline)

Reading filescat, head, tail, grep, find, ls
Git read operationsgit status, git log, git diff, git show
Build & installnpm install, pip install, cargo build, make
Running testspytest, jest, go test
Process & disk infops, top, df, du

Always denied (destructive)

Recursive deletionrm -rf (non-temp paths)
Disk-level operationsdd, mkfs, fdisk
Database dropsDROP TABLE, DROP DATABASE
System file writesCommands overwriting system files

Requires your approval (elevated risk)

  • Writing to system directories
  • Modifying config files outside the project
  • Network requests to external hosts
  • Elevated-risk commands that aren't clearly destructive

Dependency vulnerability scanning (optional, paid)

When the command is a package install with an exact-pinned version — pip install requests==2.19.0, npm i lodash@4.17.4 (also uv, poetry, yarn, pnpm) — AgentGuards can look the package up against the grype vulnerability database and deny the install if it has a known CVE at or above your severity threshold, naming the CVE and the fixed version. Below-threshold findings warn; unpinned installs are left alone. The threshold defaults to critical — lower it to high, medium or low for a stricter gate. It is on by default on every plan; toggle Dependency scan (grype) on the dashboard Guardrail checks page to turn it on or off. If the scanner cannot answer in time the install proceeds and is recorded as unscanned — never as clean.

3

PostToolUse

After a command runs (meaning it was allowed or you approved it), AgentGuards remembers that you approved that tool for this session. Next time Claude uses the same binary, you won't be asked again.

This memory is per-session and clears automatically when the session ends.

PostToolUse also scans content pulled by the built-in WebFetch and WebSearch tools, and scans Write/Edit/MultiEdit file writes for SAST findings and secrets (matcher Bash|WebFetch|WebSearch|Write|Edit|MultiEdit). Fetched content is checked with use_case="web_fetch"; if AgentGuards flags it — for example an indirect prompt injection planted in a webpage — the result is withheld from Claude so it never acts on the poisoned content.

When AgentGuards is unreachable

By default the hook is fail-closed: if the AgentGuards service can't be reached, prompts and commands are blocked until the connection is restored.

Set AGENTGUARDS_FAIL_OPEN=true to switch to fail-open mode, where unreachable means allow-through. Use this only if uptime matters more than enforcement.

Summary

HookFires whenOutcome
UserPromptSubmitYou submit any messageBlock or allow the prompt
PreToolUseClaude wants to run a shell commandAllow, deny, or ask you to approve
PostToolUseA command finishes, or a web fetch/search returnsRemember the approval; scan fetched web content and withhold it if flagged

Related