How Hooks Work
AgentGuards hooks sit between you and your AI coding agent. Before Claude responds to a message, and before it runs any shell command, the hooks check whether the action is safe. Nothing to call manually — it all happens in the background.
UserPromptSubmit
Every message you type is scanned before Claude ever sees it. If the message contains a jailbreak attempt, prompt injection, PII, or other policy violations, it is blocked and Claude never responds.
Gets blocked
- Attempts to override Claude's instructions
- Requests to exfiltrate secrets or credentials
- Messages containing PII (emails, passwords, card numbers)
- Jailbreak patterns
Passes through
- Normal coding requests
- Questions, explanations, code reviews
- Anything that doesn't trip a guardrail
If a message is blocked you'll see an AgentGuards explanation. Rephrase and try again.
PreToolUse
When Claude wants to run a shell command, AgentGuards evaluates it first and returns one of three decisions.
| Decision | What happens |
|---|---|
| Allow | Command runs immediately, no prompt |
| Deny | Command is hard-blocked — it never runs |
| Ask | Claude Code pauses and asks you to approve it |
Always allowed (safe baseline)
cat, head, tail, grep, find, lsgit status, git log, git diff, git shownpm install, pip install, cargo build, makepytest, jest, go testps, top, df, duAlways denied (destructive)
rm -rf (non-temp paths)dd, mkfs, fdiskDROP TABLE, DROP DATABASECommands overwriting system filesRequires your approval (elevated risk)
- Writing to system directories
- Modifying config files outside the project
- Network requests to external hosts
- Elevated-risk commands that aren't clearly destructive
Dependency vulnerability scanning (optional, paid)
When the command is a package install with an exact-pinned version — pip install requests==2.19.0, npm i lodash@4.17.4 (also uv, poetry, yarn, pnpm) — AgentGuards can look the package up against the grype vulnerability database and deny the install if it has a known CVE at or above your severity threshold, naming the CVE and the fixed version. Below-threshold findings warn; unpinned installs are left alone. The threshold defaults to critical — lower it to high, medium or low for a stricter gate. It is on by default on every plan; toggle Dependency scan (grype) on the dashboard Guardrail checks page to turn it on or off. If the scanner cannot answer in time the install proceeds and is recorded as unscanned — never as clean.
PostToolUse
After a command runs (meaning it was allowed or you approved it), AgentGuards remembers that you approved that tool for this session. Next time Claude uses the same binary, you won't be asked again.
This memory is per-session and clears automatically when the session ends.
PostToolUse also scans content pulled by the built-in WebFetch and WebSearch tools, and scans Write/Edit/MultiEdit file writes for SAST findings and secrets (matcher Bash|WebFetch|WebSearch|Write|Edit|MultiEdit). Fetched content is checked with use_case="web_fetch"; if AgentGuards flags it — for example an indirect prompt injection planted in a webpage — the result is withheld from Claude so it never acts on the poisoned content.
When AgentGuards is unreachable
By default the hook is fail-closed: if the AgentGuards service can't be reached, prompts and commands are blocked until the connection is restored.
Set AGENTGUARDS_FAIL_OPEN=true to switch to fail-open mode, where unreachable means allow-through. Use this only if uptime matters more than enforcement.
Summary
| Hook | Fires when | Outcome |
|---|---|---|
| UserPromptSubmit | You submit any message | Block or allow the prompt |
| PreToolUse | Claude wants to run a shell command | Allow, deny, or ask you to approve |
| PostToolUse | A command finishes, or a web fetch/search returns | Remember the approval; scan fetched web content and withhold it if flagged |