AgentGuards

Frequently asked questions

Everything you need to know about how AgentGuards works, integrates, and is metered.

Where is the service hosted?
The website and API run on AWS (eu-north-1 region), served over HTTPS through an Application Load Balancer with ACM certificates. The machine-learning classifiers run on a dedicated server we operate, reached only by our API over TLS.
How do I integrate AgentGuards?
For Claude Code and OpenAI Codex it's one command: curl -fsSL https://agentguards.co/install.sh | sh (on Windows, irm https://agentguards.co/install.ps1 | iex). It opens your browser to approve the sign-in — no key to copy — then you restart your agent. Gemini CLI, Copilot CLI and OpenCode have a short guided setup, and your own app can call the REST API or the Gateway. See what works where →
What checks does AgentGuards run?
Every prompt goes through up to 10 layered checks: prompt injection, jailbreak detection, PII detection, secret detection, data exfiltration, toxicity, restricted topics, web-content injection, an LLM semantic check, and a PromptGuard ML classifier — checks short-circuit on the first confirmed threat. On top of that, AgentGuards scores every shell command your agent tries to run and blocks destructive ones (like rm -rf /), screens the web pages it fetches for hidden instructions, scans every file it writes or edits for hardcoded secrets and vulnerable code (SAST via semgrep + gitleaks), and checks the packages it installs for known vulnerabilities (grype). Web-page, code and dependency scanning are on by default on every plan.
Does AgentGuards store my prompts?
No. We scan prompt content in memory and discard it immediately. Only metadata — request counts, token counts, blocked event types — is persisted for your usage dashboard.
Which LLMs and frameworks are supported?
AgentGuards is model-agnostic. It sits in front of any LLM call as an HTTP guardrail. Native integrations exist for Claude Code, OpenAI Codex, Gemini CLI, GitHub Copilot (CLI and VS Code) and OpenCode.
What happens when a threat is detected?
The request is blocked before it reaches the model. You receive a JSON response with decision: "block" and a per-check breakdown showing which check triggered and why. Your LLM is never called.
Can I configure which checks run?
Yes. Every check can be toggled on or off per tenant from your dashboard. Individual and higher plans can also edit custom detection patterns.
What counts as a request?
Each guardrail evaluation — an input check, output validation, action authorization, policy evaluation, or gateway completion — is one metered request.
Do AgentGuards requests use up my Claude Code (or other LLM) credits?
Depends on which integration you use. With the hook-based integration (Claude Code, Gemini CLI, Codex hooks), our check runs as a plain HTTP call outside the model's context — Claude never sees it, so it burns zero of your model credits. With the MCP server integration, calling our tools costs a small amount of Claude context/tokens, the same as any MCP tool call, since the model has to load the tool schema and read the JSON result — that's what shows up under "MCP servers" in Claude Code's /usage breakdown. Either way, the security check itself never invokes an LLM on our end (it's regex plus a self-hosted classifier), and it only counts once against your AgentGuards request quota — a separate meter from your LLM credits. The one exception is the optional LLM-judge check, which — if you enable it — runs on your own OpenAI key, so any token cost there stays on your account, not ours.
Does AgentGuards work in Claude Desktop's Cowork tab?
Not today — we tested it directly rather than assumed. A plugin's hooks are supposed to run on every message regardless of setup, and in Cowork sessions they don't, even with a real API key configured and Desktop fully restarted. The terminal, Claude Desktop's Code tab, and cloud sessions started at claude.ai/code all work normally. See the full breakdown →
Is the free plan really free? Do I need a credit card?
Yes — no credit card to start. The Free plan is a one-time credit of 1,000 requests with 1 API key: enough to wire AgentGuards into your tools and run it against real traffic. You have 30 days from signing up to spend it, and it doesn't refill each month. Once it's used up or the 30 days are out, you choose what happens next: upgrade, keep working without screening, or uninstall.
What happens when I hit my limit?
On a paid plan, requests beyond your monthly quota pause until the next cycle or an upgrade. On the free plan there's no next cycle — the credit is one-time. We email you when you've used 80% of it, and once it's gone you choose what happens next: upgrade, keep working without screening (with one reminder a day), or uninstall. Until you choose, requests are blocked. Your dashboard always shows usage and what's left.
Can I bring my own model key?
Yes. The optional LLM-judge check runs on your own OpenAI key, so that spend stays on your account.
Do you offer on-prem?
Enterprise can deploy in your own VPC or on-prem, with SLA, SSO, and a security review.