What AgentGuards Checks?
Your team's AI coding assistant can now do a lot on its own. Tools like Claude Code, OpenAI Codex, GitHub Copilot CLI and Gemini CLI read web pages, run commands and edit code, right on a developer's laptop. That's why they save so much time. It also means they can be tricked, or simply get something wrong, with real access to your machines.
You don't need to understand how these tools work inside to see the risk. Picture a fast, eager new hire who does exactly what they read, including a note someone slipped into a document. That's roughly where things stand.
AgentGuards sits next to the assistant and checks what goes in and what comes out. Here's what it looks at, in plain terms.
What AgentGuards checks
1. Attempts to hijack the assistant. Some text is written to take control of an AI. It might say "ignore your previous instructions and do this instead." The industry calls this prompt injection. Think of a forged memo from "the boss." We screen for it using pattern rules plus a trained classifier, which is a small model built to spot this kind of text.
2. Hidden instructions in web pages. When the assistant reads a web page, that page can carry instructions a person would never see: white text on a white background, or notes tucked into the page's code. We strip that hidden text out and hold back pages that are trying to take over the assistant. This is on by default for hosted accounts.
3. Passwords and keys that shouldn't travel. Developers handle secrets all day: cloud passwords, API keys (the logins one piece of software uses to talk to another), database addresses. We catch common formats before they reach the AI. We also check outgoing commands, so a key can't quietly be sent to an outside website. It's like a mailroom that won't let an envelope out with the office keys inside.
4. Personal data. Email addresses, phone numbers, card numbers and similar details get flagged before they go to the AI model.
5. Risky commands. Everyday commands run untouched. Risky ones, like running something with administrator rights, wait for a person to say yes. Once approved, the answer is remembered for that session, so nobody gets asked twice. Clearly destructive ones, like wiping someone's home folder, are stopped outright.
6. Code the assistant writes. Every file it creates or edits is scanned for security mistakes and for passwords written straight into the code. Under the hood this uses two well-known open-source scanners, Semgrep and Gitleaks. It's on by default.
7. Software packages it installs. Modern software is built from thousands of ready-made parts. When the assistant installs a specific version of one, we check it against a database of known security holes and stop the install if the problem is serious. This is also on by default.
8. A record of every decision. Your dashboard shows what was checked, what was allowed and what was stopped. We don't keep the text of your prompts. They're scanned in memory and thrown away, and only counts and event types are stored.
Exactly what's covered differs a little from one tool to the next. The compatibility page shows it tool by tool.
How it's different
It doesn't depend on the AI remembering
Some safety tools are options the AI can choose to use, which means it can also skip them. AgentGuards runs as hooks: automatic checkpoints the coding tool itself triggers at fixed moments, like a turnstile every visitor passes through. The AI can't skip it, and nobody has to remember to switch it on.
A few other things set it apart:
- It doesn't use up your AI budget. The hook checks happen outside the AI's conversation, so they don't spend the tokens (the units AI providers bill by) you pay for.
- One setup, several tools. It works with Claude Code, Codex, Copilot CLI, Gemini CLI and OpenCode. If your team uses a mix, you manage the rules in one dashboard.
- Setup takes a single command. For Claude Code and Codex, one line in the terminal installs everything and signs you in through your browser. Nobody has to copy and paste a key. The other tools have a short guided setup in the dashboard.
- Every check is on every plan, including free. Plans differ only in how many requests you get, how many keys you can create, and the level of support. See plans.
- You stay in control. Any check can be turned on or off, and you can add your own rules from the dashboard.
- When something is stopped, it says why. In plain English, so a developer can decide quickly instead of digging through cryptic errors.
For companies that can't send any data outside their own cloud, there's also a self-hosted version that runs in your own AWS account.
What you get out of it
- Fewer nasty surprises. A leaked key or a hijacked assistant is the kind of incident that eats a week. AgentGuards is built to catch the common ones before they happen.
- A trail you can show someone. When a customer, auditor or your own team asks "what did the AI actually do?", there's an answer in the dashboard.
- Your team keeps its speed. Routine work isn't interrupted. Developers only hear from AgentGuards when something actually needs a human.
- Less worry for whoever is responsible. If you're the one who signed off on the team using AI agents, you're no longer just hoping nothing goes wrong.
Try it
It takes about two minutes, and you don't need a credit card. On macOS or Linux, a developer runs:
curl -fsSL https://agentguards.co/install.sh | sh
On Windows (PowerShell):
irm https://agentguards.co/install.ps1 | iex
It finds Claude Code and Codex on the machine, installs the protection, and opens the browser to sign in. For other tools, or a step-by-step walkthrough, start at getting started.