Gateway API
Route every LLM call through AgentGuards: input guardrails run before the request reaches the model, and output validation lets you inspect the response before your users see it. One API key. No infrastructure to manage.
How it works
Quota: each endpoint counts separately against your request quota. Calling both /v1/gateway/complete and /v1/outputs/validate for one turn uses 2 requests, not 1 — plan your usage accordingly, especially against the free plan's one-time 1k credit. You only need the second call to screen output you produced elsewhere: /v1/gateway/complete already validates its own response before returning it.
Authentication
Every request requires your API token as the X-API-Key header. Get it from Dashboard → API Keys. The tenant is derived from the token automatically — no X-Tenant-ID header needed.
You also need a provider key. The Gateway calls the model with your OpenAI or Anthropic key — add one in Dashboard → Settings before your first call, or requests are refused with a missing_provider_api_key error.
POST /v1/gateway/complete
Runs input guardrails, then proxies the request to the configured LLM provider and returns the response.
| Field | Type | Required | Description |
|---|---|---|---|
| messages | array | yes | OpenAI-format message array |
| model_provider | string | no | openai · anthropic · gemini (defaults to your tenant config) |
| model_name | string | no | e.g. gpt-4o-mini, claude-3-5-haiku |
| use_case | string | no | Label for audit logs (e.g. customer-support) |
| stream | boolean | no | SSE streaming — OpenAI only, default false |
| metadata | object | no | Arbitrary key/value passed through to audit logs |
Response
correlation_id, tenant_id, model, output — where output is the raw OpenAI-format chat completion object (access the text at output.choices[0].message.content).
curl -X POST https://prod.agentguards.co/v1/gateway/complete \
-H "X-API-Key: ag_your_token_here" \
-H "Content-Type: application/json" \
-d '{
"messages": [
{ "role": "system", "content": "You are a concise assistant." },
{ "role": "user", "content": "Summarise the OWASP top 10." }
],
"model_provider": "openai",
"model_name": "gpt-4o-mini"
}'POST /v1/outputs/validate
Runs output guardrails on the model's response. Call this after getting the LLM output and before returning it to users.
| Field | Type | Required | Description |
|---|---|---|---|
| output_text | string | yes | The LLM response to validate |
| context_text | string | no | Retrieved context (for RAG hallucination checks) |
| include_quality_risk_score | boolean | no | Adds quality_risk_score to response metadata |
Decision values
| Decision | Meaning |
|---|---|
| pass | Output is clean — safe to return to the user |
| repair | Minor issues found — consider cleaning before returning |
| reject | Output violates policy — do not return to the user |
| escalate | High-confidence violation — log and alert |
curl -X POST https://prod.agentguards.co/v1/outputs/validate \
-H "X-API-Key: ag_your_token_here" \
-H "Content-Type: application/json" \
-d '{
"output_text": "The OWASP Top 10 covers injection, broken auth ...",
"include_quality_risk_score": true
}'Node.js (Express)
Full end-to-end integration — input guardrails, LLM call, and output validation in a single route handler.
import express from "express";
const app = express();
app.use(express.json());
const AG_URL = process.env.AGENTGUARD_URL || "https://prod.agentguards.co";
const AG_APIKEY = process.env.AGENTGUARD_API_KEY || "";
const HEADERS = {
"Content-Type": "application/json",
"X-API-Key": AG_APIKEY,
};
app.post("/chat", async (req, res) => {
try {
const userText = String(req.body?.message ?? "");
// 1. Route through the gateway (input guardrails + LLM call)
const gwRes = await fetch(AG_URL + "/v1/gateway/complete", {
method: "POST",
headers: HEADERS,
body: JSON.stringify({
messages: [
{ role: "system", content: "You are a concise assistant." },
{ role: "user", content: userText },
],
use_case: "customer-support",
model_provider: "openai",
model_name: "gpt-4o-mini",
stream: false,
metadata: { source: "my-node-api" },
}),
});
if (!gwRes.ok) {
const err = await gwRes.text();
return res.status(gwRes.status).json({ error: "gateway_failed", detail: err });
}
const gwData = await gwRes.json();
const message = gwData.output?.choices?.[0]?.message?.content ?? "";
// 2. Validate the output before returning it to the user
const valRes = await fetch(AG_URL + "/v1/outputs/validate", {
method: "POST",
headers: HEADERS,
body: JSON.stringify({
output_text: message,
include_quality_risk_score: true,
}),
});
const validation = await valRes.json();
if (validation.decision === "reject" || validation.decision === "escalate") {
return res.status(422).json({
error: "output_blocked",
decision: validation.decision,
checks: validation.checks,
});
}
return res.json({
correlation_id: gwData.correlation_id,
model: gwData.model,
message,
quality: validation.metadata,
});
} catch (e) {
return res.status(500).json({ error: "internal_error", detail: String(e) });
}
});
app.listen(3000, () => console.log("Server on :3000"));Python (httpx)
import os
import httpx
AG_URL = os.environ.get("AGENTGUARD_URL", "https://prod.agentguards.co")
AG_APIKEY = os.environ.get("AGENTGUARD_API_KEY", "")
HEADERS = {
"Content-Type": "application/json",
"X-API-Key": AG_APIKEY,
}
def chat(user_text: str) -> dict:
# 1. Gateway call — input guardrails + LLM
gw = httpx.post(
f"{AG_URL}/v1/gateway/complete",
headers=HEADERS,
json={
"messages": [
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": user_text},
],
"model_provider": "openai",
"model_name": "gpt-4o-mini",
"stream": False,
},
timeout=30,
)
gw.raise_for_status()
gw_data = gw.json()
message = gw_data["output"]["choices"][0]["message"]["content"]
# 2. Output validation
val = httpx.post(
f"{AG_URL}/v1/outputs/validate",
headers=HEADERS,
json={
"output_text": message,
"include_quality_risk_score": True,
},
timeout=10,
)
val.raise_for_status()
validation = val.json()
if validation["decision"] in ("reject", "escalate"):
raise ValueError(f"Output blocked: {validation['decision']}")
return {
"correlation_id": gw_data["correlation_id"],
"model": gw_data["model"],
"message": message,
"quality": validation.get("metadata", {}),
}Streaming
Set stream: true with model_provider: "openai" to receive an SSE stream. Reassemble the full text client-side, then call POST /v1/outputs/validate on the complete response before rendering it. Streaming is not supported for Anthropic or Gemini providers.