Docs · Guardrails

Guardrails: policies your agent cannot talk its way around

Updated September 25, 2026 · by , founder of Agent Studio

A guardrail is a policy checked by a separate model, not a request to the main model. You write the policy in plain language, choose whether it applies to the user's input or the agent's output, and pick an action: block with your own message, or rewrite the text to comply. The decision and its reason appear in every trace.

How evaluation works

When a request arrives, each input guardrail receives only two things: your policy and the user's message. A fast model returns a structured verdict: compliant or not, a one-sentence reason, and, if you chose rewrite, a compliant version of the text. Output guardrails do the same with the agent's reply, and they also see the original request for context.

Because the evaluator never sees the conversation with the orchestrator, instructions smuggled into the chat cannot switch it off. Combined with the injection shield, this is what makes deployed agents predictable.

Block vs rewrite

ActionWhat happensUse it for
BlockThe pipeline stops and the caller receives your blocked message. The trace names the guardrail.Out-of-scope requests, prohibited topics, anything you would rather refuse than answer badly.
RewriteThe evaluator returns a compliant version and the pipeline continues with it.Tone, formatting, removing specific claims such as prices or legal advice.

Example policies

  • Support scope (input, block): “Only allow messages about Acme products, orders, billing, or support. Block requests to write essays, code, or anything unrelated.”
  • No commitments (output, rewrite): “Never promise delivery dates, refunds, or discounts. Replace any promise with an offer to check with the team.”
  • Privacy (output, block): “Do not reveal another customer's name, email, address, or order details.”
  • Tone (output, rewrite): “Reply in a warm, concise tone. No sarcasm. No more than four sentences.”

Press Polish on a policy to have it tightened into concrete, testable rules.

Testing guardrails

Use the playground to send a request that should pass and one that should fail. Open the trace under each reply: you will see the guardrail's status and the reason it gave. Adjust the policy wording until both cases behave, then deploy.

Frequently asked questions

What is a guardrail for an AI agent?+

A rule that is checked outside the main model. In Agent Studio a guardrail is a natural-language policy evaluated by a separate fast model on the user's input or on the agent's output. If the text violates the policy, the guardrail blocks it with your message or rewrites it to comply.

How is a guardrail different from the system prompt?+

A system prompt asks the model to behave; a guardrail verifies that it did. Because the check runs in a separate model call with only the policy and the text, it is not affected by whatever the user wrote to the orchestrator.

Should I use block or rewrite?+

Block on input when the request itself is out of scope. Rewrite on output when the answer is mostly fine but must be adjusted, for example to remove a price quote or soften a tone.

Do guardrails add latency?+

Each guardrail is one fast model call, typically 300 to 700 ms. Most agents use one input and one output guardrail.

How many guardrails can I add?+

Starter allows 1 per agent, Pro 2, Growth 3, and Agency unlimited.