Glossary · Agent architecture

Guardrails (LLM)

Checks that run around a language model or agent to block, change or flag unsafe inputs, outputs and tool calls before they cause harm.

Guardrails are checks that run around a language model or agent, screening its inputs, outputs and actions, and blocking, changing or flagging anything that breaks a rule.

Where they run.

  • Input: before the model processes a request, for example to reject off-topic or malicious prompts.
  • Output: before a response reaches the user or the next system.
  • Tool calls: before and after a tool executes, to check arguments and results.
  • Retrieval and dialog: NVIDIA NeMo Guardrails, an open-source toolkit, also defines rails for retrieved content and for conversation flow, written in its Colang language.

One implementation. The OpenAI Agents SDK has input guardrails, which apply to the first agent in a run, output guardrails for the final agent, and tool guardrails around each function tool. A guardrail returns a result with a tripwire. When the tripwire triggers, the SDK raises an exception and halts the run. Input guardrails run in parallel with the agent by default, which saves time, or in blocking mode, so a rejected request never reaches the model and spends no tokens. Google ADK offers callbacks at set points in an agent’s run for the same kind of checks.

What they can and cannot do. Many guardrails are themselves model-based classifiers, so they lower risk without removing it. The OWASP Top 10 for Agentic Applications (December 2025) lists risks such as agent goal hijack, tool misuse, and identity and privilege abuse, which a filter on text alone may miss. Deterministic limits are the complement: tools with the least privilege they need, spending caps, allow-lists, and human approval for consequential steps.

Between agents. When input comes from another organization’s agent, it is untrusted content like any other. Messages, data parts and artifacts can all carry instructions aimed at your model. Guardrails on inbound tasks work alongside identity checks on who sent them and authorization checks on what they may ask for.

Neighbouring terms. Prompt injection is the main attack guardrails screen for. Human-in-the-loop asks a person where a guardrail decides automatically.

Sources

  1. OpenAI Agents SDK documentation: Guardrails (accessed )
  2. NVIDIA NeMo Guardrails documentation (accessed )
  3. OWASP GenAI Security Project: OWASP Top 10 for Agentic Applications (9 December 2025) (accessed )
  4. Agent Development Kit documentation: Technical overview (callbacks) (accessed )