Generative AI guardrails
Amazon Bedrock Guardrails policies, the OWASP Top 10 for LLM Applications mapped to AWS controls, least privilege for agents, preventing data leakage, and logging model invocations.
Exam tasks: 3.2 (protections and guardrails for generative AI applications, applying the OWASP Top 10 for LLM Applications)
The decision: which risk sits in the model's input, which in its output, and which in what the application lets the model do, and which AWS control handles each without trusting the model to police itself?
Where controls sit in a GenAI request
The model is not a security boundary. Anything the model can reach, a clever prompt can ask it to reach. Put authorization in IAM, the tools and the data layer, and use guardrails to filter what goes in and out.
Amazon Bedrock Guardrails
| Policy | Checks | Typical action |
|---|---|---|
| Content filters | Hate, insults, sexual, violence, misconduct in text and images, with a strength per category | Block |
| Prompt attack (a content filter category) | Jailbreaks, prompt injection and prompt leakage attempts in the input | Block |
| Denied topics | Topics you describe in natural language, such as investment advice or a competitor's products | Block |
| Word filters | Exact words and phrases, plus a managed profanity list | Block |
| Sensitive information filters | Built-in PII types (email, card numbers, national IDs) and your own regex patterns | Block or mask (anonymize), separately for input and output |
| Contextual grounding checks | Whether the answer is supported by the source passages (grounding) and answers the question (relevance), against thresholds you set | Block or flag hallucinated RAG answers |
| Automated Reasoning checks | Whether an answer is consistent with formal rules you define from a policy document | Validate and explain |
- A guardrail runs on the input before the model and on the output after it. You choose the message returned when something is blocked.
- Attach it to
InvokeModel,Converse, Bedrock Agents and Knowledge Bases, or call the ApplyGuardrail API to check text for any model, including ones hosted outside Bedrock. - Input tagging marks which part of the prompt is user input, so the prompt attack filter checks only that part and not your system prompt.
- The Standard tier extends detection to more languages, code elements and prompt leakage. Classic is the original tier.
- Create a version of a guardrail before production use. Drafts change as you edit.
- Enforce use of a guardrail with the IAM condition key
bedrock:GuardrailIdentifieron inference actions, or organization-wide with an Amazon Bedrock policy in AWS Organizations, which applies a versioned guardrail from the management account to every model call in the targeted OUs and accounts.
Exam signal
"Users are trying to make the chatbot ignore its instructions" is the prompt attack filter. "The support bot must never discuss legal advice" is a denied topic. "Redact customers' phone numbers from summaries" is a sensitive information filter set to mask. "Answers must be based only on the retrieved documents" is a contextual grounding check.
Guardrails as authorization
A denied topic that blocks "show me other customers' orders" doesn't stop a RAG app from retrieving another tenant's documents. Tenant isolation belongs in the retrieval layer (knowledge base metadata filters built from the caller's verified identity, or separate indexes) and in IAM, not in a filter on the model's text.
OWASP Top 10 for LLM Applications (2025) on AWS
| Risk | What goes wrong | AWS controls |
|---|---|---|
| LLM01 Prompt injection | Direct or indirect (in a retrieved page or email) instructions override yours | Guardrails prompt attack filter with input tagging, treat retrieved content as data, least-privilege tools |
| LLM02 Sensitive information disclosure | Model reveals PII, secrets or other users' data | Sensitive information filters, Macie on training and RAG sources, metadata filters per tenant, don't put secrets in prompts |
| LLM03 Supply chain | Compromised model, plugin, library or dataset | Bedrock managed models, private model import from trusted sources, Inspector and ECR scanning of app dependencies |
| LLM04 Data and model poisoning | Tampered training or knowledge base data | Write-restricted S3 sources, versioning and Object Lock, CloudTrail data events, review before ingest |
| LLM05 Improper output handling | Model output executed as code, SQL or HTML | Validate and encode output, parameterized queries, WAF on downstream apps, never pass output to a shell |
| LLM06 Excessive agency | Agent can take more actions than the task needs | Narrow action groups, least-privilege Lambda execution roles, user confirmation for writes, human approval steps |
| LLM07 System prompt leakage | Prompt reveals internal rules, keys or logic | Keep secrets out of prompts, Standard tier prompt leakage detection, enforce rules in code |
| LLM08 Vector and embedding weaknesses | Cross-tenant retrieval, poisoned embeddings | Per-tenant indexes or metadata filters, encryption with KMS, access control on the vector store |
| LLM09 Misinformation | Confident, wrong answers | Contextual grounding checks, Automated Reasoning checks, citations, human review for high-impact answers |
| LLM10 Unbounded consumption | Token floods, model extraction, runaway cost | API Gateway usage plans and throttling, WAF rate-based rules, Bedrock quotas, max token limits, budgets and alarms |
Least privilege for agents
- An agent acts with the permissions of its tools. Each Bedrock Agents action group calls a Lambda function or API, so scope that function's execution role to the exact resources and actions.
- Pass the end user's identity to tools and authorize per user, rather than giving the agent a role that can read everyone's data.
- Turn on user confirmation (or return of control) for actions that change state, such as refunds, deletions or emails.
- The agent's own service role should allow only the specific models, knowledge bases and guardrails it uses.
- Limit iterations and tool calls so a looping agent can't run up cost or spam an API.
Keeping data from leaking
- Bedrock doesn't use your prompts or completions to train its models, and model providers don't see them.
- Reach Bedrock through an interface VPC endpoint (PrivateLink) and restrict it with an endpoint policy that allows only approved model ARNs. Deny other models with SCPs.
- Encrypt custom models, knowledge bases, agent sessions and logs with customer managed KMS keys.
- Classify the S3 data that feeds knowledge bases with Macie before ingestion. See Sensitive data.
Logging model invocations
- CloudTrail records Bedrock API calls (who invoked which model, when), not the prompt text.
- Model invocation logging is off by default. Turn it on per Region to capture full request and response bodies to CloudWatch Logs and/or S3. Encrypt with KMS and restrict read access, because the logs contain exactly the sensitive data you're protecting.
- Blocked content also appears in plain text in invocation logs. Mask sensitive data with guardrails or a CloudWatch Logs data protection policy if the log readers shouldn't see it.
- Guardrail metrics and trace output show which policy intervened, which is how you tune thresholds.
Scenarios
Masking replaces the sensitive values and keeps the rest of the summary, and the prompt attack filter catches attempts to override instructions or leak the system prompt in the user's part of the input. A denied topic would block whole summaries, and a word filter only matches exact words. A log data protection policy protects logs, not what agents see. IAM can't inspect prompt content.
An Organizations Bedrock policy enforces the guardrail on model inference calls across the targeted accounts without relying on application code. Code reviews and logging only detect missing guardrails after the fact. WAF inspects HTTP requests but doesn't apply Bedrock's content, PII or prompt attack policies, and teams can call Bedrock directly around the API.
The flaw is excessive agency: the tool trusted the model's claim about who was asking. Authorizing in the tool with the real caller identity fixes it, and user confirmation adds an explicit check before a state-changing action. A denied topic is a text filter that rewording can bypass. Grounding checks target hallucinations, and a broader role makes things worse.
Further reading
Vulnerability and patch management
Amazon Inspector for EC2, ECR, Lambda and code repositories, GuardDuty Runtime Monitoring, Systems Manager Patch Manager and State Manager, and security scanning inside CI/CD pipelines with Amazon Q Developer and Inspector.
Network controls
Security groups and network ACLs, ephemeral ports, AWS Network Firewall with Suricata rules and domain lists, Route 53 Resolver DNS Firewall, Gateway Load Balancer with third-party IDS/IPS, and egress filtering.