Asterrr's Handbook

Generative AI guardrails

Amazon Bedrock Guardrails policies, the OWASP Top 10 for LLM Applications mapped to AWS controls, least privilege for agents, preventing data leakage, and logging model invocations.

Exam tasks: 3.2 (protections and guardrails for generative AI applications, applying the OWASP Top 10 for LLM Applications)

The decision: which risk sits in the model's input, which in its output, and which in what the application lets the model do, and which AWS control handles each without trusting the model to police itself?

Where controls sit in a GenAI request

The model is not a security boundary. Anything the model can reach, a clever prompt can ask it to reach. Put authorization in IAM, the tools and the data layer, and use guardrails to filter what goes in and out.

Amazon Bedrock Guardrails

PolicyChecksTypical action
Content filtersHate, insults, sexual, violence, misconduct in text and images, with a strength per categoryBlock
Prompt attack (a content filter category)Jailbreaks, prompt injection and prompt leakage attempts in the inputBlock
Denied topicsTopics you describe in natural language, such as investment advice or a competitor's productsBlock
Word filtersExact words and phrases, plus a managed profanity listBlock
Sensitive information filtersBuilt-in PII types (email, card numbers, national IDs) and your own regex patternsBlock or mask (anonymize), separately for input and output
Contextual grounding checksWhether the answer is supported by the source passages (grounding) and answers the question (relevance), against thresholds you setBlock or flag hallucinated RAG answers
Automated Reasoning checksWhether an answer is consistent with formal rules you define from a policy documentValidate and explain
  • A guardrail runs on the input before the model and on the output after it. You choose the message returned when something is blocked.
  • Attach it to InvokeModel, Converse, Bedrock Agents and Knowledge Bases, or call the ApplyGuardrail API to check text for any model, including ones hosted outside Bedrock.
  • Input tagging marks which part of the prompt is user input, so the prompt attack filter checks only that part and not your system prompt.
  • The Standard tier extends detection to more languages, code elements and prompt leakage. Classic is the original tier.
  • Create a version of a guardrail before production use. Drafts change as you edit.
  • Enforce use of a guardrail with the IAM condition key bedrock:GuardrailIdentifier on inference actions, or organization-wide with an Amazon Bedrock policy in AWS Organizations, which applies a versioned guardrail from the management account to every model call in the targeted OUs and accounts.

Exam signal

"Users are trying to make the chatbot ignore its instructions" is the prompt attack filter. "The support bot must never discuss legal advice" is a denied topic. "Redact customers' phone numbers from summaries" is a sensitive information filter set to mask. "Answers must be based only on the retrieved documents" is a contextual grounding check.

Guardrails as authorization

A denied topic that blocks "show me other customers' orders" doesn't stop a RAG app from retrieving another tenant's documents. Tenant isolation belongs in the retrieval layer (knowledge base metadata filters built from the caller's verified identity, or separate indexes) and in IAM, not in a filter on the model's text.

OWASP Top 10 for LLM Applications (2025) on AWS

RiskWhat goes wrongAWS controls
LLM01 Prompt injectionDirect or indirect (in a retrieved page or email) instructions override yoursGuardrails prompt attack filter with input tagging, treat retrieved content as data, least-privilege tools
LLM02 Sensitive information disclosureModel reveals PII, secrets or other users' dataSensitive information filters, Macie on training and RAG sources, metadata filters per tenant, don't put secrets in prompts
LLM03 Supply chainCompromised model, plugin, library or datasetBedrock managed models, private model import from trusted sources, Inspector and ECR scanning of app dependencies
LLM04 Data and model poisoningTampered training or knowledge base dataWrite-restricted S3 sources, versioning and Object Lock, CloudTrail data events, review before ingest
LLM05 Improper output handlingModel output executed as code, SQL or HTMLValidate and encode output, parameterized queries, WAF on downstream apps, never pass output to a shell
LLM06 Excessive agencyAgent can take more actions than the task needsNarrow action groups, least-privilege Lambda execution roles, user confirmation for writes, human approval steps
LLM07 System prompt leakagePrompt reveals internal rules, keys or logicKeep secrets out of prompts, Standard tier prompt leakage detection, enforce rules in code
LLM08 Vector and embedding weaknessesCross-tenant retrieval, poisoned embeddingsPer-tenant indexes or metadata filters, encryption with KMS, access control on the vector store
LLM09 MisinformationConfident, wrong answersContextual grounding checks, Automated Reasoning checks, citations, human review for high-impact answers
LLM10 Unbounded consumptionToken floods, model extraction, runaway costAPI Gateway usage plans and throttling, WAF rate-based rules, Bedrock quotas, max token limits, budgets and alarms

Least privilege for agents

  • An agent acts with the permissions of its tools. Each Bedrock Agents action group calls a Lambda function or API, so scope that function's execution role to the exact resources and actions.
  • Pass the end user's identity to tools and authorize per user, rather than giving the agent a role that can read everyone's data.
  • Turn on user confirmation (or return of control) for actions that change state, such as refunds, deletions or emails.
  • The agent's own service role should allow only the specific models, knowledge bases and guardrails it uses.
  • Limit iterations and tool calls so a looping agent can't run up cost or spam an API.

Keeping data from leaking

  • Bedrock doesn't use your prompts or completions to train its models, and model providers don't see them.
  • Reach Bedrock through an interface VPC endpoint (PrivateLink) and restrict it with an endpoint policy that allows only approved model ARNs. Deny other models with SCPs.
  • Encrypt custom models, knowledge bases, agent sessions and logs with customer managed KMS keys.
  • Classify the S3 data that feeds knowledge bases with Macie before ingestion. See Sensitive data.

Logging model invocations

  • CloudTrail records Bedrock API calls (who invoked which model, when), not the prompt text.
  • Model invocation logging is off by default. Turn it on per Region to capture full request and response bodies to CloudWatch Logs and/or S3. Encrypt with KMS and restrict read access, because the logs contain exactly the sensitive data you're protecting.
  • Blocked content also appears in plain text in invocation logs. Mask sensitive data with guardrails or a CloudWatch Logs data protection policy if the log readers shouldn't see it.
  • Guardrail metrics and trace output show which policy intervened, which is how you tune thresholds.
Off
Default state of Bedrock model invocation logging. Turn it on per Region.
Input + output
Guardrails evaluate both, with separate actions for each.
Block or mask
The two actions for sensitive information filters.
10
Risks in the OWASP Top 10 for LLM Applications, LLM01 to LLM10.

Scenarios

Scenario
A travel booking assistant built on Amazon Bedrock summarizes customer chat transcripts for support agents. Summaries must not show customers' passport numbers or card numbers, but the rest of the summary must still be shown. Some users also paste text telling the assistant to ignore its rules and reveal its system prompt. What should the security engineer configure?
Scenario
A bank has 40 AWS accounts with teams building Bedrock applications. The CISO requires that every model invocation in the organization pass through the bank's approved guardrail, regardless of how each team writes its code. What is the most effective way to enforce this?
Scenario · choose 2
An HR agent built with Amazon Bedrock Agents answers employee questions and can update employees' bank details through a Lambda action group. During testing, an employee got the agent to change a colleague's bank details by claiming to be their manager. Which TWO changes best address the root cause?

Further reading

On this page