Chuyển đến nội dung chính

Bài 10: AWS Responsible AI Tools — Clarify, A2I & Guardrails

Amazon SageMaker Clarify (bias detection, explainability). Amazon Augmented AI (A2I) — Human-in-the-loop. Amazon Bedrock Guardrails deep dive. Content moderation.

AWS Responsible AI Tools

AWS Responsible AI Tools: SageMaker Clarify, Amazon A2I và Bedrock Guardrails

1. Amazon SageMaker Clarify

SageMaker Clarify helps detect bias in data and models, and provides model explainability — the go-to AWS service for Responsible AI.

1.1. Three Core Capabilities

CapabilityWhenWhat it does
Pre-training bias detectionBefore trainingDetects imbalances in training data across demographic groups
Post-training bias detectionAfter trainingDetects bias in model predictions (e.g., different accuracy across groups)
Explainability (SHAP)After trainingShows feature contributions to each prediction

1.2. Key Bias Metrics in Clarify

MetricPre/Post trainingWhat it measures
Class Imbalance (CI)Pre-trainingDistribution of classes across groups
Difference in Proportions (DPL)Pre-trainingProportion of positive labels across groups
KL DivergencePre-trainingDistribution divergence between groups
Disparate Impact (DI)Post-trainingRatio of positive predictions across groups
Accuracy Difference (AD)Post-trainingAccuracy gap between groups
Treatment Equality (TE)Post-trainingRatio of FP to FN across groups

1.3. Clarify Workflow

1. Configure Clarify Job
   ├── Specify sensitive attributes (gender, age, race)
   ├── Define facets (groups to compare)
   └── Choose bias metrics to compute

2. Run Pre-training Analysis
   ├── Upload training dataset
   └── Get bias report on data distribution

3. Train Model

4. Run Post-training Analysis
   ├── Compare predictions across groups
   └── Get SHAP values for explainability

5. Monitor with SageMaker Model Monitor
   └── Detect bias drift over time in production

Exam tip: "Which AWS service can detect if a model makes more errors for one racial group vs another?" → SageMaker Clarify (post-training bias detection).

2. Amazon Augmented AI (Amazon A2I)

Amazon A2I cung cấp human-in-the-loop (HITL) workflows cho AI predictions — đặc biệt quan trọng khi model confidence thấp hoặc high-stakes decisions.

2.1. How A2I Works

AI Prediction Flow with A2I:
                                          ┌─────────────┐
User Request → AI Model → Confident?  YES → Return result
                              │
                              NO (below threshold)
                              ↓
                     ┌──────────────────┐
                     │  Create Human    │
                     │  Review Task     │
                     │  (A2I Workflow)   │
                     └────────┬─────────┘
                              ↓
                     ┌──────────────────┐
                     │  Human Reviewer  │  ← AWS Mechanical Turk
                     │  Reviews &       │  ← Private workforce
                     │  Corrects        │  ← Third-party vendor
                     └────────┬─────────┘
                              ↓
                     Return human-verified result

2.2. A2I Components

ComponentPurpose
Human review workflowDefines when and how to trigger human review
Worker task templateUI for human reviewers to make decisions
WorkforceWho does the review (private, Mechanical Turk, vendor)
Activation conditionsConfidence threshold triggers (e.g., < 95%)

2.3. Built-in A2I Integrations

ServiceA2I Use Case
Amazon TextractReview low-confidence document extractions
Amazon RekognitionReview low-confidence content moderation
Custom ML modelsAny SageMaker model can trigger A2I

Exam tip: "A healthcare company needs a human to review AI diagnoses when the model is less than 90% confident" → Amazon A2I with activation condition set to confidence < 90%.

3. Amazon Bedrock Guardrails — Deep Dive

3.1. Guardrail Policies

PolicyHow it worksConfiguration
Content filtersBlock by category + severityNone / Low / Medium / High for each category (Hate, Insults, Sexual, Violence, Misconduct)
Denied topicsDefine topics to blockNatural language description + sample phrases
Word filtersBlock specific wordsCustom word/phrase list + profanity filter toggle
Sensitive info (PII)Detect PII with actionBLOCK or ANONYMIZE for each PII type (SSN, email, phone, name, address, ...)
Contextual groundingCheck if answer is groundedGrounding threshold (0-1) + relevance threshold

3.2. How Guardrails Process Requests

User Input
    ↓
[INPUT GUARDRAILS]
    ├── Content filter check
    ├── Denied topic check
    ├── Word filter check
    ├── PII detection → BLOCK or ANONYMIZE
    ↓ (if passes all checks)
Foundation Model generates response
    ↓
[OUTPUT GUARDRAILS]
    ├── Content filter check
    ├── Denied topic check
    ├── Word filter check
    ├── PII detection → BLOCK or ANONYMIZE
    ├── Contextual grounding check
    ↓ (if passes all checks)
Response returned to user

If blocked → Return configured "blocked" message

3.3. Guardrails vs System Prompts

AspectSystem PromptGuardrails
EnforcementSoft — model may ignoreHard — enforced by the platform
Bypass riskCan be bypassed via prompt injectionCannot be bypassed by prompts
PII handlingModel asked to not output PIIProgrammatic detection & redaction
AuditabilityLimitedFull logging and metrics

Exam tip: "A company needs to GUARANTEE that PII is never in model responses" → Bedrock Guardrails (not system prompts, which can be bypassed).

4. Content Moderation on AWS

ServiceContent TypeUse Case
Amazon RekognitionImages & VideoDetect inappropriate content, faces, text
Amazon ComprehendTextToxicity detection, sentiment analysis
Bedrock GuardrailsFM input/outputFilter harmful content in GenAI apps
Amazon A2IAnyHuman review for edge cases

5. AI Governance

5.1. Governance Framework

AreaWhat to implement
PolicyOrganization-wide AI ethics guidelines
Risk assessmentEvaluate risks before deploying AI systems
MonitoringContinuous monitoring for bias, performance drift
Audit trailLog all model decisions for accountability
Human oversightHuman-in-the-loop for high-stakes decisions
DocumentationModel cards, AI Service Cards

5.2. SageMaker ML Governance

  • SageMaker Model Cards: Document model details and intended use
  • SageMaker Model Dashboard: Centralized view of all model status
  • SageMaker Model Monitor: Detect data drift, model quality degradation
  • SageMaker Role Manager: Fine-grained access control for ML

6. Practice Questions

Q1: An insurance company wants to ensure their claim approval model treats customers of all ages fairly. Which AWS service should they use to detect age-based bias in the model's predictions?

  • A) Amazon Rekognition
  • B) Amazon SageMaker Clarify ✓
  • C) Amazon Bedrock Guardrails
  • D) Amazon Comprehend

Explanation: SageMaker Clarify can run post-training bias analysis comparing model predictions across age groups, using metrics like Disparate Impact and Accuracy Difference.

Q2: A document processing application using Amazon Textract needs human review when the extracted data has low confidence. Which AWS service provides this capability?

  • A) Amazon SageMaker Ground Truth
  • B) Amazon Augmented AI (A2I) ✓
  • C) Amazon Mechanical Turk directly
  • D) Amazon Bedrock Agents

Explanation: Amazon A2I has built-in integration with Amazon Textract and can automatically trigger human review workflows when extraction confidence falls below a defined threshold.

Q3: A chatbot must NEVER reveal customer credit card numbers in its responses, even if the data exists in the knowledge base. Which approach provides the STRONGEST guarantee?

  • A) Add "never output credit card numbers" to the system prompt
  • B) Fine-tune the model to not output PII
  • C) Use Amazon Bedrock Guardrails with PII filters set to BLOCK ✓
  • D) Remove credit card numbers from the knowledge base

Explanation: Bedrock Guardrails with PII filters provide programmatic detection and blocking of credit card numbers in both input and output — this cannot be bypassed by prompt injection, unlike system prompts.