Responsible AI Pillars và các điểm Bias xâm nhập trong ML Pipeline
1. What is Responsible AI?
Responsible AI là framework đảm bảo AI systems được phát triển và sử dụng một cách ethical, fair, transparent, và accountable.
1.1. Pillars of Responsible AI
| Pillar | Definition | Example |
|---|---|---|
| Fairness | Treat all groups equitably | Loan approval model doesn't discriminate by race |
| Explainability | Understand why model made a decision | "Your loan was denied because debt-to-income ratio > 0.5" |
| Transparency | Clear about AI capabilities & limitations | Disclose when content is AI-generated |
| Privacy | Protect personal data | Don't train on PII without consent |
| Safety | Prevent harmful outputs | Content filters, guardrails |
| Robustness | Reliable under adversarial conditions | Resist prompt injection attacks |
| Governance | Oversight and accountability | Human review for high-stakes decisions |
2. Understanding Bias in AI
2.1. Types of Bias
| Bias Type | What | Example |
|---|---|---|
| Selection bias | Training data doesn't represent population | Hiring model trained only on tech company data |
| Measurement bias | Inconsistent data collection | Different image quality across demographic groups |
| Confirmation bias | Model reinforces existing patterns | Recommender shows only what users already like |
| Label bias | Human labelers introduce biases | Inconsistent sentiment labels across annotators |
| Algorithmic bias | Model architecture amplifies bias | Optimizing for accuracy favors majority group |
| Recall bias | Overrepresented historical patterns | More arrest data in certain areas → predicts more crime there |
| Sampling bias | Non-random data collection | Online survey misses elderly population |
2.2. Where Bias Can Enter the ML Lifecycle
Data Collection Data Processing Model Training Evaluation Deployment
↓ ↓ ↓ ↓ ↓
Selection bias Feature engineering Algorithmic Evaluation Feedback
Sampling bias Missing values bias metric bias loop bias
Measurement Encoding choices Optimization User bias
bias objective
Exam tip: "Where can bias be introduced in an ML pipeline?" → At every stage — data collection, preprocessing, model training, evaluation, and deployment. This is why monitoring throughout the lifecycle is critical.
3. Fairness Metrics
3.1. Key Fairness Concepts
| Concept | Definition |
|---|---|
| Demographic parity | Positive outcomes at same rate across groups |
| Equal opportunity | Equal true positive rates across groups |
| Equalized odds | Equal TPR and FPR across groups |
| Individual fairness | Similar individuals get similar outcomes |
| Disparate impact | Ratio of positive outcomes between groups (80% rule) |
3.2. Detecting Bias
- Pre-training: Analyze training data distribution across demographic groups
- Post-training: Compare model predictions across groups
- Runtime: Monitor live predictions for drift in fairness metrics
4. Model Explainability
Explainability = ability to understand why a model made a specific prediction.
4.1. Explainability Techniques
| Technique | Type | What it does |
|---|---|---|
| SHAP (SHapley Additive exPlanations) | Model-agnostic | Shows contribution of each feature to prediction |
| LIME (Local Interpretable Model-agnostic Explanations) | Model-agnostic | Explains individual predictions by approximating locally |
| Feature importance | Model-specific | Ranks features by their impact on model output |
| Attention visualization | Transformer-specific | Shows which tokens the model focused on |
| Partial Dependence Plots | Model-agnostic | Shows how a feature affects predictions |
SHAP Example:
Loan Application: DENIED
Feature Contributions:
Debt-to-income ratio: +0.42 (pushes toward DENY)
Credit score: +0.28 (pushes toward DENY)
Employment years: -0.15 (pushes toward APPROVE)
Loan amount: +0.08 (pushes toward DENY)
Age: -0.03 (neutral)
─────────────
Base (avg prediction): 0.45
Final prediction: 0.45 + 0.42 + 0.28 - 0.15 + 0.08 - 0.03 = 1.05 → DENY
Exam tip: "How to explain why an ML model denied a loan application?" → SHAP values — shows the contribution of each feature to the individual prediction. SageMaker Clarify provides this on AWS.
5. Transparency in AI
5.1. AWS AI Service Cards
AI Service Cards are public documentation from AWS that provide transparency about AWS AI services:
- Intended use cases: What the service is designed for
- Limitations: Known limitations and failure modes
- Design choices: How the model was built
- Best practices: Recommended usage patterns
- Fairness considerations: Known demographic performance differences
Available for: Amazon Rekognition, Textract, Comprehend, Transcribe, etc.
5.2. Model Cards
Model Cards (from SageMaker) are internal documentation you create for your own models:
- Model description and intended use
- Training data details
- Performance metrics across subgroups
- Ethical considerations
- Limitations and risks
5.3. Transparency Best Practices
| Practice | How |
|---|---|
| Disclose AI usage | Tell users when they're interacting with AI |
| Source attribution | Cite sources in RAG applications |
| Confidence scores | Show model confidence to users |
| Limitations disclosure | Document what the model can't do |
| Watermarking | Mark AI-generated content (images, text) |
6. Toxicity & Harmful Content
Types of Harmful Content:
- Hate speech: Content targeting protected groups
- Violence: Graphic or promoting violence
- Sexual content: Explicit or inappropriate
- Self-harm: Promoting self-harm or suicide
- Misinformation: Factually incorrect content presented as fact
- Prompt injection: Malicious prompts that override system instructions
Mitigation Strategies:
- Content filters: Automated detection and blocking (Bedrock Guardrails)
- Human review: Human-in-the-loop for high-risk content
- Input sanitization: Validate and sanitize user inputs
- Output filtering: Check model outputs before showing to users
- Red teaming: Adversarial testing before deployment
7. Practice Questions
Q1: A hiring AI system consistently ranks male candidates higher than equally qualified female candidates. Which type of bias is MOST likely present?
- A) Measurement bias
- B) Selection bias in training data ✓
- C) Confirmation bias
- D) Recall bias
Explanation: If the training data contained historical hiring decisions that favored male candidates, the model would learn and reproduce that selection bias. The training data didn't represent the qualified population fairly.
Q2: A bank is required by regulators to explain why each loan application was approved or denied. Which AWS service feature can provide per-prediction explanations?
- A) Amazon Bedrock Guardrails
- B) Amazon SageMaker Clarify with SHAP values ✓
- C) Amazon Comprehend sentiment analysis
- D) AWS AI Service Cards
Explanation: SageMaker Clarify computes SHAP values that show the contribution of each feature to individual predictions, providing the explainability required by regulators.
Q3: Which AWS resource provides public documentation about the intended use cases, limitations, and fairness considerations of AWS AI services?
- A) SageMaker Model Cards
- B) AWS AI Service Cards ✓
- C) Amazon Bedrock Model Evaluation
- D) AWS Trusted Advisor
Explanation: AWS AI Service Cards are public documents that provide transparency about the design, limitations, and best practices for AWS AI services like Rekognition, Textract, and Comprehend. Model Cards are for your own custom models.