Chuyển đến nội dung chính

Lesson 9: Responsible AI & Security

Google Responsible AI principles. Vertex AI Explainability (SHAP, IG). Fairness indicators. Privacy: differential privacy, federated learning. IAM, VPC-SC, CMEK for ML workloads.

📝 Exam Prep — Lesson 9 Lesson 9: Responsible AI & Security

Google Cloud Professional Machine Learning Engineer Exam Prep

Part 5: Responsible AI & Review

xdev.asia

1. Google's Responsible AI Principles

PrincipleKey Requirement
Socially BeneficialBenefits society and individuals
Avoid Unfair BiasTest fairness across demographic groups
SafetyTest across diverse scenarios, continuous evaluation
AccountableAppropriate human oversight and control
Privacy PreservingProtect training data privacy
Scientific ExcellenceRigorous research standards
Available for Beneficial UsesPrimary benefit criteria

2. Vertex AI Explainability

Vertex AI Explainability provides feature attribution scores — explaining why the model made a specific prediction.

MethodForHow
SHAP (Shapley Values)Tabular modelsGame theory: each feature's contribution
Integrated Gradients (IG)Neural networks (image, text)Gradient accumulation from baseline to input
XRAIImage modelsPixel-region attribution (better UX than IG)
Sampled ShapleyLarge tabular datasetsApproximate SHAP, faster

Exam tip: "Explain why a loan was denied" → SHAP for tabular models. "Highlight which image regions drove classification" → Integrated Gradients or XRAI. Vertex AI Explainability must be enabled when deploying the endpoint.

3. Fairness & Bias Detection

Tool/ConceptDescription
Fairness IndicatorsGCP tool: evaluate model fairness metrics across demographic slices
What-If ToolInteractive exploration of model behavior, counterfactuals
Demographic parityModel predicts same rate across demographic groups
Equal opportunitySame recall/TPR across groups
Data slice evaluationEvaluate metrics per gender, race, age in TFX Evaluator

4. Privacy Techniques

TechniqueDescription
Differential PrivacyAdd statistical noise to training data/model, prevents individual data re-identification
Federated LearningTrain on distributed data without centralizing raw data — model updates only
Data AnonymizationRemove PII before training (Cloud DLP API)

5. Security Controls for ML Workloads

ControlPurpose
IAM rolesLeast-privilege access for ML service accounts
VPC Service Controls (VPC-SC)Security perimeter: prevent data exfiltration from BigQuery, GCS
CMEK (Customer-Managed Encryption Keys)Control encryption keys via Cloud KMS
Private IP for Vertex AITraining and endpoints use private networking
Cloud Audit LogsWho accessed what data, when (Data Access + Admin Activity)
VPC Service Controls Perimeter:

┌────── Security Perimeter ─────────┐
│  BigQuery  │  Cloud Storage       │
│  Vertex AI │  Cloud KMS           │
│  Dataflow  │  Secret Manager      │
└──────────────────────────────────┘
         │ (no exfiltration outside perimeter)
         ✗ Unauthorized access blocked

6. Practice Questions

Q1: A financial services company deployed a loan approval ML model. Regulators require the company to explain why specific loan applications were denied. Which Vertex AI feature provides per-prediction feature importance scores for tabular models?

  • A) Vertex AI Experiments
  • B) Vertex AI Explainability with SHAP ✓
  • C) Vertex AI Model Monitoring
  • D) Fairness Indicators

Explanation: Vertex AI Explainability with Shapley Values (SHAP) assigns an importance score to each feature for each individual prediction, explaining why a specific loan was denied by attributing the model's decision to specific input features like credit_score, income, debt_ratio.

Q2: A healthcare company needs to train ML models on patient data distributed across multiple hospitals. Data privacy regulations prohibit centralizing raw patient records. Which privacy-preserving ML approach should they use?

  • A) Differential Privacy with central training
  • B) Federated Learning ✓
  • C) Data anonymization + BigQuery ML
  • D) Cloud DLP de-identification

Explanation: Federated Learning trains models on distributed data without moving raw data to a central location. Each hospital trains locally on its own data; only model updates (gradients) are shared and aggregated. Raw patient records never leave the hospital's environment.

Q3: A company processes sensitive financial data in BigQuery for ML training. They need to prevent data from being moved outside an approved security boundary to unauthorized GCP projects. Which GCP feature should they implement?

  • A) Cloud KMS CMEK encryption
  • B) VPC Service Controls (VPC-SC) perimeter ✓
  • C) IAM role deny policies
  • D) Cloud Armor WAF

Explanation: VPC Service Controls creates a security perimeter around GCP services (BigQuery, Cloud Storage, Vertex AI). It prevents data exfiltration by blocking requests that would move data outside the defined perimeter, even from authenticated users. CMEK provides encryption control but doesn't prevent exfiltration.