1. Google's Responsible AI Principles
| Principle | Key Requirement |
|---|---|
| Socially Beneficial | Benefits society and individuals |
| Avoid Unfair Bias | Test fairness across demographic groups |
| Safety | Test across diverse scenarios, continuous evaluation |
| Accountable | Appropriate human oversight and control |
| Privacy Preserving | Protect training data privacy |
| Scientific Excellence | Rigorous research standards |
| Available for Beneficial Uses | Primary benefit criteria |
2. Vertex AI Explainability
Vertex AI Explainability cung cấp feature attribution scores — giải thích tại sao model đưa ra prediction nào đó.
| Method | For | How |
|---|---|---|
| SHAP (Shapley Values) | Tabular models | Game theory: contribution của mỗi feature |
| Integrated Gradients (IG) | Neural networks (image, text) | Gradient accumulation from baseline to input |
| XRAI | Image models | Pixel-region attribution (better UX than IG) |
| Sampled Shapley | Large tabular datasets | Approximate SHAP, faster |
Exam tip: "Explain why a loan was denied" → SHAP for tabular models. "Highlight which image regions drove classification" → Integrated Gradients or XRAI. Vertex AI Explainability phải được enable lúc deploy endpoint.
3. Fairness & Bias Detection
| Tool/Concept | Description |
|---|---|
| Fairness Indicators | GCP tool: evaluate model fairness metrics across demographic slices |
| What-If Tool | Interactive exploration of model behavior, counterfactuals |
| Demographic parity | Model predicts same rate across demographic groups |
| Equal opportunity | Same recall/TPR across groups |
| Data slice evaluation | Evaluate metrics per gender, race, age in TFX Evaluator |
4. Privacy Techniques
| Technique | Description |
|---|---|
| Differential Privacy | Add statistical noise to training data/model, prevents individual data re-identification |
| Federated Learning | Train on distributed data without centralizing raw data — model updates only |
| Data Anonymization | Remove PII before training (Cloud DLP API) |
5. Security Controls for ML Workloads
| Control | Purpose |
|---|---|
| IAM roles | Least-privilege access for ML service accounts |
| VPC Service Controls (VPC-SC) | Security perimeter: prevent data exfiltration from BigQuery, GCS |
| CMEK (Customer-Managed Encryption Keys) | Control encryption keys via Cloud KMS |
| Private IP for Vertex AI | Training and endpoints use private networking |
| Cloud Audit Logs | Who accessed what data, when (Data Access + Admin Activity) |
VPC Service Controls Perimeter:
┌────── Security Perimeter ─────────┐
│ BigQuery │ Cloud Storage │
│ Vertex AI │ Cloud KMS │
│ Dataflow │ Secret Manager │
└──────────────────────────────────┘
│ (no exfiltration outside perimeter)
✗ Unauthorized access blocked
6. Practice Questions
Q1: A financial services company deployed a loan approval ML model. Regulators require the company to explain why specific loan applications were denied. Which Vertex AI feature provides per-prediction feature importance scores for tabular models?
- A) Vertex AI Experiments
- B) Vertex AI Explainability with SHAP ✓
- C) Vertex AI Model Monitoring
- D) Fairness Indicators
Explanation: Vertex AI Explainability with Shapley Values (SHAP) assigns an importance score to each feature for each individual prediction, explaining why a specific loan was denied by attributing the model's decision to specific input features like credit_score, income, debt_ratio.
Q2: A healthcare company needs to train ML models on patient data distributed across multiple hospitals. Data privacy regulations prohibit centralizing raw patient records. Which privacy-preserving ML approach should they use?
- A) Differential Privacy with central training
- B) Federated Learning ✓
- C) Data anonymization + BigQuery ML
- D) Cloud DLP de-identification
Explanation: Federated Learning trains models on distributed data without moving raw data to a central location. Each hospital trains locally on its own data; only model updates (gradients) are shared and aggregated. Raw patient records never leave the hospital's environment.
Q3: A company processes sensitive financial data in BigQuery for ML training. They need to prevent data from being moved outside an approved security boundary to unauthorized GCP projects. Which GCP feature should they implement?
- A) Cloud KMS CMEK encryption
- B) VPC Service Controls (VPC-SC) perimeter ✓
- C) IAM role deny policies
- D) Cloud Armor WAF
Explanation: VPC Service Controls creates a security perimeter around GCP services (BigQuery, Cloud Storage, Vertex AI). It prevents data exfiltration by blocking requests that would move data outside the defined perimeter, even from authenticated users. CMEK provides encryption control but doesn't prevent exfiltration.