Chuyển đến nội dung chính

Lesson 4: Feature Engineering & Vertex AI Feature Store

Feature engineering techniques. BigQuery for feature computation. Vertex AI Feature Store: online/offline serving. Feature monitoring, training/serving consistency.

Vertex AI Feature Store

Feature Engineering & Vertex AI Feature Store: create, store, and reuse features for ML

1. Feature Engineering Techniques

TechniqueWhen to UseExample
Normalization (Min-Max)Bounded range required (0-1)Image pixels, probabilities
Standardization (Z-score)Normal-ish distribution, no boundsCustomer age, transaction amount
Log TransformSkewed distributions (price, salary)Log(price) for housing
One-Hot EncodingNominal categorical (no order)Country, brand, color
Label EncodingOrdinal categorical (has order)Low/Medium/High → 0/1/2
Feature CrossingCapture interaction between featurescity × day_of_week
BucketizingConvert continuous to categoricalAge → age_group
EmbeddingsHigh-cardinality categoricalUserID, ProductID

2. Handling Missing Values

StrategyWhen
Mean/Median imputationNumerical, low missingness rate
Mode imputationCategorical features
Model-based imputationHigh missingness, complex patterns
Indicator variableMissingness itself is informative (add is_missing flag)
Drop rowsMissing target / very few rows affected
Drop column>80% missing

3. Training-Serving Skew

Training-serving skew is a critical issue: features are computed differently between training and serving, causing the model to perform poorly in production despite good test metrics.

Training-Serving Skew Example:

TRAINING TIME:
  avg_purchase_last_30d = mean(all purchases in batch)  ← computed over full period

SERVING TIME:
  avg_purchase_last_30d = mean(last 5 purchases)        ← computed differently!

Result: Feature distribution mismatch → poor predictions

SOLUTION: Vertex AI Feature Store
  Same feature serve logic used at training AND serving time

4. Vertex AI Feature Store

ComponentDescription
Feature StoreCentralized repository for ML features
Entity TypeCategory of things you track (User, Product)
FeatureNamed attribute of an entity (user.avg_spend)
Online StoreLow-latency serving (ms) for real-time predictions
Offline StoreBigQuery-backed, for batch training data retrieval
Vertex AI Feature Store Architecture:

Feature Ingestion (Batch or Streaming)
        ↓
┌──── Feature Store ────────────────┐
│  Offline Store (BigQuery)          │  ← Training data export
│  Online Store (Bigtable-backed)    │  ← Serving (ms latency)
└───────────────────────────────────┘
        ↑ Same features ↑
  Training      Inference
  Pipeline      Endpoint

5. BigQuery for Feature Engineering

BigQuery is the best tool on GCP for computing aggregate features from large datasets.

Feature PatternBigQuery Approach
Rolling window aggregatesWindow functions: AVG() OVER (PARTITION BY ... ORDER BY ... ROWS BETWEEN ...)
User activity countsCOUNT() GROUP BY user_id
Categorical encodingCASE WHEN ... or ML.ONE_HOT_ENCODE()
Hash embedding (high cardinality)FARM_FINGERPRINT() mod N
Feature normalizationML.STANDARD_SCALER() in BigQuery ML

Exam tip: When a question mentions "training-serving consistency" or "feature reuse across multiple models" → Vertex AI Feature Store. When it mentions "compute features from BigQuery data at scale" → BigQuery window functions + scheduled queries.

6. Feature Drift Monitoring

TypeWhat ChangesDetection Method
Feature SkewTraining vs serving feature distribution differsCompare training baseline vs serving stats
Feature DriftServing features change over timeMonitor serving feature distributions daily
Label DriftTarget variable distribution changesTrack prediction distribution shifts

7. Practice Questions

Q1: A team's ML model has excellent accuracy during testing but performs poorly in production. Investigations reveal that the average purchase feature is calculated differently in training (using historical batch data) vs. serving (using real-time lookups). What is this problem called and how should it be solved?

  • A) Model drift — retrain the model more frequently
  • B) Training-serving skew — use Vertex AI Feature Store ✓
  • C) Data leakage — remove the purchase feature
  • D) Overfitting — add dropout layers

Explanation: Training-serving skew occurs when features are computed differently at training and serving time. Vertex AI Feature Store solves this by providing a single source of truth for feature computation, ensuring the same logic is used for both training data export and online serving.

Q2: A feature has values ranging from $10 to $10,000,000 with a heavily right-skewed distribution. Which transformation is MOST appropriate before using this feature in a linear model?

  • A) One-Hot Encoding
  • B) Min-Max Normalization
  • C) Log transformation ✓
  • D) Label Encoding

Explanation: Log transformation compresses the scale of highly skewed distributions, making them more normal-like and suitable for linear models. Min-Max normalization would still preserve the skew. One-hot encoding is for categorical data.

Q3: Which Vertex AI Feature Store store type is optimized for serving features to real-time prediction endpoints with millisecond latency?

  • A) Offline Store (BigQuery)
  • B) Online Store (Bigtable-backed) ✓
  • C) Feature Catalog
  • D) Cloud Memorystore

Explanation: The Online Store in Vertex AI Feature Store is backed by Bigtable and designed for sub-100ms latency lookups, serving fresh feature values to real-time prediction endpoints. The Offline Store uses BigQuery and is for batch training data retrieval.