Chuyển đến nội dung chính

Bài 6: BigQuery ML & TensorFlow on GCP

BigQuery ML: CREATE MODEL syntax, supported models. TensorFlow Extended (TFX) pipeline components. TFServing, TFLite. Model optimization techniques.

BigQuery ML & TFX Pipeline

BigQuery ML và TFX Pipeline: train models bằng SQL, optimize model, và production ML pipelines

1. BigQuery ML (BQML)

BigQuery ML cho phép data analysts train và serve ML models bằng SQL trong BigQuery — không cần export data, không cần biết framework ML.

BigQuery ML Workflow:

1. CREATE MODEL → train
2. ML.EVALUATE() → evaluate metrics
3. ML.PREDICT() → generate predictions
4. ML.EXPLAIN_PREDICT() → SHAP-based explanations
5. EXPORT MODEL → export to Cloud Storage (TF SavedModel format)
Model TypeBQML OptionTask
Linear RegressionLINEAR_REGRegression
Logistic RegressionLOGISTIC_REGBinary/Multiclass classification
K-MeansKMEANSClustering
XGBoostBOOSTED_TREE_CLASSIFIER / BOOSTED_TREE_REGRESSORTabular classification/regression
Random ForestRANDOM_FOREST_CLASSIFIER / RANDOM_FOREST_REGRESSORTabular classification/regression
DNNDNN_CLASSIFIER / DNN_REGRESSORComplex patterns
Wide & DeepWIDE_AND_DEEP_CLASSIFIERRecommendations (memorization + generalization)
AutoMLAUTOML_CLASSIFIER / AUTOML_REGRESSORAutomated model selection
Time SeriesARIMA_PLUSForecasting
Matrix FactorizationMATRIX_FACTORIZATIONCollaborative filtering

Exam tip: BQML ARIMA_PLUS tự động xử lý seasonality, holiday effects, trend decomposition. Khi đề hỏi "forecast using BigQuery data" → ARIMA_PLUS. Khi hỏi "recommendation system in BigQuery" → MATRIX_FACTORIZATION.

2. TensorFlow Extended (TFX)

TFX là production ML pipeline library dành cho TensorFlow. Cung cấp standard components cho mỗi bước trong ML lifecycle.

TFX ComponentPurpose
ExampleGenIngest data từ CSV, BigQuery, Avro, Parquet
StatisticsGenCompute statistics về training data
SchemaGenInfer schema từ statistics
ExampleValidatorDetect anomalies: missing, distribution skew
TransformFeature engineering (Apache Beam-based)
TrainerTrain TF model (EvalSpec + TrainSpec)
TunerHyperparameter tuning (KerasTuner)
EvaluatorEvaluate model against baseline
ModelValidatorValidate model meets quality thresholds
PusherPush model to serving (TF Serving, Vertex AI)
TFX Pipeline (simplified):

ExampleGen → StatisticsGen → SchemaGen → ExampleValidator
                ↓
            Transform (feature engineering)
                ↓
            Trainer (model training)
                ↓
            Evaluator (metrics vs baseline)
                ↓ (if pass)
            Pusher → TF Serving / Vertex AI Endpoint

3. TF Serving & TFLite

OptionUse Case
TF ServingHigh-performance serving trên server/cloud (gRPC or REST)
TFLiteMobile devices, edge devices, microcontrollers
TF.jsBrowser-based inference

4. Model Optimization Techniques

TechniqueDescriptionTrade-off
QuantizationFloat32 → INT8 weights4x smaller, ~2x faster, slight accuracy loss
PruningRemove low-weight connectionsSmaller model, preserve accuracy
Knowledge DistillationTrain small "student" model from large "teacher"Smaller + fast, slight accuracy loss
TensorRTNVIDIA GPU optimization (layer fusion)3-5x inference speedup on NVIDIA GPUs

5. Practice Questions

Q1: A data analyst team needs to build a sales forecasting model on data already in BigQuery. They are comfortable with SQL but have no Python/ML framework experience. Which BigQuery ML model type should they use for time series forecasting?

  • A) KMEANS
  • B) LOGISTIC_REG
  • C) ARIMA_PLUS ✓
  • D) MATRIX_FACTORIZATION

Explanation: BigQuery ML ARIMA_PLUS is designed for time series forecasting and automatically handles seasonality, trend, and holiday effects. It can be trained with a simple CREATE MODEL statement in SQL, requiring no Python expertise.

Q2: A TFX pipeline is detecting that the distribution of the "age" feature in new production data differs significantly from the training data distribution. Which TFX component is responsible for detecting this anomaly?

  • A) StatisticsGen
  • B) SchemaGen
  • C) ExampleValidator ✓
  • D) Transform

Explanation: ExampleValidator compares data statistics against the expected schema and flags anomalies including distribution skew (significant difference between training and serving data distributions). StatisticsGen computes statistics; SchemaGen creates the schema; Transform does feature engineering.

Q3: A team needs to deploy a TensorFlow image classification model to mobile devices with limited compute resources. They need to reduce model size by 4x with minimal accuracy loss. Which technique should they apply?

  • A) Knowledge Distillation
  • B) Model Pruning
  • C) Post-training quantization (INT8) ✓
  • D) TensorRT optimization

Explanation: Post-training quantization converts Float32 weights to INT8, reducing model size by approximately 4x and improving inference speed by 2x, with minimal accuracy loss for most models. TFLite supports INT8 quantization for mobile/edge deployment. TensorRT is for NVIDIA GPUs, not mobile.