Introduction
Generative AI is the branch of AI that has the ability to generate new content — images, videos, sounds, text, code — instead of just classifying or predicting. This is the leap from "cognitive" AI to "creative" AI.
💡 Discriminative models answer "what is this?" — Generative models answers "how to create new ones?"
1. Discriminative vs Generative Models
┌──────────────────────────────────────────────────────────┐
│ DISCRIMINATIVE MODELS │
│ Input: x → P(y|x) → Label/Class │
│ "Ảnh này là mèo hay chó?" │
│ Ví dụ: CNN classifier, SVM, Logistic Regression │
├──────────────────────────────────────────────────────────┤
│ GENERATIVE MODELS │
│ Noise z → P(x) → Synthetic Data │
│ "Tạo một ảnh mèo mới" │
│ Ví dụ: GAN, VAE, Diffusion, GPT │
└──────────────────────────────────────────────────────────┘
| Features | Discriminative | Generative |
|---|---|---|
| Goal | Learn P(y|x) | Learn P(x) or P(x|z) |
| Output | Label, score | New data |
| Example | ResNet, BERT classifier | GAN, Stable Diffusion |
| Application | Classification, detection | Generation, synthesis |
2. History of Generative AI
| Year | Milestone | Model |
|---|---|---|
| 2013 | VAE was born | Kingma & Welling |
| 2014 | GAN: "The coolest idea in ML" | Ian Goodfellow |
| 2015 | DCGAN: stable GAN training | Radford et al. |
| 2018 | StyleGAN: realistic faces | NVIDIA |
| 2020 | DDPM: Diffusion Models | Ho et al. |
| 2021 | DALL-E, CLIP | OpenAI |
| 2022 | Stable Diffusion, Midjourney | Stability AI |
| 2023 | SDXL, DALL-E 3, Midjourney v5 | Multiple |
| 2024 | Sora, Flux, SD3 | OpenAI, BFL |
| 2025-26 | Video gene maturity, 3D gene | Multiple |
3. Taxonomy — Types of Generative Models
3.1 GAN — Generative Adversarial Networks
# Concept: 2 networks chơi game
# Generator: tạo fake data
# Discriminator: phân biệt real vs fake
# Generator
z = torch.randn(batch_size, latent_dim) # random noise
fake_images = generator(z) # tạo ảnh giả
# Discriminator
real_score = discriminator(real_images) # → 1 (real)
fake_score = discriminator(fake_images) # → 0 (fake)
# Training: Generator cố gắng "lừa" Discriminator
3.2 VAE — Variational Autoencoders
# Concept: Encode → Latent Space → Decode
# Input image → encoder → μ, σ → sample z → decoder → reconstructed image
# Ưu điểm: latent space có cấu trúc, có thể interpolation
# Nhược điểm: ảnh thường bị blurry
3.3 Diffusion Models
# Concept: Thêm noise dần → Học cách bỏ noise
# Forward: image → noisy → noisier → ... → pure noise
# Reverse: pure noise → less noisy → ... → clean image
# Ưu điểm: chất lượng cao nhất hiện tại
# Nhược điểm: chậm hơn GAN (nhiều steps)
3.4 Autoregressive Models
# Concept: Tạo từng phần một, dựa vào context trước
# GPT: tạo text token by token
# PixelCNN: tạo image pixel by pixel
# DALL-E 1: text → image tokens (autoregressive)
3.5 Flow-based Models
# Concept: Invertible transformations
# z → f1 → f2 → ... → x (exact likelihood)
# Normalizing Flows: RealNVP, Glow
# Ưu điểm: exact log-likelihood
# Nhược điểm: architecture constraints
4. Generative AI 2026 practical applications
Image Generation
- Marketing: Create advertising images, banners, product mockups
- Design: UI/UX prototyping, concept art
- E-commerce: Product photography, virtual try-on
Video Generation
- Content creation: Social media videos, ads
- Education: Teaching material, simulations
- Entertainment: SFX, animation
3D & Spatial
- Gaming: 3D asset generation
- Architecture: Interior design visualization
- AR/VR: Virtual environments
Audio & Music
- Podcasting: Voice synthesis, editing
- Music: Background music, jingles
- Accessibility: Text-to-speech
5. Generative AI Landscape 2026
┌──────────────────────────────────────────────────────┐
│ GENERATIVE AI STACK │
│ │
│ ┌──────────────┐ ┌──────────────┐ ┌───────────┐ │
│ │ Text-to- │ │ Text-to- │ │ Text-to- │ │
│ │ Image │ │ Video │ │ 3D │ │
│ │ │ │ │ │ │ │
│ │ • SD, SDXL │ │ • Sora │ │ • Point-E │ │
│ │ • DALL-E 3 │ │ • Runway │ │ • Magic3D │ │
│ │ • Midjourney│ │ • Kling │ │ • 3D GS │ │
│ │ • Flux │ │ • Pika │ │ │ │
│ └──────────────┘ └──────────────┘ └───────────┘ │
│ │
│ ┌──────────────────────────────────────────────┐ │
│ │ Foundation Models │ │
│ │ Diffusion Models, Transformers, GANs │ │
│ └──────────────────────────────────────────────┘ │
│ │
│ ┌──────────────────────────────────────────────┐ │
│ │ Infrastructure │ │
│ │ GPU Cloud, Model Serving, Storage │ │
│ └──────────────────────────────────────────────┘ │
└──────────────────────────────────────────────────────┘
6. Set up the development environment
# Tạo virtual environment
python -m venv genai-env
source genai-env/bin/activate
# Cài đặt core packages
pip install torch torchvision torchaudio
pip install diffusers transformers accelerate
pip install Pillow opencv-python matplotlib
# Verify GPU
python -c "import torch; print(f'CUDA: {torch.cuda.is_available()}')"
python -c "import diffusers; print(f'Diffusers: {diffusers.__version__}')"
# Quick test: generate image với Stable Diffusion
from diffusers import StableDiffusionPipeline
import torch
pipe = StableDiffusionPipeline.from_pretrained(
"stabilityai/stable-diffusion-xl-base-1.0",
torch_dtype=torch.float16,
variant="fp16"
)
pipe.to("cuda")
image = pipe("a cat wearing sunglasses, digital art").images[0]
image.save("first_genai_image.png")
print("✅ Generated first image!")
Summary
| Concept | Description |
|---|---|
| Generative AI | AI creates new content (image, video, audio, text) |
| GAN | Generator vs Discriminator — advanced training |
| VAE | Encode-decode via latent space — structured generation |
| Diffusion | Noise → denoise step-by-step — highest quality |
| Autoregressive | Partial sequential generation — GPT, PixelCNN |
📌 Next article: Deep dive GAN — Generative Adversarial Networks from zero, training dynamics, and important variants.