Transformer Architecture — Encoder stack, Decoder stack và các biến thể BERT/GPT/T5
1. Transformer Architecture
Transformer là kiến trúc neural network đã cách mạng hoá NLP, được giới thiệu trong paper "Attention Is All You Need" (2017). Hầu hết LLMs hiện tại đều dựa trên Transformer.
1.1. Self-Attention Mechanism
Self-attention cho phép model xem xét mối quan hệ giữa tất cả các từ trong input, bất kể khoảng cách.
Input: "The cat sat on the mat because it was tired"
Self-attention answers: What does "it" refer to?
→ Attends to "cat" (high attention score)
→ Not "mat" (low attention score)
Traditional RNN would struggle with this long-range dependency.
1.2. Encoder-Decoder Architecture
Original Transformer:
┌──────────────────────────┐
│ ENCODER │ ← Understands input
│ (Self-Attention + │
│ Feed-Forward layers) │
├──────────────────────────┤
│ DECODER │ ← Generates output
│ (Masked Self-Attention +│
│ Cross-Attention + │
│ Feed-Forward layers) │
└──────────────────────────┘
1.3. Three Types of Transformers
| Type | Architecture | Best For | Models |
|---|---|---|---|
| Encoder-only | Encoder | Understanding text (classification, NER, sentiment) | BERT, RoBERTa, DistilBERT |
| Decoder-only | Decoder | Generating text (chatbot, content creation) | GPT-4, Claude, Llama |
| Encoder-Decoder | Both | Sequence-to-sequence (translation, summarization) | T5, BART |
Exam tip: "Which architecture is best for text generation?" → Decoder-only (GPT, Claude). "Which architecture is best for text classification?" → Encoder-only (BERT).
2. Large Language Models (LLMs)
LLMs là Foundation Models specifically для text — trained on massive text corpora to understand and generate human language.
2.1. LLM Capabilities
| Capability | Description | Example |
|---|---|---|
| Text Generation | Create new text content | Articles, emails, stories |
| Summarization | Condense long text | Document summaries |
| Translation | Convert between languages | English → Vietnamese |
| Q&A | Answer questions | Customer support, FAQ |
| Code Generation | Write and explain code | Amazon Q Developer |
| Text Classification | Categorize text | Sentiment analysis |
| Reasoning | Logical analysis | Math problems, step-by-step reasoning |
2.2. LLM Limitations
- Knowledge cutoff: Doesn't know events after training data cutoff date
- Hallucination: Can generate false information confidently
- Context window limit: Can't process unlimited text
- No real-time data: Can't access internet or live data (unless augmented)
- Expensive: Large models need significant compute for inference
- Bias: Can reflect biases in training data
3. Embeddings & Vector Representations
Embeddings biến text (hoặc images, audio) thành numerical vectors mà machines hiểu được. Các text có ý nghĩa tương tự sẽ có vectors gần nhau trong không gian nhiều chiều.
Text: "King" → [0.23, 0.87, -0.12, 0.45, ...]
Text: "Queen" → [0.21, 0.89, -0.15, 0.43, ...] ← Close vectors!
Text: "Banana" → [0.91, -0.32, 0.67, -0.88, ...] ← Far away
Relationship: King - Man + Woman ≈ Queen
Why Embeddings Matter for the Exam:
- Semantic search: Find similar documents based on meaning (not just keywords)
- RAG: Convert documents to embeddings, store in vector DB, retrieve relevant context
- Clustering: Group similar documents/sentences
- Amazon Titan Embeddings: AWS model specifically for creating text embeddings
Vector Databases
Store and search embeddings efficiently:
| Vector DB | Notes |
|---|---|
| Amazon OpenSearch Serverless | AWS-managed vector search |
| Amazon Aurora (pgvector) | PostgreSQL with vector extension |
| Pinecone | Popular third-party vector DB |
| Amazon Bedrock Knowledge Bases | Managed RAG — handles vector storage internally |
4. Multi-modal Models
Multi-modal models có thể xử lý và tạo nội dung từ nhiều loại data types (text + images + audio + video).
Examples on AWS:
| Model | Modalities | What it can do |
|---|---|---|
| Claude 3 (Anthropic) | Text + Image input → Text output | Describe images, analyze charts, visual Q&A |
| Amazon Titan Image Generator | Text → Image | Create images from text descriptions |
| Amazon Titan Multimodal Embeddings | Text + Image → Vectors | Search across text and images |
| Stable Diffusion (Stability AI) | Text → Image | Generate and edit images |
Multi-modal Use Cases for Exam:
- "Analyze product images and generate descriptions" → Multi-modal model (Claude 3 Vision)
- "Generate product images from text descriptions" → Text-to-image (Titan Image Generator, Stable Diffusion)
- "Search across both text documents and images" → Multi-modal embeddings
5. Diffusion Models
Diffusion models (như Stable Diffusion) hoạt động bằng cách:
- Forward process: Gradually add noise to an image until it becomes pure noise
- Reverse process: Learn to remove noise step by step, generating a new image
Training (Forward):
Clean Image → Add Noise → Add More Noise → ... → Pure Noise
Generation (Reverse):
Pure Noise → Remove Noise → Remove More Noise → ... → New Image
(guided by text prompt)
Exam tip: Bạn không cần biết math chi tiết, chỉ cần hiểu concept: diffusion models tạo images bằng cách từ từ khử noise có guided bởi text prompt.
6. Pre-training vs Fine-tuning vs Prompting
| Method | What | Data Needed | Cost | When to Use |
|---|---|---|---|---|
| Pre-training | Train from scratch | Billions of examples | $$$$ | Creating new FM (done by providers) |
| Fine-tuning | Further train existing FM | Thousands of examples | $$ | Domain-specific knowledge |
| Prompt Engineering | Craft better inputs | None (few examples) | $ | Quick adaptation, no training needed |
| RAG | Augment with external data | Knowledge base | $ | Access current/proprietary data |
Decision Tree for Exam:
Need the model to know specific domain knowledge?
├── Is the knowledge in documents you can provide?
│ └── YES → RAG (Bedrock Knowledge Bases)
│ └── NO, model needs to learn patterns →
│ ├── Have thousands of training examples? → Fine-tuning
│ └── Only a few examples? → Few-shot prompting
├── General knowledge is enough? → Prompt Engineering (zero/few-shot)
7. Practice Questions
Q1: A company wants to search for relevant information across both product images and text descriptions. Which type of model would be MOST suitable?
- A) A text-only LLM
- B) A multi-modal embedding model ✓
- C) A diffusion model
- D) A RNN model
Explanation: Multi-modal embedding models can create vector representations of both text and images in the same vector space, enabling cross-modal search.
Q2: Which Transformer architecture is BEST suited for text generation tasks such as chatbots and content creation?
- A) Encoder-only (BERT)
- B) Decoder-only (GPT, Claude) ✓
- C) Encoder-decoder (T5)
- D) Convolutional Neural Network (CNN)
Explanation: Decoder-only architectures generate text one token at a time (autoregressive) and are the basis for most modern chatbots and text generators.
Q3: What is the purpose of text embeddings in the context of generative AI applications?
- A) To compress files for storage
- B) To convert text into numerical vectors that capture semantic meaning ✓
- C) To encrypt text for security
- D) To translate text between languages
Explanation: Embeddings are numerical vector representations of text that capture semantic meaning. Similar texts have similar vectors, enabling semantic search, RAG, and clustering.