Transformer Architecture — Encoder stack, Decoder stack, and the BERT/GPT/T5 variants
1. Transformer Architecture
The Transformer is a neural network architecture that revolutionized NLP, introduced in the paper "Attention Is All You Need" (2017). Nearly all current LLMs are based on the Transformer.
1.1. Self-Attention Mechanism
Self-attention allows the model to consider the relationships between all words in the input, regardless of distance.
Input: "The cat sat on the mat because it was tired"
Self-attention answers: What does "it" refer to?
→ Attends to "cat" (high attention score)
→ Not "mat" (low attention score)
Traditional RNN would struggle with this long-range dependency.
1.2. Encoder-Decoder Architecture
Original Transformer:
┌──────────────────────────┐
│ ENCODER │ ← Understands input
│ (Self-Attention + │
│ Feed-Forward layers) │
├──────────────────────────┤
│ DECODER │ ← Generates output
│ (Masked Self-Attention +│
│ Cross-Attention + │
│ Feed-Forward layers) │
└──────────────────────────┘
1.3. Three Types of Transformers
| Type | Architecture | Best For | Models |
|---|---|---|---|
| Encoder-only | Encoder | Understanding text (classification, NER, sentiment) | BERT, RoBERTa, DistilBERT |
| Decoder-only | Decoder | Generating text (chatbot, content creation) | GPT-4, Claude, Llama |
| Encoder-Decoder | Both | Sequence-to-sequence (translation, summarization) | T5, BART |
Exam tip: "Which architecture is best for text generation?" → Decoder-only (GPT, Claude). "Which architecture is best for text classification?" → Encoder-only (BERT).
2. Large Language Models (LLMs)
LLMs are Foundation Models specifically for text — trained on massive text corpora to understand and generate human language.
2.1. LLM Capabilities
| Capability | Description | Example |
|---|---|---|
| Text Generation | Create new text content | Articles, emails, stories |
| Summarization | Condense long text | Document summaries |
| Translation | Convert between languages | English → Vietnamese |
| Q&A | Answer questions | Customer support, FAQ |
| Code Generation | Write and explain code | Amazon Q Developer |
| Text Classification | Categorize text | Sentiment analysis |
| Reasoning | Logical analysis | Math problems, step-by-step reasoning |
2.2. LLM Limitations
- Knowledge cutoff: Doesn't know events after training data cutoff date
- Hallucination: Can generate false information confidently
- Context window limit: Can't process unlimited text
- No real-time data: Can't access internet or live data (unless augmented)
- Expensive: Large models need significant compute for inference
- Bias: Can reflect biases in training data
3. Embeddings & Vector Representations
Embeddings convert text (or images, audio) into numerical vectors that machines can understand. Texts with similar meanings will have vectors close to each other in multi-dimensional space.
Text: "King" → [0.23, 0.87, -0.12, 0.45, ...]
Text: "Queen" → [0.21, 0.89, -0.15, 0.43, ...] ← Close vectors!
Text: "Banana" → [0.91, -0.32, 0.67, -0.88, ...] ← Far away
Relationship: King - Man + Woman ≈ Queen
Why Embeddings Matter for the Exam:
- Semantic search: Find similar documents based on meaning (not just keywords)
- RAG: Convert documents to embeddings, store in vector DB, retrieve relevant context
- Clustering: Group similar documents/sentences
- Amazon Titan Embeddings: AWS model specifically for creating text embeddings
Vector Databases
Store and search embeddings efficiently:
| Vector DB | Notes |
|---|---|
| Amazon OpenSearch Serverless | AWS-managed vector search |
| Amazon Aurora (pgvector) | PostgreSQL with vector extension |
| Pinecone | Popular third-party vector DB |
| Amazon Bedrock Knowledge Bases | Managed RAG — handles vector storage internally |
4. Multi-modal Models
Multi-modal models can process and generate content from multiple data types (text + images + audio + video).
Examples on AWS:
| Model | Modalities | What it can do |
|---|---|---|
| Claude 3 (Anthropic) | Text + Image input → Text output | Describe images, analyze charts, visual Q&A |
| Amazon Titan Image Generator | Text → Image | Create images from text descriptions |
| Amazon Titan Multimodal Embeddings | Text + Image → Vectors | Search across text and images |
| Stable Diffusion (Stability AI) | Text → Image | Generate and edit images |
Multi-modal Use Cases for Exam:
- "Analyze product images and generate descriptions" → Multi-modal model (Claude 3 Vision)
- "Generate product images from text descriptions" → Text-to-image (Titan Image Generator, Stable Diffusion)
- "Search across both text documents and images" → Multi-modal embeddings
5. Diffusion Models
Diffusion models (like Stable Diffusion) work by:
- Forward process: Gradually add noise to an image until it becomes pure noise
- Reverse process: Learn to remove noise step by step, generating a new image
Training (Forward):
Clean Image → Add Noise → Add More Noise → ... → Pure Noise
Generation (Reverse):
Pure Noise → Remove Noise → Remove More Noise → ... → New Image
(guided by text prompt)
Exam tip: You don't need to know the detailed math — just understand the concept: diffusion models create images by gradually removing noise guided by a text prompt.
6. Pre-training vs Fine-tuning vs Prompting
| Method | What | Data Needed | Cost | When to Use |
|---|---|---|---|---|
| Pre-training | Train from scratch | Billions of examples | $$$$ | Creating new FM (done by providers) |
| Fine-tuning | Further train existing FM | Thousands of examples | $$ | Domain-specific knowledge |
| Prompt Engineering | Craft better inputs | None (few examples) | $ | Quick adaptation, no training needed |
| RAG | Augment with external data | Knowledge base | $ | Access current/proprietary data |
Decision Tree for Exam:
Need the model to know specific domain knowledge?
├── Is the knowledge in documents you can provide?
│ └── YES → RAG (Bedrock Knowledge Bases)
│ └── NO, model needs to learn patterns →
│ ├── Have thousands of training examples? → Fine-tuning
│ └── Only a few examples? → Few-shot prompting
├── General knowledge is enough? → Prompt Engineering (zero/few-shot)
7. Practice Questions
Q1: A company wants to search for relevant information across both product images and text descriptions. Which type of model would be MOST suitable?
- A) A text-only LLM
- B) A multi-modal embedding model ✓
- C) A diffusion model
- D) A RNN model
Explanation: Multi-modal embedding models can create vector representations of both text and images in the same vector space, enabling cross-modal search.
Q2: Which Transformer architecture is BEST suited for text generation tasks such as chatbots and content creation?
- A) Encoder-only (BERT)
- B) Decoder-only (GPT, Claude) ✓
- C) Encoder-decoder (T5)
- D) Convolutional Neural Network (CNN)
Explanation: Decoder-only architectures generate text one token at a time (autoregressive) and are the basis for most modern chatbots and text generators.
Q3: What is the purpose of text embeddings in the context of generative AI applications?
- A) To compress files for storage
- B) To convert text into numerical vectors that capture semantic meaning ✓
- C) To encrypt text for security
- D) To translate text between languages
Explanation: Embeddings are numerical vector representations of text that capture semantic meaning. Similar texts have similar vectors, enabling semantic search, RAG, and clustering.