Chuyển đến nội dung chính

Lesson 4: LLMs, Transformers & Multi-modal Models

Transformer architecture: attention mechanism, self-attention. GPT (decoder-only), BERT (encoder-only), T5 (encoder-decoder). Multi-modal models. Hallucination: causes and mitigation. Embeddings and vector representations.

Transformer Architecture

Transformer Architecture — Encoder stack, Decoder stack, and the BERT/GPT/T5 variants

1. Transformer Architecture

The Transformer is a neural network architecture that revolutionized NLP, introduced in the paper "Attention Is All You Need" (2017). Nearly all current LLMs are based on the Transformer.

1.1. Self-Attention Mechanism

Self-attention allows the model to consider the relationships between all words in the input, regardless of distance.

Input: "The cat sat on the mat because it was tired"

Self-attention answers: What does "it" refer to?
→ Attends to "cat" (high attention score)
→ Not "mat" (low attention score)

Traditional RNN would struggle with this long-range dependency.

1.2. Encoder-Decoder Architecture

Original Transformer:
┌──────────────────────────┐
│        ENCODER           │  ← Understands input
│  (Self-Attention +       │
│   Feed-Forward layers)   │
├──────────────────────────┤
│        DECODER           │  ← Generates output
│  (Masked Self-Attention +│
│   Cross-Attention +      │
│   Feed-Forward layers)   │
└──────────────────────────┘

1.3. Three Types of Transformers

TypeArchitectureBest ForModels
Encoder-onlyEncoderUnderstanding text (classification, NER, sentiment)BERT, RoBERTa, DistilBERT
Decoder-onlyDecoderGenerating text (chatbot, content creation)GPT-4, Claude, Llama
Encoder-DecoderBothSequence-to-sequence (translation, summarization)T5, BART

Exam tip: "Which architecture is best for text generation?" → Decoder-only (GPT, Claude). "Which architecture is best for text classification?" → Encoder-only (BERT).

2. Large Language Models (LLMs)

LLMs are Foundation Models specifically for text — trained on massive text corpora to understand and generate human language.

2.1. LLM Capabilities

CapabilityDescriptionExample
Text GenerationCreate new text contentArticles, emails, stories
SummarizationCondense long textDocument summaries
TranslationConvert between languagesEnglish → Vietnamese
Q&AAnswer questionsCustomer support, FAQ
Code GenerationWrite and explain codeAmazon Q Developer
Text ClassificationCategorize textSentiment analysis
ReasoningLogical analysisMath problems, step-by-step reasoning

2.2. LLM Limitations

  • Knowledge cutoff: Doesn't know events after training data cutoff date
  • Hallucination: Can generate false information confidently
  • Context window limit: Can't process unlimited text
  • No real-time data: Can't access internet or live data (unless augmented)
  • Expensive: Large models need significant compute for inference
  • Bias: Can reflect biases in training data

3. Embeddings & Vector Representations

Embeddings convert text (or images, audio) into numerical vectors that machines can understand. Texts with similar meanings will have vectors close to each other in multi-dimensional space.

Text: "King"     → [0.23, 0.87, -0.12, 0.45, ...]
Text: "Queen"    → [0.21, 0.89, -0.15, 0.43, ...]  ← Close vectors!
Text: "Banana"   → [0.91, -0.32, 0.67, -0.88, ...] ← Far away

Relationship: King - Man + Woman ≈ Queen

Why Embeddings Matter for the Exam:

  • Semantic search: Find similar documents based on meaning (not just keywords)
  • RAG: Convert documents to embeddings, store in vector DB, retrieve relevant context
  • Clustering: Group similar documents/sentences
  • Amazon Titan Embeddings: AWS model specifically for creating text embeddings

Vector Databases

Store and search embeddings efficiently:

Vector DBNotes
Amazon OpenSearch ServerlessAWS-managed vector search
Amazon Aurora (pgvector)PostgreSQL with vector extension
PineconePopular third-party vector DB
Amazon Bedrock Knowledge BasesManaged RAG — handles vector storage internally

4. Multi-modal Models

Multi-modal models can process and generate content from multiple data types (text + images + audio + video).

Examples on AWS:

ModelModalitiesWhat it can do
Claude 3 (Anthropic)Text + Image input → Text outputDescribe images, analyze charts, visual Q&A
Amazon Titan Image GeneratorText → ImageCreate images from text descriptions
Amazon Titan Multimodal EmbeddingsText + Image → VectorsSearch across text and images
Stable Diffusion (Stability AI)Text → ImageGenerate and edit images

Multi-modal Use Cases for Exam:

  • "Analyze product images and generate descriptions" → Multi-modal model (Claude 3 Vision)
  • "Generate product images from text descriptions" → Text-to-image (Titan Image Generator, Stable Diffusion)
  • "Search across both text documents and images" → Multi-modal embeddings

5. Diffusion Models

Diffusion models (like Stable Diffusion) work by:

  1. Forward process: Gradually add noise to an image until it becomes pure noise
  2. Reverse process: Learn to remove noise step by step, generating a new image
Training (Forward):
Clean Image → Add Noise → Add More Noise → ... → Pure Noise

Generation (Reverse):
Pure Noise → Remove Noise → Remove More Noise → ... → New Image
                           (guided by text prompt)

Exam tip: You don't need to know the detailed math — just understand the concept: diffusion models create images by gradually removing noise guided by a text prompt.

6. Pre-training vs Fine-tuning vs Prompting

MethodWhatData NeededCostWhen to Use
Pre-trainingTrain from scratchBillions of examples$$$$Creating new FM (done by providers)
Fine-tuningFurther train existing FMThousands of examples$$Domain-specific knowledge
Prompt EngineeringCraft better inputsNone (few examples)$Quick adaptation, no training needed
RAGAugment with external dataKnowledge base$Access current/proprietary data

Decision Tree for Exam:

Need the model to know specific domain knowledge?
├── Is the knowledge in documents you can provide?
│   └── YES → RAG (Bedrock Knowledge Bases)
│   └── NO, model needs to learn patterns →
│       ├── Have thousands of training examples? → Fine-tuning
│       └── Only a few examples? → Few-shot prompting
├── General knowledge is enough? → Prompt Engineering (zero/few-shot)

7. Practice Questions

Q1: A company wants to search for relevant information across both product images and text descriptions. Which type of model would be MOST suitable?

  • A) A text-only LLM
  • B) A multi-modal embedding model ✓
  • C) A diffusion model
  • D) A RNN model

Explanation: Multi-modal embedding models can create vector representations of both text and images in the same vector space, enabling cross-modal search.

Q2: Which Transformer architecture is BEST suited for text generation tasks such as chatbots and content creation?

  • A) Encoder-only (BERT)
  • B) Decoder-only (GPT, Claude) ✓
  • C) Encoder-decoder (T5)
  • D) Convolutional Neural Network (CNN)

Explanation: Decoder-only architectures generate text one token at a time (autoregressive) and are the basis for most modern chatbots and text generators.

Q3: What is the purpose of text embeddings in the context of generative AI applications?

  • A) To compress files for storage
  • B) To convert text into numerical vectors that capture semantic meaning ✓
  • C) To encrypt text for security
  • D) To translate text between languages

Explanation: Embeddings are numerical vector representations of text that capture semantic meaning. Similar texts have similar vectors, enabling semantic search, RAG, and clustering.