Chuyển đến nội dung chính

Bài 4: LLMs, Transformers & Multi-modal Models

Transformer architecture: attention mechanism, self-attention. GPT (decoder-only), BERT (encoder-only), T5 (encoder-decoder). Multi-modal models. Hallucination: causes and mitigation. Embeddings và vector representations.

Transformer Architecture

Transformer Architecture — Encoder stack, Decoder stack và các biến thể BERT/GPT/T5

1. Transformer Architecture

Transformer là kiến trúc neural network đã cách mạng hoá NLP, được giới thiệu trong paper "Attention Is All You Need" (2017). Hầu hết LLMs hiện tại đều dựa trên Transformer.

1.1. Self-Attention Mechanism

Self-attention cho phép model xem xét mối quan hệ giữa tất cả các từ trong input, bất kể khoảng cách.

Input: "The cat sat on the mat because it was tired"

Self-attention answers: What does "it" refer to?
→ Attends to "cat" (high attention score)
→ Not "mat" (low attention score)

Traditional RNN would struggle with this long-range dependency.

1.2. Encoder-Decoder Architecture

Original Transformer:
┌──────────────────────────┐
│        ENCODER           │  ← Understands input
│  (Self-Attention +       │
│   Feed-Forward layers)   │
├──────────────────────────┤
│        DECODER           │  ← Generates output
│  (Masked Self-Attention +│
│   Cross-Attention +      │
│   Feed-Forward layers)   │
└──────────────────────────┘

1.3. Three Types of Transformers

TypeArchitectureBest ForModels
Encoder-onlyEncoderUnderstanding text (classification, NER, sentiment)BERT, RoBERTa, DistilBERT
Decoder-onlyDecoderGenerating text (chatbot, content creation)GPT-4, Claude, Llama
Encoder-DecoderBothSequence-to-sequence (translation, summarization)T5, BART

Exam tip: "Which architecture is best for text generation?" → Decoder-only (GPT, Claude). "Which architecture is best for text classification?" → Encoder-only (BERT).

2. Large Language Models (LLMs)

LLMs là Foundation Models specifically для text — trained on massive text corpora to understand and generate human language.

2.1. LLM Capabilities

CapabilityDescriptionExample
Text GenerationCreate new text contentArticles, emails, stories
SummarizationCondense long textDocument summaries
TranslationConvert between languagesEnglish → Vietnamese
Q&AAnswer questionsCustomer support, FAQ
Code GenerationWrite and explain codeAmazon Q Developer
Text ClassificationCategorize textSentiment analysis
ReasoningLogical analysisMath problems, step-by-step reasoning

2.2. LLM Limitations

  • Knowledge cutoff: Doesn't know events after training data cutoff date
  • Hallucination: Can generate false information confidently
  • Context window limit: Can't process unlimited text
  • No real-time data: Can't access internet or live data (unless augmented)
  • Expensive: Large models need significant compute for inference
  • Bias: Can reflect biases in training data

3. Embeddings & Vector Representations

Embeddings biến text (hoặc images, audio) thành numerical vectors mà machines hiểu được. Các text có ý nghĩa tương tự sẽ có vectors gần nhau trong không gian nhiều chiều.

Text: "King"     → [0.23, 0.87, -0.12, 0.45, ...]
Text: "Queen"    → [0.21, 0.89, -0.15, 0.43, ...]  ← Close vectors!
Text: "Banana"   → [0.91, -0.32, 0.67, -0.88, ...] ← Far away

Relationship: King - Man + Woman ≈ Queen

Why Embeddings Matter for the Exam:

  • Semantic search: Find similar documents based on meaning (not just keywords)
  • RAG: Convert documents to embeddings, store in vector DB, retrieve relevant context
  • Clustering: Group similar documents/sentences
  • Amazon Titan Embeddings: AWS model specifically for creating text embeddings

Vector Databases

Store and search embeddings efficiently:

Vector DBNotes
Amazon OpenSearch ServerlessAWS-managed vector search
Amazon Aurora (pgvector)PostgreSQL with vector extension
PineconePopular third-party vector DB
Amazon Bedrock Knowledge BasesManaged RAG — handles vector storage internally

4. Multi-modal Models

Multi-modal models có thể xử lý và tạo nội dung từ nhiều loại data types (text + images + audio + video).

Examples on AWS:

ModelModalitiesWhat it can do
Claude 3 (Anthropic)Text + Image input → Text outputDescribe images, analyze charts, visual Q&A
Amazon Titan Image GeneratorText → ImageCreate images from text descriptions
Amazon Titan Multimodal EmbeddingsText + Image → VectorsSearch across text and images
Stable Diffusion (Stability AI)Text → ImageGenerate and edit images

Multi-modal Use Cases for Exam:

  • "Analyze product images and generate descriptions" → Multi-modal model (Claude 3 Vision)
  • "Generate product images from text descriptions" → Text-to-image (Titan Image Generator, Stable Diffusion)
  • "Search across both text documents and images" → Multi-modal embeddings

5. Diffusion Models

Diffusion models (như Stable Diffusion) hoạt động bằng cách:

  1. Forward process: Gradually add noise to an image until it becomes pure noise
  2. Reverse process: Learn to remove noise step by step, generating a new image
Training (Forward):
Clean Image → Add Noise → Add More Noise → ... → Pure Noise

Generation (Reverse):
Pure Noise → Remove Noise → Remove More Noise → ... → New Image
                           (guided by text prompt)

Exam tip: Bạn không cần biết math chi tiết, chỉ cần hiểu concept: diffusion models tạo images bằng cách từ từ khử noise có guided bởi text prompt.

6. Pre-training vs Fine-tuning vs Prompting

MethodWhatData NeededCostWhen to Use
Pre-trainingTrain from scratchBillions of examples$$$$Creating new FM (done by providers)
Fine-tuningFurther train existing FMThousands of examples$$Domain-specific knowledge
Prompt EngineeringCraft better inputsNone (few examples)$Quick adaptation, no training needed
RAGAugment with external dataKnowledge base$Access current/proprietary data

Decision Tree for Exam:

Need the model to know specific domain knowledge?
├── Is the knowledge in documents you can provide?
│   └── YES → RAG (Bedrock Knowledge Bases)
│   └── NO, model needs to learn patterns →
│       ├── Have thousands of training examples? → Fine-tuning
│       └── Only a few examples? → Few-shot prompting
├── General knowledge is enough? → Prompt Engineering (zero/few-shot)

7. Practice Questions

Q1: A company wants to search for relevant information across both product images and text descriptions. Which type of model would be MOST suitable?

  • A) A text-only LLM
  • B) A multi-modal embedding model ✓
  • C) A diffusion model
  • D) A RNN model

Explanation: Multi-modal embedding models can create vector representations of both text and images in the same vector space, enabling cross-modal search.

Q2: Which Transformer architecture is BEST suited for text generation tasks such as chatbots and content creation?

  • A) Encoder-only (BERT)
  • B) Decoder-only (GPT, Claude) ✓
  • C) Encoder-decoder (T5)
  • D) Convolutional Neural Network (CNN)

Explanation: Decoder-only architectures generate text one token at a time (autoregressive) and are the basis for most modern chatbots and text generators.

Q3: What is the purpose of text embeddings in the context of generative AI applications?

  • A) To compress files for storage
  • B) To convert text into numerical vectors that capture semantic meaning ✓
  • C) To encrypt text for security
  • D) To translate text between languages

Explanation: Embeddings are numerical vector representations of text that capture semantic meaning. Similar texts have similar vectors, enabling semantic search, RAG, and clustering.