Chuyển đến nội dung chính

Lesson 7: Fine-tuning & Model Customization

Pre-training vs Fine-tuning vs RLHF. PEFT & LoRA. Continued Pre-training. Amazon Bedrock Custom Models. Training data preparation, evaluation, deployment.

Model Customization Spectrum

Model Customization Spectrum: from Prompt Engineering to Pre-training from Scratch

1. Model Customization Spectrum

There are many ways to customize FM behavior, from simple to complex:

Least Effort                                    Most Effort
──────────────────────────────────────────────────────────
Prompt       Few-shot      RAG       Fine-      Continued    Pre-
Engineering  Prompting               tuning     Pre-training training
──────────────────────────────────────────────────────────
No training                   ←                 →    Full training
$ cheapest                    ←                 →    $$$$ most expensive
Minutes                       ←                 →    Weeks/Months

2. Fine-tuning

Fine-tuning = further training an existing FM on your specific dataset to improve performance on your domain/task.

2.1. When to Fine-tune?

Fine-tune When...DON'T Fine-tune When...
Need specific style, tone, or formatJust need factual Q&A (use RAG)
Domain-specific language patternsTask works well with prompting
Improve accuracy on specific tasksDon't have labeled training data
Reduce prompt size (internalize instructions)Data changes frequently (use RAG)
Need consistent output formatBudget is limited

2.2. Types of Fine-tuning

TypeWhatData FormatUse Case
Instruction fine-tuningTrain on prompt-response pairs{"prompt": "...", "completion": "..."}Follow instructions better
Domain adaptationTrain on domain textDomain documents (medical, legal)Learn domain terminology
Task-specificTrain on specific task examplesTask input-output pairsClassification, extraction

3. PEFT & LoRA

3.1. Parameter-Efficient Fine-Tuning (PEFT)

Full fine-tuning updates ALL model parameters — expensive and needs lots of GPU memory. PEFT methods update only a small subset of parameters.

Full Fine-tuning:
  Model: 7 billion parameters
  Updated: 7 billion parameters (100%)
  GPU Memory: Very high
  Cost: $$$$

PEFT (LoRA):
  Model: 7 billion parameters
  Updated: ~10 million parameters (0.1%)
  GPU Memory: Much lower
  Cost: $$

3.2. LoRA (Low-Rank Adaptation)

LoRA adds small trainable matrices to model layers instead of updating all weights:

  • Freezes original model weights
  • Adds small "adapter" matrices (rank decomposition)
  • Only trains these small adapters
  • At inference: merge adapters with original weights

Exam tip: "Which technique reduces the cost of fine-tuning while maintaining quality?" → LoRA / PEFT. Key concept: train a small percentage of parameters instead of all.

4. Continued Pre-training

Continued Pre-training trains the FM on large amounts of unlabeled domain data — teaching the model new vocabulary and concepts before fine-tuning on task-specific data.

Workflow:
Base FM → Continued Pre-training → Fine-tuning → Evaluation
           (domain corpus,           (labeled        (test on
            unlabeled)                task data)       holdout)

Example:
Base Claude → Train on 100K medical papers → Fine-tune on
              (continued pre-training)       medical Q&A pairs
              Learns: medical terminology,   Learns: how to
              drug names, procedures         answer clinical questions

Continued Pre-training vs Fine-tuning:

AspectContinued Pre-trainingFine-tuning
DataLarge, unlabeled domain textSmaller, labeled task data
GoalLearn domain knowledgeLearn task-specific behavior
CostMore expensive (larger data)Less expensive
WhenModel lacks domain vocabularyModel needs to do specific tasks

5. RLHF (Reinforcement Learning from Human Feedback)

RLHF is used to align model outputs with human preferences — making outputs more helpful, truthful, and harmless.

RLHF Pipeline:
1. Collect human feedback    2. Train reward model    3. Optimize with RL
   "Which response is           Learns: what humans      FM generates →
    better? A or B?"             prefer                   reward model scores →
                                                          update FM weights

RLHF is mainly done by FM providers (Anthropic, Meta, Amazon) — not typically by end users. But you should know the concept for the exam.

6. Amazon Bedrock Custom Models

Bedrock offers two customization approaches:

6.1. Fine-tuning in Bedrock

FeatureDetail
Supported modelsAmazon Titan, Meta Llama, Cohere
Data formatJSONL with prompt-completion pairs
Data locationAmazon S3
OutputCustom model version in Bedrock
Provisioned ThroughputRequired to use fine-tuned model

6.2. Continued Pre-training in Bedrock

FeatureDetail
Supported modelsAmazon Titan, Meta Llama, Cohere
Data formatPlain text files (unlabeled)
Use caseDomain adaptation before fine-tuning

6.3. Training Data Preparation

// Fine-tuning data format (JSONL):
{"prompt": "What is the recommended dosage of Drug X?", "completion": "The recommended dosage of Drug X is 500mg twice daily for adults."}
{"prompt": "List side effects of Drug X.", "completion": "Common side effects include headache, nausea, and dizziness."}

6.4. Model Evaluation in Bedrock

Amazon Bedrock Model Evaluation allows you to compare models:

  • Automatic evaluation: Built-in metrics (accuracy, robustness, toxicity)
  • Human evaluation: Human reviewers rate model outputs
  • Compare models: Side-by-side comparison of different FMs

Exam tip: "How to compare the quality of two foundation models for a specific use case?" → Amazon Bedrock Model Evaluation. Supports both automatic metrics and human evaluation.

7. Training Data Best Practices

PracticeWhy
High-quality dataGarbage in = garbage out
Diverse examplesPrevent overfitting to narrow patterns
Balanced classesAvoid bias toward majority class
Clean dataRemove duplicates, errors, PII
Sufficient quantityTypically 1000+ for fine-tuning
Train/validation splitEvaluate on unseen data
Format consistencySame structure for all examples

8. Summary: When to Use What

ScenarioBest Approach
Simple task, model already good at itPrompt Engineering
Need model to follow a specific patternFew-shot Prompting
Need answers from company documentsRAG
Need specific style/tone/formatFine-tuning
Model doesn't know domain vocabularyContinued Pre-training + Fine-tuning
Align with human preferencesRLHF (done by FM providers)

9. Practice Questions

Q1: A legal firm wants their AI assistant to generate legal documents in a specific firm-approved writing style. They have 5,000 examples of approved documents. Which customization approach is MOST appropriate?

  • A) RAG with a knowledge base
  • B) Zero-shot prompting
  • C) Fine-tuning on the approved document examples ✓
  • D) Continued pre-training on legal textbooks

Explanation: Fine-tuning is ideal for teaching a model a specific writing style with labeled examples. RAG is for retrieving information, not learning styles. Continued pre-training would teach legal concepts but not the firm's specific style.

Q2: Which technique allows fine-tuning a large language model while updating only a small fraction of the model's parameters?

  • A) Full fine-tuning
  • B) LoRA (Low-Rank Adaptation) ✓
  • C) Continued pre-training
  • D) RLHF

Explanation: LoRA is a PEFT (Parameter-Efficient Fine-Tuning) method that adds small trainable adapter matrices while freezing the original model weights — typically updating less than 1% of total parameters.

Q3: A company fine-tuned a foundation model, but the model performs well on training data and poorly on new data. What is this problem called?

  • A) Underfitting
  • B) Overfitting ✓
  • C) High bias
  • D) Data drift

Explanation: Overfitting occurs when a model memorizes training data instead of learning general patterns. Solutions include: more training data, regularization, lower learning rate, early stopping, or data augmentation.