Model Customization Spectrum: from Prompt Engineering to Pre-training from Scratch
1. Model Customization Spectrum
There are many ways to customize FM behavior, from simple to complex:
Least Effort Most Effort
──────────────────────────────────────────────────────────
Prompt Few-shot RAG Fine- Continued Pre-
Engineering Prompting tuning Pre-training training
──────────────────────────────────────────────────────────
No training ← → Full training
$ cheapest ← → $$$$ most expensive
Minutes ← → Weeks/Months
2. Fine-tuning
Fine-tuning = further training an existing FM on your specific dataset to improve performance on your domain/task.
2.1. When to Fine-tune?
| Fine-tune When... | DON'T Fine-tune When... |
|---|---|
| Need specific style, tone, or format | Just need factual Q&A (use RAG) |
| Domain-specific language patterns | Task works well with prompting |
| Improve accuracy on specific tasks | Don't have labeled training data |
| Reduce prompt size (internalize instructions) | Data changes frequently (use RAG) |
| Need consistent output format | Budget is limited |
2.2. Types of Fine-tuning
| Type | What | Data Format | Use Case |
|---|---|---|---|
| Instruction fine-tuning | Train on prompt-response pairs | {"prompt": "...", "completion": "..."} | Follow instructions better |
| Domain adaptation | Train on domain text | Domain documents (medical, legal) | Learn domain terminology |
| Task-specific | Train on specific task examples | Task input-output pairs | Classification, extraction |
3. PEFT & LoRA
3.1. Parameter-Efficient Fine-Tuning (PEFT)
Full fine-tuning updates ALL model parameters — expensive and needs lots of GPU memory. PEFT methods update only a small subset of parameters.
Full Fine-tuning:
Model: 7 billion parameters
Updated: 7 billion parameters (100%)
GPU Memory: Very high
Cost: $$$$
PEFT (LoRA):
Model: 7 billion parameters
Updated: ~10 million parameters (0.1%)
GPU Memory: Much lower
Cost: $$
3.2. LoRA (Low-Rank Adaptation)
LoRA adds small trainable matrices to model layers instead of updating all weights:
- Freezes original model weights
- Adds small "adapter" matrices (rank decomposition)
- Only trains these small adapters
- At inference: merge adapters with original weights
Exam tip: "Which technique reduces the cost of fine-tuning while maintaining quality?" → LoRA / PEFT. Key concept: train a small percentage of parameters instead of all.
4. Continued Pre-training
Continued Pre-training trains the FM on large amounts of unlabeled domain data — teaching the model new vocabulary and concepts before fine-tuning on task-specific data.
Workflow:
Base FM → Continued Pre-training → Fine-tuning → Evaluation
(domain corpus, (labeled (test on
unlabeled) task data) holdout)
Example:
Base Claude → Train on 100K medical papers → Fine-tune on
(continued pre-training) medical Q&A pairs
Learns: medical terminology, Learns: how to
drug names, procedures answer clinical questions
Continued Pre-training vs Fine-tuning:
| Aspect | Continued Pre-training | Fine-tuning |
|---|---|---|
| Data | Large, unlabeled domain text | Smaller, labeled task data |
| Goal | Learn domain knowledge | Learn task-specific behavior |
| Cost | More expensive (larger data) | Less expensive |
| When | Model lacks domain vocabulary | Model needs to do specific tasks |
5. RLHF (Reinforcement Learning from Human Feedback)
RLHF is used to align model outputs with human preferences — making outputs more helpful, truthful, and harmless.
RLHF Pipeline:
1. Collect human feedback 2. Train reward model 3. Optimize with RL
"Which response is Learns: what humans FM generates →
better? A or B?" prefer reward model scores →
update FM weights
RLHF is mainly done by FM providers (Anthropic, Meta, Amazon) — not typically by end users. But you should know the concept for the exam.
6. Amazon Bedrock Custom Models
Bedrock offers two customization approaches:
6.1. Fine-tuning in Bedrock
| Feature | Detail |
|---|---|
| Supported models | Amazon Titan, Meta Llama, Cohere |
| Data format | JSONL with prompt-completion pairs |
| Data location | Amazon S3 |
| Output | Custom model version in Bedrock |
| Provisioned Throughput | Required to use fine-tuned model |
6.2. Continued Pre-training in Bedrock
| Feature | Detail |
|---|---|
| Supported models | Amazon Titan, Meta Llama, Cohere |
| Data format | Plain text files (unlabeled) |
| Use case | Domain adaptation before fine-tuning |
6.3. Training Data Preparation
// Fine-tuning data format (JSONL):
{"prompt": "What is the recommended dosage of Drug X?", "completion": "The recommended dosage of Drug X is 500mg twice daily for adults."}
{"prompt": "List side effects of Drug X.", "completion": "Common side effects include headache, nausea, and dizziness."}
6.4. Model Evaluation in Bedrock
Amazon Bedrock Model Evaluation allows you to compare models:
- Automatic evaluation: Built-in metrics (accuracy, robustness, toxicity)
- Human evaluation: Human reviewers rate model outputs
- Compare models: Side-by-side comparison of different FMs
Exam tip: "How to compare the quality of two foundation models for a specific use case?" → Amazon Bedrock Model Evaluation. Supports both automatic metrics and human evaluation.
7. Training Data Best Practices
| Practice | Why |
|---|---|
| High-quality data | Garbage in = garbage out |
| Diverse examples | Prevent overfitting to narrow patterns |
| Balanced classes | Avoid bias toward majority class |
| Clean data | Remove duplicates, errors, PII |
| Sufficient quantity | Typically 1000+ for fine-tuning |
| Train/validation split | Evaluate on unseen data |
| Format consistency | Same structure for all examples |
8. Summary: When to Use What
| Scenario | Best Approach |
|---|---|
| Simple task, model already good at it | Prompt Engineering |
| Need model to follow a specific pattern | Few-shot Prompting |
| Need answers from company documents | RAG |
| Need specific style/tone/format | Fine-tuning |
| Model doesn't know domain vocabulary | Continued Pre-training + Fine-tuning |
| Align with human preferences | RLHF (done by FM providers) |
9. Practice Questions
Q1: A legal firm wants their AI assistant to generate legal documents in a specific firm-approved writing style. They have 5,000 examples of approved documents. Which customization approach is MOST appropriate?
- A) RAG with a knowledge base
- B) Zero-shot prompting
- C) Fine-tuning on the approved document examples ✓
- D) Continued pre-training on legal textbooks
Explanation: Fine-tuning is ideal for teaching a model a specific writing style with labeled examples. RAG is for retrieving information, not learning styles. Continued pre-training would teach legal concepts but not the firm's specific style.
Q2: Which technique allows fine-tuning a large language model while updating only a small fraction of the model's parameters?
- A) Full fine-tuning
- B) LoRA (Low-Rank Adaptation) ✓
- C) Continued pre-training
- D) RLHF
Explanation: LoRA is a PEFT (Parameter-Efficient Fine-Tuning) method that adds small trainable adapter matrices while freezing the original model weights — typically updating less than 1% of total parameters.
Q3: A company fine-tuned a foundation model, but the model performs well on training data and poorly on new data. What is this problem called?
- A) Underfitting
- B) Overfitting ✓
- C) High bias
- D) Data drift
Explanation: Overfitting occurs when a model memorizes training data instead of learning general patterns. Solutions include: more training data, regularization, lower learning rate, early stopping, or data augmentation.