Introduction
Prompt writing is not done yet. Prompt also needs testing — just like code. Model updates, data changes, prompt drift over time. Without testing, you won't know the prompt has a problem until the user complains.
This article builds a complete Prompt Testing Framework: from simple unit tests to regression tests, automatic evaluation pipelines.
Prompt Engineering Pipeline:
Write Prompt → Test → Evaluate → Deploy → Monitor → Update → Re-test
↑ |
└──────────────────────────────────────────────────────────┘
1. Why do we need Test Prompt?
1.1 Prompt Drift
Prompt DRIFT xảy ra khi:
1. Model update → GPT-4 → GPT-4o: output format thay đổi
2. Context thay đổi → Data mới, edge cases mới
3. Prompt edit nhỏ → Sửa 1 từ, output sai hoàn toàn
4. Temperature khác → Cùng prompt, output khác mỗi lần
→ Không test = không biết prompt đang broken
1.2 Types of tests
| Test type | Purpose | When to run |
|---|---|---|
| Unit Test | 1 prompt + 1 input → check output | Each time you edit the prompt |
| Regression Test | The old test cases still pass | Before deploy |
| A/B Test | Compare 2 prompt versions | Choose better version |
| Stress Test | Edge cases, advanced inputs | Before production |
| Smoke Test | Quick check basic functionality | After deploy |
2. Unit Test for Prompt
2.1 Test case structure
"""Unit test cho prompt"""
from dataclasses import dataclass
@dataclass
class PromptTestCase:
name: str # Tên test case
input_text: str # Input cho prompt
expected: dict # Expected output criteria
tags: list[str] # Tags: ["happy_path", "edge_case", "vi", "en"]
# Ví dụ: test prompt phân loại email
test_cases = [
PromptTestCase(
name="spam_clear",
input_text="Bạn trúng thưởng 1 tỷ! Click link ngay!",
expected={"category": "spam", "confidence_min": 0.9},
tags=["happy_path", "vi"],
),
PromptTestCase(
name="business_email",
input_text="Kính gửi anh, báo cáo Q3 đính kèm.",
expected={"category": "business", "confidence_min": 0.8},
tags=["happy_path", "vi"],
),
PromptTestCase(
name="ambiguous_email",
input_text="Hey, check this out",
expected={"category_not": "spam", "has_reasoning": True},
tags=["edge_case", "en"],
),
]
2.2 Test Runner
"""Simple prompt test runner"""
import json
from openai import OpenAI
client = OpenAI()
CLASSIFICATION_PROMPT = """Phân loại email sau vào 1 trong 4 categories:
spam, business, personal, newsletter.
Output JSON: {"category": "...", "confidence": 0.0-1.0, "reasoning": "..."}
Email: {input}"""
def run_test(test_case: PromptTestCase) -> dict:
"""Chạy 1 test case"""
prompt = CLASSIFICATION_PROMPT.format(input=test_case.input_text)
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": prompt}],
response_format={"type": "json_object"},
temperature=0, # Deterministic cho test
)
result = json.loads(response.choices[0].message.content)
# Kiểm tra assertions
passed = True
failures = []
if "category" in test_case.expected:
if result["category"] != test_case.expected["category"]:
passed = False
failures.append(f"category: got {result['category']}, "
f"expected {test_case.expected['category']}")
if "confidence_min" in test_case.expected:
if result.get("confidence", 0) < test_case.expected["confidence_min"]:
passed = False
failures.append(f"confidence too low: {result.get('confidence')}")
return {"passed": passed, "failures": failures, "output": result}
# Chạy toàn bộ test suite
for tc in test_cases:
result = run_test(tc)
status = "✅ PASS" if result["passed"] else "❌ FAIL"
print(f"{status} | {tc.name}: {result.get('failures', [])}")
3. Evaluation Metrics
3.1 Important metrics
┌─────────────────────────────────────────────┐
│ PROMPT EVALUATION METRICS │
├──────────────┬──────────────────────────────┤
│ Correctness │ Output đúng theo expected? │
│ Consistency │ N lần chạy → N kết quả giống │
│ Latency │ Thời gian response │
│ Token Usage │ Input + Output tokens │
│ Cost │ $/request │
│ Format │ Output đúng format (JSON...)?│
│ Safety │ Không toxic, hallucination? │
└──────────────┴──────────────────────────────┘
3.2 LLM-as-Judge
"""Dùng LLM đánh giá output của LLM khác"""
JUDGE_PROMPT = """Bạn là evaluator. Đánh giá output dưới đây theo các tiêu chí.
= INPUT =
{input}
= PROMPT OUTPUT =
{output}
= EXPECTED =
{expected}
= TIÊU CHÍ =
1. Correctness (0-10): Output có đúng không?
2. Completeness (0-10): Có đầy đủ thông tin không?
3. Format (0-10): Đúng format yêu cầu không?
4. Relevance (0-10): Output có liên quan đến input không?
Output JSON:
{"correctness": N, "completeness": N, "format": N, "relevance": N,
"overall": N, "reasoning": "..."} """
def evaluate_with_llm(input_text, output_text, expected):
"""LLM-as-Judge evaluation"""
response = client.chat.completions.create(
model="gpt-4o", # Dùng model mạnh hơn làm judge
messages=[{"role": "user", "content": JUDGE_PROMPT.format(
input=input_text, output=output_text, expected=expected,
)}],
response_format={"type": "json_object"},
temperature=0,
)
return json.loads(response.choices[0].message.content)
3.3 Consistency Test
"""Đo consistency: chạy cùng prompt N lần"""
def test_consistency(prompt: str, n_runs: int = 5):
"""Chạy prompt N lần, đo consistency"""
results = []
for _ in range(n_runs):
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": prompt}],
temperature=0.7, # Non-zero để test variance
)
results.append(response.choices[0].message.content)
# So sánh pairwise similarity
unique = len(set(results))
consistency_score = 1.0 - (unique - 1) / n_runs
return {
"n_runs": n_runs,
"unique_outputs": unique,
"consistency_score": consistency_score,
"results": results,
}
💡 Exercise 1: Write 5 test cases for 1 prompt you are using (classification, extraction, or generation). Run test runner, record pass/fail rate.
4. Regression Test Suite
4.1 Golden Dataset
"""Golden dataset: tập test cases "đúng chuẩn" để regression test"""
import json
from pathlib import Path
GOLDEN_FILE = "tests/golden_dataset.json"
def save_golden(test_cases: list[dict]):
"""Lưu golden dataset"""
Path(GOLDEN_FILE).parent.mkdir(exist_ok=True)
with open(GOLDEN_FILE, "w") as f:
json.dump(test_cases, f, ensure_ascii=False, indent=2)
def load_golden() -> list[dict]:
"""Load golden dataset"""
with open(GOLDEN_FILE) as f:
return json.load(f)
# Golden dataset structure
golden_dataset = [
{
"id": "test_001",
"input": "Báo cáo doanh thu tháng 9: 5.2 tỷ, tăng 15% so tháng trước",
"prompt_version": "v2.1",
"expected_output": {
"category": "report",
"metrics": [{"name": "revenue", "value": 5.2, "unit": "tỷ"}],
"trend": "tăng",
},
"created_at": "2025-01-15",
"tags": ["happy_path", "vietnamese", "metrics"],
},
# ... thêm 50+ test cases
]
4.2 Regression Runner
"""Chạy regression: so sánh prompt mới vs golden results"""
def run_regression(prompt_template: str, golden: list[dict]):
"""Chạy regression test"""
results = {"passed": 0, "failed": 0, "errors": []}
for case in golden:
try:
output = run_prompt(prompt_template, case["input"])
score = evaluate_with_llm(
case["input"], output, json.dumps(case["expected_output"]),
)
if score["overall"] >= 7:
results["passed"] += 1
else:
results["failed"] += 1
results["errors"].append({
"id": case["id"],
"score": score["overall"],
"reasoning": score["reasoning"],
})
except Exception as e:
results["failed"] += 1
results["errors"].append({"id": case["id"], "error": str(e)})
total = results["passed"] + results["failed"]
results["pass_rate"] = results["passed"] / total if total > 0 else 0
print(f"Regression: {results['passed']}/{total} "
f"({results['pass_rate']:.0%}) passed")
return results
5. Automated Test Pipeline
5.1 CI/CD Integration
# .github/workflows/prompt-tests.yml
name: Prompt Testing
on:
push:
paths: ['prompts/**'] # Chạy khi sửa prompt files
jobs:
test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Setup Python
uses: actions/setup-python@v5
with:
python-version: '3.12'
- name: Install dependencies
run: pip install openai pytest
- name: Run prompt unit tests
env:
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
run: pytest tests/test_prompts.py -v
- name: Run regression tests
env:
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
run: python tests/regression_runner.py --golden tests/golden.json
- name: Upload results
uses: actions/upload-artifact@v4
with:
name: prompt-test-results
path: tests/results/
5.2 pytest Integration
"""tests/test_prompts.py — pytest cho prompts"""
import pytest
import json
from openai import OpenAI
client = OpenAI()
PROMPT_V2 = """Phân loại email: spam, business, personal, newsletter.
Output JSON: {"category": "...", "confidence": 0.0-1.0}
Email: {input}"""
@pytest.fixture
def llm():
return client
class TestEmailClassification:
def test_spam_detection(self, llm):
result = run_prompt(PROMPT_V2, "Bạn trúng thưởng 10 tỷ!")
data = json.loads(result)
assert data["category"] == "spam"
assert data["confidence"] >= 0.8
def test_business_email(self, llm):
result = run_prompt(PROMPT_V2, "Meeting Q3 review lúc 2pm")
data = json.loads(result)
assert data["category"] == "business"
def test_output_format(self, llm):
result = run_prompt(PROMPT_V2, "Hello!")
data = json.loads(result) # Phải parse được JSON
assert "category" in data
assert "confidence" in data
assert 0 <= data["confidence"] <= 1
@pytest.mark.parametrize("input_text,expected", [
("Giảm giá 90%! Mua ngay!", "spam"),
("Báo cáo tháng 9 đính kèm", "business"),
("Cuối tuần đi café không?", "personal"),
])
def test_batch(self, llm, input_text, expected):
result = run_prompt(PROMPT_V2, input_text)
data = json.loads(result)
assert data["category"] == expected
💡 Exercise 2: Create a pytest file for one of your prompts. Write at least 5 test cases including: 2 happy path, 2 edge cases, 1 adversarial. Run
pytest -v.
Summary
| Concepts | Tools / Techniques | When |
|---|---|---|
| Unit Test | Assert output criteria | Each time you edit the prompt |
| Regression | Golden dataset comparison | Before deploy |
| Consistency | N-run variance check | New prompt |
| LLM-as-Judge | GPT-4o evaluate GPT-4o-mini | Complex Evaluation |
| CI/CD | GitHub Actions + pytest | Automatically every commit |
| Metrics | Correctness, latency, cost | Monitor continuously |
General exercises
- ✅ Complete 2 small exercises (1, 2)
- Full Test Suite: Select 1 production prompt. Write 20+ test cases (golden dataset). Includes: happy path, edge case, adversarial, multilingual. Pass rate must be >= 90%.
- Dashboard: Run test suite → export JSON results → visualize pass rate, average scores, latency. Use matplotlib or Streamlit.
- Auto-Evaluator: Build LLM-as-Judge pipeline: input → model A output → model B evaluate → aggregate scores. Compare 2 versions prompt.
Next article: Prompt Versioning, A/B Testing & CI/CD — version management, production comparison, rollback.