Chuyển đến nội dung chính

Model & Prompt Evaluation Protocol: How BA Should Evaluate AI Output

Duy Tran14 min
Model & Prompt Evaluation Protocol: How BA Should Evaluate AI Output

When Dev says "the model reached 89% accuracy," BA should immediately ask: 89% on which dataset? Labeled by whom? Is it representative of production data? Without a clear protocol, 89% may mean nothing.


1. Why BA Need an Evaluation Protocol

Model evaluation is not only an ML engineer task. BA must participate because:

  • Acceptance criteria must be verified using a pre-agreed methodology
  • Bias detection requires domain context BA often has while ML engineers may not
  • Business context determines which errors are acceptable (false positive vs false negative)

2. Designing the Evaluation Test Set

2.1 Test Set Principles

PrincipleDescriptionExample
RepresentativeDistribution similar to production60% routine / 30% edge cases / 10% rare
IndependentNo overlap with training dataUse newest period not used in training
Labeled by domain expertsNot self-labeledBA + SME co-label
Sufficient sizeLarge enough for statistical significanceMinimum 200 cases/class

2.2 Test Set Composition Template

## Test Set: [AI Feature Name] v[X]

### Distribution Plan
| Category | Count | % | Source |
|----------|-------|---|--------|
| Normal cases | 300 | 60% | [system/period] |
| Edge cases | 150 | 30% | [curated by BA] |
| Adversarial | 50 | 10% | [red-team results] |
| **Total** | **500** | **100%** | |

### Label Protocol
- Labeler 1: [BA Name] (domain)
- Labeler 2: [SME Name] (subject matter)
- Conflict resolution: [Process when labels disagree]
- Inter-annotator agreement target: >= 0.85 (Cohen's Kappa)

3. Evaluation Criteria Framework

3.1 Choose the Right Metrics by Problem Type

Problem TypeMetrics BA Should RequestWhen Critical
Binary classificationAccuracy, Precision, Recall, F1When class imbalance exists
Multi-classMacro F1, Confusion MatrixWhen all classes are equally important
Generation (chatbot)BLEU, ROUGE, Human evalWhen text quality matters
Ranking/RetrievalNDCG, MRRWhen order matters
Business metricHuman override rate, first-time-right rateAlways add these

3.2 False Positive vs False Negative Trade-off

BA should decide this with stakeholders:

False PositiveFalse Negative
DefinitionAI says "yes" but reality is "no"AI says "no" but reality is "yes"
Example (fraud detection)Blocks valid transactionMisses fraudulent transaction
Business costCustomer complaints, lost revenueFinancial loss, reputational risk
Reduce first?When CX is more importantWhen risk/safety is more important

4. Evaluation Scoring Rubric

For AI generation (chatbot, summarization), use a rubric in addition to automated metrics:

## Evaluation Rubric: [Chatbot Feature]

### Dimension 1: Accuracy (0-5)
- 5: Fully accurate information
- 4: Accurate with 1-2 minor errors
- 3: Mostly correct but with notable mistakes
- 2: Many inaccuracies or important omissions
- 1: Mostly wrong or irrelevant
- 0: Complete hallucination

### Dimension 2: Relevance (0-5)
[Similar scale]

### Dimension 3: Safety (0/3/5 — non-linear)
- 5: Fully safe
- 3: Includes warning-level issues but no harm
- 0: Contains harmful content -> AUTO FAIL

### Overall Score = (D1 x 0.4) + (D2 x 0.3) + (D3 x 0.3)
### Pass threshold: >= 3.5 out of 5

5. Go/No-Go Framework

## Evaluation Sign-off: [Feature] [Date]

### Metric Results
| Metric | Target | Actual | Pass? |
|--------|--------|--------|-------|
| Accuracy on test set | >= 87% | [X]% | ☐ |
| F1 Score (edge cases) | >= 0.75 | [X] | ☐ |
| Human eval score | >= 3.5/5 | [X] | ☐ |
| False negative rate | <= 5% | [X]% | ☐ |
| Red-team: no critical failure | 100% pass | [X]% | ☐ |

### Decision
- [ ] **GO** — All metrics pass, proceed to UAT
- [ ] **CONDITIONAL GO** — [X] metrics pass, [Y] requires post-launch monitoring
- [ ] **NO-GO** — [Specific reason], requires [Action] before re-evaluation

### Sign-off
- BA: _________________ Date: _______
- PM: _________________ Date: _______
- Tech Lead: __________ Date: _______

6. Common Mistakes in Model Evaluation

Mistake 1: Evaluating on training data
-> Causes undetected overfitting; BA must request a separate test set

Mistake 2: Using only one metric
-> 90% accuracy may still mean minority class recall = 0%

Mistake 3: No human baseline
-> "AI got 85%" compared to what? How accurate are human annotators?

Mistake 4: No cohort-based testing
-> Overall accuracy looks fine while customer group X is unfairly impacted


Conclusion

An evaluation protocol is a contract between BA and engineering that defines when an AI feature is good enough to release. This document should exist BEFORE model development starts, not after build completion.