Chuyển đến nội dung chính

AI Cost & FinOps Basics for BA: Token, Latency, and Budget Control in AI Projects

Duy Tran12 min
AI Cost & FinOps Basics for BA: Token, Latency, and Budget Control in AI Projects

BA often leave "how much AI costs" to technical teams. That is a mistake. When stakeholders ask, "Is this feature worth investing in?" BA should have numbers ready, not answer "let me ask tech".


1. AI Cost Types You Need to Know

1.1 Token-based Pricing (LLM API)

Most LLM APIs charge by token (not characters, not words):

1 token ~= 4 English characters ~= 3/4 of a word
"Hello world" ~= 2 tokens
1 A4 page of text ~= 500-700 tokens

Example pricing (reference, May 2026):

ModelInputOutputWhen to use
GPT-4o mini$0.15/1M tokens$0.60/1M tokensHigh volume, simple tasks
GPT-4o$2.50/1M tokens$10/1M tokensComplex tasks, high quality
Claude Sonnet$3/1M tokens$15/1M tokensReasoning, coding, analysis
Gemini Flash$0.075/1M tokens$0.30/1M tokensLowest cost, enough for many tasks

1.2 Cost Estimation Formula

Monthly cost = Daily requests x Avg tokens/request x Cost per token x 30 days

Example: Support chatbot
- 500 requests/day
- 800 input tokens + 400 output tokens = 1200 tokens/request
- Model: GPT-4o mini ($0.15 input, $0.60 output)

Cost = 500 x [(800 x $0.15) + (400 x $0.60)] / 1,000,000 x 30
     = 500 x [($0.00012) + ($0.00024)] x 30
     = 500 x $0.00036 x 30
     = $5.40/month (very low for chatbot)

2. Latency and Compute Cost

Beyond API cost, include:

Cost TypeDescriptionWhat BA should know
Compute (GPU)Running self-hosted models$0.5-5 per GPU hour depending on type
EmbeddingVector search, RAG$0.02/1M tokens (very low)
StorageVector DB, model weights$20-100/month depending on volume
Inference latencyUser waiting time -> UX costP95 latency < 3s is commonly acceptable

3. Cloud AI API vs Self-hosted Model

CriteriaCloud APISelf-hosted
Upfront cost$0$10K-100K+ (GPU)
Variable costPay per tokenElectricity + cloud GPU
Latency0.5-3s0.1-1s (local GPU)
Data privacyData goes to vendorData stays in-house
MaintenanceNoneRequires MLOps team
Best whenMVP, low-medium volumeLarge scale, sensitive data

BA rule of thumb:

  • < 1M tokens/month -> Cloud API is almost always cheaper
  • 100M tokens/month -> Self-hosted may reach positive ROI after 6-12 months

  • Classified data (healthcare, finance) -> Self-hosted may be mandatory

4. FinOps Practices BA Can Propose

4.1 Cost Alert Setup

Ask the team to set:

Alert Level 1: Cost reaches 70% monthly budget -> Notify BA + PM
Alert Level 2: Cost reaches 90% monthly budget -> Notify leadership
Alert Level 3: Hard cap automatically disables non-critical features above 100%

4.2 Token Optimization Strategies

BA can propose these without coding:

StrategyDescriptionPotential savings
Prompt cachingReuse system prompt across requests30-50% input tokens
Model routingUse smaller model for simple tasks60-90% cost for routine tasks
Context pruningRemove unnecessary conversation history20-40% per request
Batch processingGroup non-real-time requestsUp to 50% volume discount
Output length limitCap max output tokens20-30% output cost

4.3 Cost per Business Metric

Do not track API cost alone. Track cost per business outcome:

Cost per resolved ticket = Monthly AI cost ÷ Tickets resolved by AI
Cost per qualified lead = Monthly AI cost ÷ Leads qualified by AI
Cost savings vs manual = (FTE hours saved x hourly rate) - AI cost

5. Budget Estimation Template for BA

## AI Feature Cost Estimate
**Feature:** [Name] | **Date:** [YYYY-MM] | **Owner:** [BA]

### Volume Assumptions
| Parameter | Estimate | Source |
|-----------|----------|--------|
| Daily active users | [X] | Analytics/forecast |
| Requests per user per day | [Y] | UX assumption |
| Total daily requests | [XxY] | Calculated |

### Per-Request Cost
| Component | Cost | Notes |
|-----------|------|-------|
| LLM API (input) | $[X]/1K tokens | [Model name] |
| LLM API (output) | $[X]/1K tokens | [Model name] |
| Avg tokens per request | [X] input + [Y] output | Estimated from prototype |
| **Cost per request** | **$[total]** | |

### Monthly Projection
| Scenario | Cost | Notes |
|----------|------|-------|
| Conservative (50% of estimate) | $[X] | Low adoption |
| Base case | $[Y] | Expected |
| Peak (200% of estimate) | $[Z] | Viral/spike |

### ROI Estimate
- Manual cost replaced: $[X]/month
- AI cost: $[Y]/month
- Net savings: $[X-Y]/month
- Break-even: [N] months

Conclusion

BA do not need cloud architecture expertise to estimate AI cost. Understanding token pricing, knowing when to use cloud vs self-hosted, and setting cost alerts are enough for BA to contribute to business cases and prevent budget overruns.

When stakeholders ask "How much does this AI feature cost?" you can answer within 30 minutes using this template.