Step 1 — Describe the feature
What goes in (text, file, record, audio) and exactly what comes out (format, length, where it lands).
User role, tasks per user per day or month, peak volume.
Annoying (user edits it), costly (money, time, trust) or dangerous (health, legal, safety)? This sets how much review and evaluation you need.
Real-time (under 1 s), interactive (under 10 s) or background (minutes are fine)?
Step 2 — Choose the simplest approach that can work
Answer the questions in order. Stop at the first "yes" — that is your starting approach. You can always move up a level later; moving down is rare because nobody wants to delete work.
- Can the logic be written as rules, math or a lookup? → Build it in plain code. No model. It is cheaper, faster and testable.
- Does the base model already know enough, and the task is about reasoning, writing or extracting from the input you give it? → Prompt an API model with a clear output schema.
- Does it need to read or change structured data (database rows, CRM, calendar, an API)? → Prompting + tool (function) calling. Do not embed rows into a vector database.
- Does it need knowledge from many documents the model has not seen (help center, contracts, wiki, PDFs)? → RAG: retrieve relevant passages, then prompt with them. Show sources.
- Is prompting plateauing on format, style or a narrow classification — and do you have 500+ high-quality examples? → Consider fine-tuning, usually a smaller model, to cut cost or latency. Keep the prompted version as the baseline.
- None of the above? → Training your own model is almost never an MVP decision. Revisit the problem framing first.
| Approach | Use when | Avoid when | Data you need | First version |
|---|---|---|---|---|
| Plain code | Rules, math, lookups, routing | Inputs are free text with endless variety | Rules and test cases | Days |
| Prompting | Writing, summarizing, extracting, classifying from given input | Answers depend on private knowledge | 20–50 sample inputs for testing | 1–3 days |
| Prompt + tools | Reading or acting on structured systems | You can't give the model safe, scoped permissions | API access, a list of allowed actions | 3–10 days |
| RAG | Answering from a large, changing body of documents | Data is small enough to fit in the prompt, or is tabular | Clean documents, access rules, refresh plan | 1–3 weeks |
| Fine-tuning | Narrow, repetitive task where prompting plateaus; cost or latency must drop | You have under ~500 good examples or the task keeps changing | 500–5,000 labelled examples | 2–4 weeks + data work |
| Own model | Research-grade problem, unique data at scale | Almost every MVP | A lot | Months |
Which level did you stop at, and what would make you move up a level later?
Step 3 — Estimate the effort
Ranges are working days for one experienced designer-engineer with modern AI-assisted tooling. S = narrow feature with clean inputs; M = typical; L = messy data, many edge cases or high cost of error. Cross out the rows you don't need and write your own estimate in the last column.
| Workstream | S | M | L | Yours |
|---|---|---|---|---|
| Prompt, output schema and examples | 1 | 2–3 | 5–8 | |
| Golden set and eval harness | 1 | 2–3 | 5+ | |
| UI for AI states (streaming, edit, errors, sources) | 2 | 4–6 | 8–12 | |
| Tool / API integrations (per system) | 1–2 | 3–5 | 8+ | |
| Retrieval: ingest, chunk, embed, index, refresh | — | 4–7 | 10–15 | |
| Guardrails: moderation, injection, permissions | 0.5 | 1–2 | 4+ | |
| Cost controls, logging, quotas | 0.5 | 1–2 | 3 | |
| Fine-tuning: data cleaning, training, comparison | — | 5–10 | 15+ | |
| Total (add 20% for unknowns) |
Step 4 — Check the unit economics
Cost per task = (input tokens × input price + output tokens × output price) ÷ 1,000,000. Then cost per user per month = cost per task × tasks per user per month. Measure tokens on real inputs; take prices from your provider's current price list.
| Line | Example (illustrative prices) | Yours |
|---|---|---|
| Input tokens per task (prompt + context) | 3,000 | |
| Output tokens per task | 500 | |
| Price per 1M input / output tokens | $3 / $15 | |
| Cost per task | $0.009 + $0.0075 = $0.0165 | |
| Tasks per user per month (p95 user) | 400 | |
| AI cost per heavy user per month | $6.60 | |
| Plan price per user per month | $29 | |
| AI cost as % of price (aim under 20–25%) | 23% |
If the heavy-user cost is above a quarter of the price: shorten the context, cache repeated work, route easy tasks to a smaller model, or meter usage above a quota.
Step 5 — Go / no-go
Worked example — "Draft a reply" in a support inbox
Feature: for a B2B SaaS helpdesk, draft a reply to an incoming ticket that the agent edits and sends. ~120 tickets per agent per week; wrong answer is costly (customer gets wrong instructions) but always reviewed by an agent; interactive latency (under 10 s).
- Step 2: Level 1 (rules) no — tickets are free text. Level 2 (prompting) not enough — answers depend on the product's help center (400 articles) and the customer's plan. Plan and account data → tools (look up plan, seats, last invoice). Help center → RAG with article links shown to the agent. No fine-tuning: content changes weekly.
- Step 3: prompt + schema 2–3 d, eval set (60 historical tickets with the agent's actual reply) 3 d, UI for draft + sources + edit 5 d, 2 tool integrations 4 d, retrieval over help center 5 d, guardrails 1 d, logging and quotas 1 d → 21–22 d + 20% ≈ 5 weeks.
- Step 4: ~4,000 input + 400 output tokens per draft → about $0.018 at the illustrative prices above; 480 drafts per agent per month → ~$8.60 per agent per month against a $49 seat → 18%. Go.
- Step 5: success = agents send ≥ 50% of drafts with light edits within 4 weeks; fallback = "draft unavailable" and the normal editor; kill switch per workspace.
EngiBoard
Vigilo
AI Dent