Free template 04 of 05 · for founders

AI feature scoping: code, prompt, RAG or fine-tune?

Most AI features are over-engineered before they are tested. Teams reach for fine-tuning when a better prompt would do, or build a vector database for data that already lives in a SQL table. This worksheet walks one feature through five steps: describe it, choose the simplest approach that can work, estimate the effort, check the unit economics, and decide go or no-go.

All templates60–90 minutes per feature5 steps + worked examplePDF + Markdown

Take it with you

Everything on this page is free to read. Leave your email to get the print-ready PDF and a Markdown file you can paste into Notion or import into Google Docs.

  • PDFDesigned for printing and filling in by hand
  • MDEditable copy for Notion, Google Docs, Obsidian

Step 1 — Describe the feature

Input → output

What goes in (text, file, record, audio) and exactly what comes out (format, length, where it lands).

Who uses it and how often

User role, tasks per user per day or month, peak volume.

Cost of a wrong answer

Annoying (user edits it), costly (money, time, trust) or dangerous (health, legal, safety)? This sets how much review and evaluation you need.

Latency budget

Real-time (under 1 s), interactive (under 10 s) or background (minutes are fine)?

Step 2 — Choose the simplest approach that can work

Answer the questions in order. Stop at the first "yes" — that is your starting approach. You can always move up a level later; moving down is rare because nobody wants to delete work.

  1. Can the logic be written as rules, math or a lookup? → Build it in plain code. No model. It is cheaper, faster and testable.
  2. Does the base model already know enough, and the task is about reasoning, writing or extracting from the input you give it? → Prompt an API model with a clear output schema.
  3. Does it need to read or change structured data (database rows, CRM, calendar, an API)? → Prompting + tool (function) calling. Do not embed rows into a vector database.
  4. Does it need knowledge from many documents the model has not seen (help center, contracts, wiki, PDFs)? → RAG: retrieve relevant passages, then prompt with them. Show sources.
  5. Is prompting plateauing on format, style or a narrow classification — and do you have 500+ high-quality examples? → Consider fine-tuning, usually a smaller model, to cut cost or latency. Keep the prompted version as the baseline.
  6. None of the above? → Training your own model is almost never an MVP decision. Revisit the problem framing first.
ApproachUse whenAvoid whenData you needFirst version
Plain codeRules, math, lookups, routingInputs are free text with endless varietyRules and test casesDays
PromptingWriting, summarizing, extracting, classifying from given inputAnswers depend on private knowledge20–50 sample inputs for testing1–3 days
Prompt + toolsReading or acting on structured systemsYou can't give the model safe, scoped permissionsAPI access, a list of allowed actions3–10 days
RAGAnswering from a large, changing body of documentsData is small enough to fit in the prompt, or is tabularClean documents, access rules, refresh plan1–3 weeks
Fine-tuningNarrow, repetitive task where prompting plateaus; cost or latency must dropYou have under ~500 good examples or the task keeps changing500–5,000 labelled examples2–4 weeks + data work
Own modelResearch-grade problem, unique data at scaleAlmost every MVPA lotMonths
Your approach and why

Which level did you stop at, and what would make you move up a level later?

Step 3 — Estimate the effort

Ranges are working days for one experienced designer-engineer with modern AI-assisted tooling. S = narrow feature with clean inputs; M = typical; L = messy data, many edge cases or high cost of error. Cross out the rows you don't need and write your own estimate in the last column.

WorkstreamSMLYours
Prompt, output schema and examples12–35–8 
Golden set and eval harness12–35+ 
UI for AI states (streaming, edit, errors, sources)24–68–12 
Tool / API integrations (per system)1–23–58+ 
Retrieval: ingest, chunk, embed, index, refresh—4–710–15 
Guardrails: moderation, injection, permissions0.51–24+ 
Cost controls, logging, quotas0.51–23 
Fine-tuning: data cleaning, training, comparison—5–1015+ 
Total (add 20% for unknowns)    

Step 4 — Check the unit economics

Cost per task = (input tokens × input price + output tokens × output price) ÷ 1,000,000. Then cost per user per month = cost per task × tasks per user per month. Measure tokens on real inputs; take prices from your provider's current price list.

LineExample (illustrative prices)Yours
Input tokens per task (prompt + context)3,000 
Output tokens per task500 
Price per 1M input / output tokens$3 / $15 
Cost per task$0.009 + $0.0075 = $0.0165 
Tasks per user per month (p95 user)400 
AI cost per heavy user per month$6.60 
Plan price per user per month$29 
AI cost as % of price (aim under 20–25%)23% 

If the heavy-user cost is above a quarter of the price: shorten the context, cache repeated work, route easy tasks to a smaller model, or meter usage above a quota.

Step 5 — Go / no-go

Worked example — "Draft a reply" in a support inbox

Feature: for a B2B SaaS helpdesk, draft a reply to an incoming ticket that the agent edits and sends. ~120 tickets per agent per week; wrong answer is costly (customer gets wrong instructions) but always reviewed by an agent; interactive latency (under 10 s).

  • Step 2: Level 1 (rules) no — tickets are free text. Level 2 (prompting) not enough — answers depend on the product's help center (400 articles) and the customer's plan. Plan and account data → tools (look up plan, seats, last invoice). Help center → RAG with article links shown to the agent. No fine-tuning: content changes weekly.
  • Step 3: prompt + schema 2–3 d, eval set (60 historical tickets with the agent's actual reply) 3 d, UI for draft + sources + edit 5 d, 2 tool integrations 4 d, retrieval over help center 5 d, guardrails 1 d, logging and quotas 1 d → 21–22 d + 20% ≈ 5 weeks.
  • Step 4: ~4,000 input + 400 output tokens per draft → about $0.018 at the illustrative prices above; 480 drafts per agent per month → ~$8.60 per agent per month against a $49 seat → 18%. Go.
  • Step 5: success = agents send ≥ 50% of drafts with light edits within 4 weeks; fallback = "draft unavailable" and the normal editor; kill switch per workspace.

Want this done for you?

Stepikin Studio designs and builds AI products end to end — brief, UX for AI, frontend, launch — and ships an MVP in weeks, not quarters. Send what you have, even half a brief, and get a fixed estimate.

Get an estimateBook a call

More templates for founders

All templates

AI products we shipped

All projects

Create bold.
Deliver better.

See our workGet in touch