How to ship your first AI feature in 14 days — without rebuilding your product.
A step-by-step sprint plan we use with clients: pick the right use case, ground the model in your data, ship behind a flag and measure what matters.

Start with one job
The first AI feature fails when it tries to be the product. A chat box bolted onto everything, a promise to “answer any question,” and a knowledge base nobody owns. Two weeks later the demo is impressive and the support team does not trust it.
We start smaller. One job the product already does, done often, with a clear wrong answer. Password resets, order status, a quote from a price list, a summary of a document the user just uploaded. If we cannot name the job in a sentence, we are not ready to build.
Key takeaway: pick one job the product already does. Scope is what makes a first AI feature reliable.
Fourteen days, four moves
The sprint on our wall is four moves, not fourteen features. Scope, ground, flag, ship. Days one to three are only scope: the job, the user, the failure, and the sentence the feature is allowed to say. If that sentence needs a new product, we stop.
1–3
Days to lock the job
4–8
Days to ground it
9–14
Flag it, then ship
Days four to eight are grounding: the documents, rows, or policies the answer must come from. Days nine to eleven put the feature behind a flag for a small set of real users. Days twelve to fourteen are the ship: a measured rollout, a way to turn it off, and a person who reads what it said.
Ground it before you prompt it
A clever prompt on an empty context window will invent. We do not treat that as a writing problem. We treat it as a data problem. The model only sees the passages, records, or tool results we handed it for this question, and the answer has to point at one of them.
That is why the second move is “ground,” not “tune.” Fine-tuning a voice is a later decision. The first version should be able to say “I don’t have that” and mean it. If the source is missing, stale, or belongs to another customer, the feature declines. A confident wrong answer costs more than a short hand-off.
Ship it behind a flag
We do not launch an AI feature to everyone on day fourteen. The flag is the product. A named cohort, a log of questions and answers, and a weekly review with the person who owns the job. If the feature cannot be switched off without a deploy, it is not ready.
“Fourteen days is enough to prove one job. It is not enough to rebuild the product around a model.”Mujtaba Asif, CTO
What we measure is boring on purpose: did it answer from the source, did a human have to take over, and did the user finish the job. Vanity scores on a demo set do not survive contact with Tuesday afternoon. If those three numbers are healthy for the cohort, we widen the flag. If they are not, we narrow the job before we touch the model.
Get the next build note
in your inbox.
One practical email a month. No fluff, unsubscribe anytime.


