Evals before features: shipping AI that doesn't guess
We write the test conversations before we write the prompt. It's slower for a week and much faster for a year.
(02) — The note
Written by the people who did the work. Numbers come from client analytics, shared with permission.
An AI feature is only as good as the worst answer it gives in front of a customer. So we start with the answers.
Write the exam first
For Quillmate we collected 412 real questions from support tickets and wrote the answer a good agent would give, with the page it should cite.
Run it every night
Every prompt change, model change and retrieval tweak runs the full suite. If the grounded-answer score drops, the change doesn’t ship.
Know when to hand off
The best answer is sometimes ‘let me get a person’. We test for that too.
(03) — Keep reading
More from the journal
Process
7 min
How we quote a build in two weeks
A fixed quote needs a fixed scope. Here is the discovery sprint we run before we name a price, and what you get at the end of it.
Case notes
6 min
The checkout that paid for itself in a month
Fernloop's new checkout lifted conversion from 3.1% to 4.6%. The design was the easy part. Here is what made the numbers move.
AI
9 min
Evals before features: shipping AI that doesn't guess
We write the test conversations before we write the prompt. It's slower for a week and much faster for a year.


