A labelled set, a threshold agreed in writing, and a gate in CI. Built before the feature, not after the first complaint.
Without evaluation, every change to an AI system is a guess with a demo attached. The team argues about whether the new prompt is better, someone tries four examples, and the version that ships is the one whose author argued hardest. An evaluation harness turns that into a number that either went up or down. It is the single practice that separates a system you can maintain from one you can only admire.
Measurements, not results. These are the readouts we put in place so that you end up with numbers about your own system. There are no values on this page because a value here would be somebody else’s.
Named as plain text. None of these is a partnership, an endorsement or a default: the right one is chosen per engagement, usually the one your team already runs.
engineering / the other five
Send the architecture and the failure you are seeing. You get a written read from an engineer within a working day.