Arguments we have had often enough to be worth writing down. No borrowed benchmarks and no client numbers: if a piece needed a figure to make its point, it needed an engagement behind it and it belongs in the work section instead.
N1 · Evaluation
Without a labelled set, every change to an AI system is a guess with a demo attached, and the version that ships is the one whose author argued hardest.
6 min · 5 sectionsN2 · Retrieval
When a system answers badly it is usually reading the wrong thing. Changing the model is the most expensive way to avoid looking at your corpus.
7 min · 6 sectionsN3 · Operations
Per month is the wrong unit for AI spend. Per task is the unit, and you have to design for it before you can measure it.
6 min · 5 sectionsN4 · Architecture
Every vendor says their architecture is model agnostic. It is a testable property, and the test is how much a provider switch actually costs you.
5 min · 5 sectionsSay so. The version of this conversation where somebody pushes back on the method is the one that produces a good scope.