Explicit state, typed tools, bounded loops and a replayable trace. An agent is a distributed system, so we build it like one.
Most agent failures are not reasoning failures. They are the failures any distributed system has: a step retried and wrote twice, a loop that never terminated, a tool call that silently returned an error string, state that nobody could inspect after the fact. We build agents as state machines with explicit transitions, because that is the version you can debug at nine on a Tuesday morning when it has done something odd.
Measurements, not results. These are the readouts we put in place so that you end up with numbers about your own system. There are no values on this page because a value here would be somebody else’s.
Named as plain text. None of these is a partnership, an endorsement or a default: the right one is chosen per engagement, usually the one your team already runs.
engineering / the other five
Send the architecture and the failure you are seeing. You get a written read from an engineer within a working day.