A multi step workflow that currently runs on people, rebuilt as a system that acts, checks itself, and hands off cleanly when it should not act.
Price
$150k to $300k
Duration
12 to 16 weeks
How it is priced
Phased, with a go or no go at the end of the design phase. The range moves on the number of systems the workflow has to write to, the compliance surface, and how much of the process exists only in people's heads at the start.
Who is on it
A delivery lead, senior engineers on orchestration and integration, an engineer on evaluation and guardrails, and a designer on the human review surfaces.
Written for
Mid market and enterprise operations owners where a process spans several systems, runs at volume, and cannot be automated with rules because the inputs are messy.
An agent that can act is a different risk profile from a tool that can answer. The work is less about reasoning and more about authority: what the system is allowed to do without asking, what it must escalate, how a human sees why it decided what it decided, and what the audit trail looks like when someone asks in nine months. We design the escalation path before the happy path, because the escalation path is what determines whether the system is allowed to stay switched on.
Stated here rather than discovered in month three. Anything on this list can be scoped separately, and we will say what it would take.
One call with the people who know the problem, then a written brief back to you within two working days: what we understood, what we think the hard part is, what is explicitly out of scope, and which of the four engagements fits. If none of them fits we say that instead of reshaping your problem to match a price list.
Output: A written brief and a named engagement, or an honest noBefore any prompt is written we build a labelled set from your real examples with your domain experts, split by category, with a held back portion. The pass threshold goes in writing, with the consequence of missing it agreed at the same time. This is the week that makes every later argument about quality a number rather than an opinion.
Output: A labelled evaluation set, a harness, and an agreed thresholdOne path through the system, end to end, running in your cloud account, usually inside the second week. Not a prototype on a laptop. Deploying early is how integration risk, credential problems and data access surprises surface while there is still time to change the plan.
Output: A deployed slice on your infrastructure and a working pipelineEvery week ends with a working demonstration against the evaluation set rather than a status document. You see the score, the failures, and what we are doing about them. Anything at risk is raised in the week it becomes at risk, not in the week it becomes a problem.
Output: Shipped increments, a score per week, and a visible risk listGuardrails on the output path, escalation payloads, tracing, cost and latency instrumentation, alerts wired to people who can act, and runbooks for the two or three ways this specific system fails. Where the system can act, it shadow runs against live traffic before it is granted authority.
Output: Guardrails, dashboards, alerts, runbooks and shadow run resultsYour engineers make the last change while we are still there to watch. Repository, infrastructure as code, evaluation suite, documentation and access all sit with you, and none of it depends on an account we control. If you want us to stay, that is an Embedded AI Pod with its own scope, not a dependency we engineered into the build.
Output: Code, IP, docs, evaluation coverage and a team that can run itThe design phase ends with the process specification, the authority model, the connector plan and a build estimate against them. If any of those says the workflow is not ready to be automated, you stop there having paid for the design phase, and you keep the specification. That is a better outcome than a build that has to be switched off in month six.
Every run writes an ordered record: the inputs it read, the retrieval results it used with their versions, the step by step state, the confidence at each decision point, the action taken, and any human override. It exports as one record per case. This is the part that takes the longest to build and the part that decides whether the system survives its first serious question.
Connectors validate their contracts and fail loudly rather than silently writing wrong data. Contract breaks page a human and pause the affected step rather than the whole workflow. That behaviour is designed and tested during the build, not discovered afterwards.
It can run inside your cloud tenancy with data never leaving your perimeter, which is how comparable work has been delivered before. Fully air gapped changes the model options substantially, so raise it in the first conversation and we will tell you what is and is not possible rather than finding out in week nine.
engineering / the other three
You get a written view of the problem within one working day, from an engineer, including the version where the answer is that this is not a build.