AI features built into a product that already has customers, without destabilising the thing they are already paying for.
Price
$60k to $120k
Duration
8 to 12 weeks
How it is priced
Time and materials against a fixed capacity, billed monthly. The range moves on how many surfaces of your product are touched and how much of your data layer has to be reshaped before retrieval is worth doing.
Who is on it
A delivery lead, senior engineers working inside your codebase and your review process, and an engineer on retrieval, evaluation and cost.
Written for
Series A to C SaaS companies adding AI to a product that already ships, where the risk is regression rather than novelty.
Adding AI to a live product is a different job from building one. The constraints are your existing schema, your existing review process, your existing latency budget and a customer base that notices. We work inside your repository, behind your feature flags, against your CI, and the first thing we ship is usually the evaluation harness rather than the feature.
Stated here rather than discovered in month three. Anything on this list can be scoped separately, and we will say what it would take.
One call with the people who know the problem, then a written brief back to you within two working days: what we understood, what we think the hard part is, what is explicitly out of scope, and which of the four engagements fits. If none of them fits we say that instead of reshaping your problem to match a price list.
Output: A written brief and a named engagement, or an honest noBefore any prompt is written we build a labelled set from your real examples with your domain experts, split by category, with a held back portion. The pass threshold goes in writing, with the consequence of missing it agreed at the same time. This is the week that makes every later argument about quality a number rather than an opinion.
Output: A labelled evaluation set, a harness, and an agreed thresholdOne path through the system, end to end, running in your cloud account, usually inside the second week. Not a prototype on a laptop. Deploying early is how integration risk, credential problems and data access surprises surface while there is still time to change the plan.
Output: A deployed slice on your infrastructure and a working pipelineEvery week ends with a working demonstration against the evaluation set rather than a status document. You see the score, the failures, and what we are doing about them. Anything at risk is raised in the week it becomes at risk, not in the week it becomes a problem.
Output: Shipped increments, a score per week, and a visible risk listGuardrails on the output path, escalation payloads, tracing, cost and latency instrumentation, alerts wired to people who can act, and runbooks for the two or three ways this specific system fails. Where the system can act, it shadow runs against live traffic before it is granted authority.
Output: Guardrails, dashboards, alerts, runbooks and shadow run resultsYour engineers make the last change while we are still there to watch. Repository, infrastructure as code, evaluation suite, documentation and access all sit with you, and none of it depends on an account we control. If you want us to stay, that is an Embedded AI Pod with its own scope, not a dependency we engineered into the build.
Output: Code, IP, docs, evaluation coverage and a team that can run itYours. Branches, pull requests, your review rules, your CI. If your team cannot review what we wrote, we have built you a dependency instead of a feature.
Feature flags on everything, a separate deployment path where your architecture allows it, and evaluation gates in CI so a regression is caught by a machine rather than a customer. We also cap how much of your engineers' time we ask for and put that number in the scope, because the hidden cost of an integration is usually your team's attention.
One interface in your codebase that every AI call goes through, prompts held as versioned artefacts rather than string literals scattered through the code, provider selection as configuration, and an evaluation harness that can run the same suite against more than one provider. Switching then costs a configuration change and an evaluation run rather than a quarter.
Yes. NDAs, background checked engineers, your access controls, your tenancy, and data flow documentation written for the reviewer rather than for us. Say early that a review is coming and we scope the paperwork into the plan instead of discovering it in week ten.
engineering / the other three
You get a written view of the problem within one working day, from an engineer, including the version where the answer is that this is not a build.