Services

MLOps and production ML

The unglamorous work that makes models real. Most machine learning dies after the prototype: no deployment path, no monitoring, no answer when the data drifts. We take models into production and keep them honest, with pipelines, monitoring, retraining and runbooks your engineers own.

Editorial illustration of a model deployment pipeline with monitoring and retraining loops.

The process we optimise

The process we optimise: keeping models that make live decisions healthy in production

A model earns its keep only once it is making live decisions and staying reliable while it does. That is a different discipline from building the model in the first place. Data shifts, behaviour changes, and a model that was accurate at launch quietly decays unless something is watching. The process we optimise is the operational one that surrounds a production model: how it is deployed, how its health is monitored, how it is retrained when it drifts, and how it is rolled back when something goes wrong. Done well, this is the unglamorous work that turns a promising prototype into infrastructure the business can depend on.

Before and after

What changes when the system is rebuilt

Before: a model that works until it doesn't

  • The model lives in a notebook on someone's laptop, with no repeatable path to production
  • It drifts silently, because nothing measures whether its predictions still match reality
  • Retraining happens by hand, whenever someone remembers or notices it has gone stale
  • There is no rollback, so a bad deployment means the live decision stays broken until it is rebuilt

After: production ML your team operates

  • The model deploys through a repeatable pipeline, the same way every time
  • Monitoring tracks performance and drift, and alerts a human before the decision quality slips
  • Retraining runs on a defined trigger with human sign-off, rather than from memory
  • A known-good version is always one rollback away, so a bad release is a quick reversal, not an outage

How we think about it

We start from the objective, then work outwards

The same discipline runs through every engagement: understand what the process is for, then design the system to serve it and measure against it.

01

Start from the decision the model is meant to keep making

We begin with the objective: which live decision this model drives, what good performance means for that decision, and what the cost is when it degrades. That tells us how closely it needs watching, how fast a problem must be caught, and what a safe rollback has to protect. Everything we build afterwards is sized against that, rather than against a generic best-practice checklist.

02

Map how the model reaches production today, honestly

We trace the real path from notebook to live decision, including the manual copy-paste steps, the retraining that happens by hand, and the absence of any monitoring or rollback. We name where the model could drift unnoticed and what would happen if it did. This is the honest audit that shows why the prototype has never been safe to lean on.

03

Design the operational system around that workflow

We build deployment pipelines, monitoring, alerting, retraining and rollback around how your team actually ships and operates, using your stack. Retraining is automated but gated by human sign-off, so nothing goes live unattended. The design goal is that your own engineers can run and reason about the system, not that it depends on us to keep breathing.

04

Measure against the objective and the decision it protects

We test the finished system against the objective from step one. Does monitoring catch a real drift before the decision degrades? Does retraining restore performance on the metric that matters? Does a rollback return to a known-good state quickly? We measure against whether the live decision stays healthy, because that, not pipeline elegance, is the point of the work.

An engagement, step by step

Representative walkthrough

The following is a representative eight to sixteen week arc, not an account of a specific client. It shows the shape production ML work typically takes when a model already exists but has never been safe to run live.

  1. Weeks 1 to 4

    Diagnose the model and the decision it drives

    We establish which live decision the model serves, what good performance means for it, and how it currently reaches production. We map where it could drift unnoticed and where a bad release would have no way back.

  2. Weeks 4 to 8

    Build the deployment pipeline

    We move the model out of the notebook and into a repeatable deployment pipeline, so it ships the same way every time. This is the point at which going live stops being a manual, error-prone event.

  3. Weeks 8 to 12

    Add monitoring, drift detection and rollback

    We instrument the model so performance and drift are measured continuously and a human is alerted before decision quality slips. We put a rollback in place so a known-good version is always one step away.

  4. Weeks 12 to 16

    Automate retraining and hand over

    We wire up retraining on a defined trigger with human sign-off, then hand the estate to your team with runbooks. Your engineers operate the monitored, retrainable model themselves, rather than depending on us to keep it alive.

The engagement ends with monitored, retrainable production ML that the client's own team operates, judged on whether the live decision the model drives stays healthy without heroics.

What you get

  • Deployment pipelines (CI/CD) for your models
  • Monitoring and alerting on performance and drift
  • Automated retraining with human sign-off
  • Runbooks and handover so your team owns the estate

How it typically runs

Typical engagement: 8 to 16 weeks, 2 engineers, rolling retainer after go-live.

1

Diagnose

Weeks 1 to 4

2

Prototype pipeline

Weeks 4 to 8

3

Build and harden

Weeks 8 to 12

4

Productionise and monitor

Weeks 12 to 16

Indicative timeline. Every project is scoped individually: book a discovery call and we will provide a detailed proposal within 48 hours.

Model-neutral by default

We evaluate Claude, GPT and open-weights models head to head for every problem, and recommend whichever wins on cost, latency and accuracy.

See how the vendor-led alternatives compare

Frequently asked questions

We already have a working model. Why is this a separate piece of work?
Because building a model and operating it in production are different disciplines. A model that scores well in a notebook can still drift silently, have no rollback and need retraining by hand. This work is about making the live decision reliable over time, which the prototype was never designed to guarantee.
Does automated retraining mean models change without anyone knowing?
No. Retraining runs on a defined trigger but is gated by human sign-off, so a person approves before a new version goes live. Automation removes the manual toil, not the human judgement.
Who runs the system after you leave?
Your team does. We build on your stack and hand over runbooks and monitoring your engineers can reason about, so the estate is yours to operate. The rolling retainer some clients keep is for depth on tap, not because the system cannot run without us.

Ready to talk?

Get in touch. We will discuss your challenge and show you what is possible.