Models that keep working after launch

Our MLOps services give every model and LLM feature versioned releases, a tested rollback and monitoring for drift, latency and cost. Changes go through written change control, so you always know which version is live and why.

  • Model registry and versioning
  • Drift, latency and cost alerts
  • Runbooks and written service levels
What we do

The plumbing that makes a model a service

A model in a notebook is an experiment. A model in production needs the same discipline as any other software, plus a few things normal software doesn't: data that shifts under it, results that degrade without throwing errors, and bills that scale with usage.

CI/CD for models

Pipelines that test code, validate data, train, evaluate against the current production model and only promote a new version when it wins. Built on GitHub Actions, GitLab CI or Cloud Build. Our CI/CD guide covers the software side.

Experiment tracking

Every training run logged with its code version, data snapshot, parameters and metrics, using MLflow, Weights & Biases or your cloud's equivalent, so "which settings gave us that result?" has an answer.

Model registry and versioning

One place that records which model version is in staging and in production, who approved it, and what it was evaluated on. Rolling back is a single, tested step.

Reproducible training

Pinned dependencies, containerised training jobs, versioned datasets and fixed random seeds where they matter. Anyone on the team can rebuild last month's model.

Monitoring

Input drift, prediction drift, latency, error rates and cost per prediction, with alerts routed to a named owner rather than a dashboard nobody opens.

Infrastructure as code

Cloud resources on Google Cloud or AWS defined in Terraform or a similar tool, reviewed like code, with separate environments for development, staging and production.

How do you release a new model version safely?

Run it in shadow first, then canary. In shadow mode the new model scores real traffic but its answers aren't used, so you compare it with production at no risk. A canary then sends a small share of live traffic to it, with automatic rollback if agreed metrics get worse.

Shadow mode is the right default for anything touching money, safety or customers' access to a service. For lower-risk models, a straight canary with a short observation window is usually enough. Either way, the rollback path is tested before it's needed.

What should you monitor once a model is live?

Monitor the inputs, the outputs, the service and the bill. Models rarely crash. They get quietly worse as the world changes, so input and prediction drift matter as much as uptime.

SignalWhat it tells youTypical response
Input driftThe data coming in no longer looks like the training data: new products, a changed form, a new customer segment.Investigate the source, then retrain or adjust features.
Prediction drift and outcome metricsThe model's outputs are shifting, or, once ground truth arrives, its accuracy is falling.Compare with the baseline, retrain, or roll back to the previous version.
Latency and errorsThe service is slow or failing, often after a traffic spike or dependency change.Scale, cache or fix the dependency. Handled as a service incident.
Cost per predictionCompute or token spend is rising faster than usage.Batch requests, use a smaller model for easy cases, cache repeat answers, tighten budgets.

LLMOps: running LLM features in production

LLM features change more often than classical models, because a prompt edit is a release. They need their own versioning, testing and cost controls.

  • Prompt and configuration versioning. Prompts, model names, temperature and retrieval settings live in version control and are deployed like code, not edited live in a console.
  • Evaluation sets. A written set of real inputs with expected behaviour, scored automatically where possible and by a reviewer where not, rerun on every prompt or model change. See our GenAI solutions for how these are built.
  • Token cost budgets. Limits per request, per user and for each month, with alerts before the invoice arrives. Smaller models handle simple requests; larger ones only where evaluation shows they're needed.
  • Guardrails. Input filtering, output checks for format and policy, and refusals for out-of-scope requests. Our AI governance and security work sets the policy these enforce.
  • Tracing. Each request logged through retrieval, tool calls and model responses, so a bad answer can be traced to a missing document, a weak prompt or the model itself.

GPU or CPU, managed or self-hosted?

Default to CPUs and managed services, and move only when measurements say so. Most teams overspend on GPUs they don't keep busy, and on clusters they don't have the people to run.

ChoiceChoose it whenAvoid it when
CPU servingClassical models, small neural networks, batch scoring, quantised models with modest trafficLarge deep learning or self-hosted LLMs with tight latency targets
GPU servingLarge vision, audio or language models, high-throughput inference you can keep busyTraffic is low or spiky and the GPU would sit idle most of the day
Managed (Vertex AI, SageMaker, Cloud Run, hosted model APIs)Small team, early product, uneven traffic, you want to ship this quarterSteady high volume where the managed premium outweighs engineering time, or data must stay on your own hardware
Self-hosted (Kubernetes, your own GPUs)Predictable heavy load, strict data residency, a team able to run it on callNobody on your side can operate it after handover

Security and access

Least-privilege service accounts for training and serving, secrets in a managed secret store rather than in code, separate projects or accounts per environment, and audit logs for who changed what. Training data containing personal information gets access controls at the column or bucket level,.

Service levels and on-call

We're ISO/IEC 20000-1 certified for IT service management, and deployed models run under the same processes: incidents logged and reviewed, changes approved with a rollback plan, and service levels written down. Our office hours are 09:00–18:00 Nepal Time (UTC+5:45), Monday to Friday. If you need cover outside those hours, we agree it explicitly in the SLA before go-live rather than assume it.

How an engagement runs

Assess what's live, then automate the riskiest step first

  1. Discovery call

    A free 30-minute call, then a scoped proposal within two working days.

  2. Assessment

    One to three weeks reviewing what's in production, how it's trained and released, what's monitored and what the cloud bill looks like. You get a prioritised list, with the riskiest manual step at the top.

  3. Foundations

    Infrastructure as code, a model registry, experiment tracking and a first automated pipeline for one model.

  4. Monitoring and safe releases

    Drift, latency and cost monitoring with alerts to named owners, then shadow and canary releases.

  5. Run or hand over

    Runbooks and training for your team, or we run the platform for you under an agreed SLA.

How do engagements work, and how fast can you start?

MLOps work usually starts with a scoped assessment and foundations project, then continues as a small team that runs the platform.

ModelTeam sizeStart timeNoticeGood for
Dedicated team3–15+ engineers2–4 weeks30 days' written noticeRunning and improving the ML platform for several models over time
Staff augmentation1–5 engineers5–10 working days30 days' written noticeAdding an MLOps or platform engineer to your existing ML team
ProjectScoped to the work1–3 week discoveryEnds on milestone acceptanceA defined assessment plus foundations: registry, pipelines and monitoring for one model
PoC or MVPSmall squad4–8 week buildEnds on milestone acceptanceTaking one prototype model or LLM feature to a monitored first release

How pricing works. We quote after a free 30-minute discovery call and send a written proposal within two working days. The price depends on the seniority mix, team size, how long the engagement runs, the stack, and any compliance scope. Full details on engagement models.

When we're not the right fit

If you have one model that's retrained twice a year, you probably don't need an MLOps platform. A documented manual process and a basic monitor will do, and we'll help you write it in a few days. If you need a team on call 24 hours a day with response times in minutes from day one, ask us how that would be staffed and priced before assuming it's included. And if your model hasn't yet shown it's worth running, spend the money on model work first.

Related work

Case study

Certifyi AI SaaS GRC platform

We delivered Certifyi from product architecture to deployment. The platform runs on AWS with Docker, CI/CD pipelines in GitHub Actions, centralised logging, metrics dashboards and alerting, alongside role-based access control, encryption and audit logging.

Read the Certifyi case study

Research unit

Dignep AI Research Unit (DAIR)

DAIR, our applied AI research wing, covers the full model lifecycle, from data collection and model research to MLOps, compliance and monitoring, with a focus on moving research prototypes into production.

About DAIR

Questions about MLOps

What do MLOps services include?

MLOps services cover everything that keeps a model useful after its first release: automated training and deployment pipelines, experiment tracking, a model registry with versioning, safe rollouts, monitoring for drift, latency and cost, and the infrastructure, access control and on-call process around it. For LLM features it also covers prompts, evaluation sets and token budgets.

When do we need MLOps?

Once a model affects real decisions or customers and will need retraining. A single model refreshed twice a year can live with a documented manual process. When you have several models, frequent data changes, or an LLM feature whose prompts change weekly, manual steps start causing outages and nobody can say which version is live.

Is LLMOps different from MLOps?

The goals are the same and the details differ. With LLMs you version prompts and retrieval settings as well as models, evaluate against a written set of test cases rather than one accuracy number, set token budgets per request and monthly, add guardrails on inputs and outputs, and trace each call through retrieval and tool use.

Should we use managed ML services or self-host?

Start managed. Vertex AI, SageMaker, Cloud Run and hosted model APIs let a small team ship without running clusters. Self-hosting on Kubernetes or your own GPUs pays off when usage is steady and high enough that the managed premium exceeds an engineer's time, or when data residency rules require it.

Do we need GPUs to serve our models?

Often not. Classical models and many small neural networks serve comfortably on CPUs, especially with batching or quantisation. GPUs are worth it for large deep learning models, self-hosted LLMs and high-throughput image or audio work. We measure latency and cost per thousand predictions on both before recommending either.

Not sure which model version is live right now?

That's the most common starting point. Book a 30-minute call and we'll go through what you run today and the first thing we would automate.

Scroll to Top