Shadow and canary releases for machine learning models

Analytics charts on a laptop screen

How shadow and canary releases work for machine learning models: traffic splitting, the metrics to watch, drift, safe rollback and a model release checklist.

Bhaskar Bhatt8 min read

A shadow release runs a new machine learning model on real production traffic without showing its predictions to anyone, so you can compare it with the current model safely. A canary release for machine learning sends a small share of real users to the new model, watches business and model metrics, and widens the rollout only if they hold. Most teams should do both, in that order.

On this page
  1. What is a shadow release for an ML model?
  2. How does a canary release work for machine learning?
  3. Shadow, canary, A/B or blue-green: which should you use?
  4. How should you split traffic for a model canary?
  5. Which metrics should you watch during a release?
  6. How do you roll back a model safely?
  7. What should a model release checklist include?
  8. When isn’t a staged release the right approach?
  9. Frequently asked questions

Offline evaluation tells you how a model performs on historical data, not how it behaves on today’s traffic inside today’s systems. Plenty of models that looked better on a test set have made things worse in production because of a feature computed differently at serving time, a slow dependency or a shift in user behaviour. Shadow and canary releases are how you find that out on a small scale, with a fast way back.

What is a shadow release for an ML model?

In a shadow release, every request (or a sample) goes to both the current model and the candidate. Users only ever get the current model’s answer. The candidate’s predictions, latency and errors are logged so you can compare them. It tests the model on real inputs with zero user impact, but it can’t measure how users react.

Shadow mode is the right first step for most model changes, especially retrained models, new architectures or new feature pipelines. It answers practical questions before any user is exposed. Does the candidate return predictions for every request, or does it fail on inputs the test set never contained? Is it within the latency budget under real load? How often does it disagree with the current model, and on what kind of input?

Disagreement is the most useful number here. Where the two models agree, there’s little risk either way. Where they disagree, pull a sample and have someone who knows the domain look at it. If the candidate is right in most disagreements, that’s strong evidence. If you can’t tell who’s right, you’ve learned your evaluation needs better labels before you go further.

Two cautions. Shadow traffic doubles inference load, so plan the capacity or sample a fraction of requests. And make sure the shadow path has no side effects: it must not send emails, write to customer records or call paid third-party APIs on the candidate’s behalf.

How does a canary release work for machine learning?

A canary routes a small, controlled share of live traffic to the new model while the rest stays on the current one. You compare the two groups on agreed metrics for a set period, then step the share up, hold, or roll back. Unlike shadow mode, users see the new model’s output, so you measure real outcomes.

Google’s SRE workbook chapter on canarying releases describes canarying as a partial, time-limited deployment of a change together with its evaluation. That second half matters. A canary isn’t just a slow rollout. It’s an experiment with a control group and a decision at the end.

For ML models, two details make canaries harder than for ordinary code. First, many model outcomes arrive late. A fraud model’s mistakes may only show up when chargebacks come in weeks later, and a recommendation model’s effect on retention takes even longer. Second, the effect of a model change is often small and noisy, so a canary at a tiny traffic share may need a long time to show a real difference. Plan the canary duration around how quickly your key outcome becomes measurable.

Shadow, canary, A/B or blue-green: which should you use?

Use shadow mode to check correctness, stability and latency with no user risk. Use a canary to confirm real-world impact on a small group with fast rollback. Use a full A/B test when you need statistically solid evidence for a product decision. Blue-green deployment switches all traffic at once and suits infrastructure changes more than model changes.

StrategyUsers see new model?Best forLimits
ShadowNoCatching serving bugs, latency problems and unexpected disagreement before exposureCan’t measure user response; extra inference cost
CanaryA small shareConfirming real outcomes with limited exposure and quick rollbackSmall samples make small effects hard to detect
A/B testA defined split, often 50/50Product decisions that need statistical evidenceExposes more users for longer; needs proper experiment design
Blue-greenAll at onceSwapping serving infrastructure with an instant switch backNo gradual exposure, so model issues hit everyone

These aren’t mutually exclusive. A typical path for a significant model change is shadow, then a small canary, staged increases and full rollout. A/B tests come in when the business wants to know the size of the improvement.

How should you split traffic for a model canary?

Split by a stable identifier such as user ID or account ID, not per request, so each user consistently gets one model. Start small, increase in planned steps with a hold period at each, and exclude groups where a mistake would be costly. Record which model served every prediction so results can be attributed correctly.

Per-request random splitting looks simpler but causes trouble. A user who sees recommendations from two models in one session gets an inconsistent experience, and you can’t cleanly attribute their behaviour to either model. Hashing a stable ID into buckets keeps assignment consistent and repeatable.

Think about who lands in the canary. If your traffic includes a few very large customers, a random split can put one of them in the canary group and dominate the results. Many teams exclude high-value or regulated accounts from early stages, or run canaries per region or segment. Whatever you choose, write it down before starting, so nobody redefines the groups after seeing the numbers.

Which metrics should you watch during a release?

Watch three layers: system health (latency, error rate, timeouts, resource use), model behaviour (prediction distribution, confidence, input feature distributions, disagreement with the old model), and business outcomes (conversion, approvals, escalations, complaints). Set rollback thresholds for each layer before the release starts, not while you’re looking at a dashboard.

  • System health is the fastest signal. A spike in errors or latency should trigger automatic rollback without a meeting.
  • Model behaviour catches problems before outcomes arrive. If the new model suddenly approves far more applications or predicts one class much more often than the old one, investigate, even if no one has complained.
  • Business outcomes are what you actually care about, but they’re slower and noisier. Compare the canary group with the control group over the same period, not with last month.

Feature drift belongs in the second layer. If the distribution of an input feature in production differs from what the model saw in training, predictions can degrade quietly. Monitoring feature distributions against a training baseline, and alerting when they move beyond an agreed threshold, catches both upstream data changes and genuine shifts in the world. That monitoring should stay on after the release, not just during it.

How do you roll back a model safely?

Keep the previous model deployed and warm until the new one has been at full traffic long enough to trust, and make rollback a routing change, not a redeployment. Version models, feature pipelines and configuration together, so reverting one doesn’t leave it paired with an incompatible other. Practise rollback before you need it.

The most common rollback failure we see is a model that can’t be separated from its features. The new model expects a new feature, the feature pipeline was updated in place, and the old model now gets inputs it was never trained on. Treat the model, its feature definitions and its preprocessing code as one versioned unit. A model registry plus versioned feature transformations makes this manageable; a shared folder of model files does not.

Automate the easy decisions. If error rate or latency crosses a threshold, the router should switch back on its own and page someone. Leave the judgement calls, like a small drop in a business metric, to a named person with authority to decide. If your team already runs progressive delivery for application code, the same tooling often extends to models. Our post on building a CI/CD pipeline for enterprise teams covers the pipeline side.

What should a model release checklist include?

It should confirm the model passed offline evaluation, the serving path matches training, shadow results were reviewed, traffic split and rollback thresholds are written down, monitoring is live, and someone owns the decision at each stage. A checklist turns a model release from a judgement made under pressure into a routine.

  1. Offline evaluation passed against the current model on the agreed test set, with results by segment, not just overall.
  2. Training and serving checked for skew. Features computed at serving time match training for a sample of real requests.
  3. Shadow run reviewed. Error rate, latency and a sample of disagreements checked by a domain expert.
  4. Canary plan written. Split key, starting share, step sizes, hold periods and excluded groups.
  5. Thresholds agreed. Automatic rollback triggers for system health; decision criteria for model and business metrics.
  6. Monitoring live. Dashboards and alerts for all three metric layers, including feature drift, split by model version.
  7. Rollback tested. The previous model is warm, and switching back has been tried in staging.
  8. Owners named. One person approves each stage, and the on-call engineer knows a release is in progress.
  9. Release recorded. Model version, data version, results and decisions logged for later audit.

When isn’t a staged release the right approach?

When the model’s predictions are reviewed by a person before anything happens, or when traffic is too low for a canary to show anything in a reasonable time. In those cases, careful offline evaluation plus a short shadow period, followed by a full switch with rollback ready, is usually enough.

Staged releases also add little for batch models that score a dataset once a week. There, compare the new model’s output with the old model’s on the same batch, review the differences and then switch. Save the full shadow-then-canary process for real-time models where mistakes reach customers directly or are expensive to undo.

Setting up the routing, monitoring and registry for this is a one-time investment that pays back on every later release. It’s a core part of our MLOps and AI infrastructure work, and our AI and ML engineering team builds the models and evaluation that feed into it.

Frequently asked questions

How long should a model canary run?

Long enough for your key outcome to become measurable and to cover normal daily and weekly patterns. For metrics that arrive within minutes, a few days per stage may be enough. For outcomes that take weeks, such as fraud losses, rely more on shadow results and model-behaviour metrics during the canary.

Can you canary a large language model change the same way?

Yes. Prompt, model version and retrieval changes can be canaried by routing a share of users to the new configuration. Output quality is harder to measure automatically, so pair the canary with sampled human or rubric-based review and watch user signals like edits, retries and escalations.

What tools do teams use for model canaries?

Traffic splitting is often handled by the serving platform, a service mesh or a feature-flag system. Model registries track versions, and monitoring tools track drift and metrics. The specific tools matter less than versioning the model with its features and keeping rollback a routing change.

Scroll to Top