ML Model Rollback
An ML model rollback restores the last validated release, but the previous model behaves as it did only if its features, preprocessing, threshold and serving runtime come back with it. The guide below covers the pin manifest, the choice between rollback, roll through, retraining and fallback, drift triggers calibrated by batch size and how to rehearse the path.
Technically reviewed by Victor Sineglazov, D.Sc.
- Roll back a manifest, not a model file Features, preprocessing, threshold, container digest, lockfile and any hosted LLM version are pinned with the artifact, because restoring weights alone creates a configuration nobody evaluated.
- Test whether the release caused the problem An automatic rollback on severe drift stops the damage, then onset against promotion, input versus prediction shift and how the previous version scores today's traffic decide whether it stays live.
- Calibrate PSI and KS cutoffs by batch size PSI and KS statistics both fire on noise as batches shrink, while KS p-values flag trivial gaps on large batches, so each statistic gets its own cutoff, set by replaying known-good history at production batch size.
- Use canary and shadow stages to keep rollback cheap A canary bounds the damage of a bad release and shadow traffic tests the full manifest before promotion, while automatic rollback works only when per-version alarms are wired to it.
- Rehearse the path before you need it Restore the previous version from its manifest alone in staging, replay recorded traffic, time the switch and check downstream consumers on every promotion.
In short: an ML model rollback returns a production prediction service to the last version that passed validation after a new release misbehaves. The previous model behaves as validated only if its feature definitions, preprocessing code, decision threshold, serving runtime and dependencies come back with it, so the rollback target is a pinned manifest rather than a model file. A rollback also fixes only problems the release caused. When the input data itself has changed, the previous model sees the same new inputs, and a roll through, a retrain or a fallback is the better move. Two failure points recur: uncalibrated PSI cutoffs misfire on small batches, where an unshifted 100-row batch crosses 0.10 roughly two times in five in our worked illustration, and a rollback target is proven restorable only when a rehearsal starts the previous version from its manifest alone.
What Does an ML Model Rollback Actually Restore?
Rollback is a production test in its own right. Google's ML Test Score rubric lists "Serving models can be rolled back." among its ML infrastructure tests and calls a model rollback procedure a key part of incident response.
A model release is rarely self-contained. In Hidden Technical Debt in Machine Learning Systems, Sculley and colleagues write: "We refer to this here as the CACE principle: Changing Anything Changes Everything." Behavior lives in the weights and in everything that shapes their inputs and outputs.
Registries capture only part of that state. The MLflow Model Registry documentation says "Each registered model version is linked to the MLflow run, logged model or notebook that produced it, enabling full reproducibility." That link records how a model was trained, not which feature pipeline, threshold or container served it. MLflow now marks deployable versions with aliases, so a rollback there reassigns an alias to an older version number, and moving a pointer is a rollback only when everything the pointer depends on is also addressable.
Orchestration has the same blind spot. The Kubernetes Deployments documentation notes that "A Deployment's revision is created when a Deployment's rollout is triggered." and that only a change to the Pod template triggers one. A threshold or feature-store pointer delivered through a mounted file or ConfigMap outside the Pod template is not restored by kubectl rollout undo. Versioning such config under content-hashed names referenced from the Pod template brings it under the revision history.
The Pin Manifest: What a Rollback Target Must Include
A pin manifest is one versioned record per model release that names every input to the release's behavior by an immutable identifier, such as a content hash, an image digest, a commit or a dated snapshot. The rollback target is the whole manifest. The table is Pharos Production's engineering framing, not a published standard.
| Item to pin | Why it must be pinned | Typical failure mode |
|---|---|---|
| Model artifact, by content hash and registry version | Evaluation results belong to exact bytes, not to a name | The alias points to a re-uploaded artifact that was never evaluated |
| Feature pipeline and feature store version | The model learned features as they were computed at training time | The old model gets features built by new definitions and scores wrong silently |
| Preprocessing code, encoders and vocabularies | Serving must transform inputs exactly as training did | A newer category encoding shifts indices, causing silent skew |
| Decision threshold and calibration | Both are fitted to one model's score distribution | The previous model runs with the newer threshold and approval or alert rates move |
| Serving container digest and runtime | Framework and native libraries affect loading and numeric behavior | A mutable tag resolves to a rebuilt image and the old artifact fails to load |
| Training data snapshot and config | The release must stay reproducible and explainable | A hotfix retrain of the old version yields a different model |
| Dependency lockfile | Transitive versions change serialization and numeric defaults | A rebuild pulls a newer library and the saved model no longer deserializes |
| Hosted LLM version, if any | The provider's model is part of the system's behavior | An undated alias moves to a new snapshot, so rollback restores your code but not the model |
The threshold row is the one most often missed, because a threshold looks like configuration. Hidden Technical Debt warns: "Thus if a model updates on new data, the old manually set threshold may be invalid." Our conclusion is that the threshold and calibration map belong to the model version, so restoring version N-1 with version N's threshold deploys a combination nobody has evaluated.
Training-serving skew is why the feature and preprocessing rows exist. Google's Rules of Machine Learning define it: "Training-serving skew is a difference between performance during training and performance during serving." The guide adds that "The best way to make sure that you train like you serve is to save the set of features used at serving time" and log them. A rollback that restores the model but keeps the new feature code reintroduces skew on purpose.
Some state has no immutable previous version to pin. A model that keeps training online, an embedding index rebuilt in place and a feature store backfilled under new definitions all overwrite their own history, so the manifest has nothing to point back to. Snapshot that state at each promotion and record the snapshot ID in the manifest, or accept that a fallback is the only safe path when such a release goes wrong.
The AWS MLOps checklist frames versioning as recovery: "Model versioning helps to track and control all changes applied to a model so that you can recover a previous version when needed." In Pharos Production's MLOps practice, every deployed version stays callable for at least 30 days after retirement, and Pharos targets reversion to a previous registry version within 5 minutes with no redeploy. The callable window keeps the endpoint alive, while the manifest keeps its features, threshold and runtime compatible. Without the manifest, the 5-minute path restores the wrong system quickly.
Rollback, Roll Through, Retrain or Fallback?
The AWS MLOps checklist on continuous deployment names the responses precisely: "In a rollback, the model reverts to a previous deployment version." Then "In a fallback, the model is replaced with a strong heuristic." And "Roll through will promote the next model to production, rolling through the previous model." Its minimum bar: "Confirm that in the ML system, you have at least one way to roll back models."
| Option | Use it when | What it does not fix |
|---|---|---|
| Rollback to the previous manifest | The new release caused the problem and the previous one still fits current inputs | A change in data the previous model also sees |
| Roll through to a corrected version | The fix is small and known, or the previous release depends on something gone | Anything quickly, since the fix still needs validation and a canary |
| Retrain | Inputs or the input-to-outcome relationship changed and no version fits | The live incident, because a pipeline run and promotion take time |
| Fallback to a heuristic or stale prediction | No model version is trustworthy and decisions must continue | Quality, which drops to the heuristic's level |
In practice the options combine: a rollback or fallback stops the damage, then a retrain or roll through becomes the durable fix. For a batch-scored model the rollback changes only the next scoring run, so predictions the bad version already persisted need a re-score with the previous version or an invalidation flag their consumers respect. NIST's AI Risk Management Framework asks for mechanisms "to supersede, disengage, or deactivate AI systems that demonstrate performance or outcomes inconsistent with intended use" and a rehearsed rollback is one concrete form of that mechanism.
Did the release cause it?
Severity sets the first move, with no engineer gate in front of it. In Pharos's drift flow, a drift score from 0.10 to 0.25 queues a retrain automatically and rolls nothing back, while a score above 0.25 returns the service to the previous version automatically and pages on-call, because stopping the damage costs less than diagnosing it on live traffic. The automatic rollback is a damage stop, not a diagnosis. It repairs only drift the release introduced, since a previous model fed the same shifted inputs cannot fix a change in the world. The rule should therefore fire once per release: a score that stays above 0.25 after the rollback is itself evidence that the data moved, so it pages without rolling back further, and version N is not promoted again until the cause is resolved. A version retrained on the shifted data is promoted with a reference window drawn from that data, or its first batches would trip the same rule. Three checks then run before version N-1 is kept long-term or any retrained version is promoted.
- Onset. A shift that starts at the traffic switch points at the release, while one that started earlier or ramps across both versions points at the data.
- Inputs versus predictions. Input feature distributions are computed before the model runs. If scores or post-threshold decisions shifted while inputs stayed stable, the model or its threshold changed. If inputs shifted and the release did not touch the feature pipeline, the world moved, unless the new model's own decisions shape the traffic it sees, in which case the onset check decides.
- The previous version on today's traffic. Score a recent sample with the previous version, which stays callable. If its predictions show the same shift, rolling back buys nothing.
If the shift started at the switch, disappears when N-1 scores today's traffic and either sits in predictions alone or traces to feature or preprocessing code the release changed, the release caused it: N-1 stays live and a corrected version later rolls through the usual canary. When the checks point at the data, the rollback has only bought time. The queued retrain becomes the durable fix, and a fallback serves if N-1 also mis-scores current traffic.
When Should a Model Roll Back? Fast and Slow Triggers
The Google Cloud MLOps architecture guide says "you need to track summary statistics of your data and monitor the online performance of your model to send notifications or roll back when values deviate from your expectations" in production. Those signals arrive on two clocks, so triggers work in two tiers.
A fast tier watches input feature distributions and the prediction distribution, compared with the reference window logged for the previous version, both available within minutes and needing no labels. Outside the automatic rollback band, Pharos pages on-call only when at least two of three signals cross their thresholds: feature distribution, prediction distribution or a business KPI. Service-level signals belong in the fast tier as well, since per-version latency and error rate catch a broken container or runtime before any distribution moves.
PSI and KS on their own scales
The Kolmogorov-Smirnov (KS) statistic is the largest gap between two cumulative distributions and lies between 0 and 1. Per its standard definition, "it is sensitive to differences in both location and shape of the empirical cumulative distribution functions of the two samples", which suits continuous features. The population stability index (PSI) sums weighted log-ratios across bins, has no upper bound and suits binned or categorical features. Pharos starts from house defaults applied to each statistic separately, 0.10 to queue a retrain and 0.25 to roll back, then tunes each cutoff per model class. A shared starting number does not make the two comparable.
Batch size drives that tuning. Yurdakul's dissertation on the statistical properties of PSI points out that the conventional 0.10 and 0.25 benchmarks carry no stated Type I or Type II error rates. It shows that, with no real shift, PSI behaves like a chi-square variable with B-1 degrees of freedom scaled by 1/N + 1/M, where B is the bin count and N and M are the sample sizes. In our own worked illustration, which simulates the exact PSI with 10 equal-probability bins and a reference window of several thousand rows or more, unshifted 100-row batches average a PSI of about 0.095 to 0.10 and roughly 40 percent of them exceed 0.10 on noise alone. At 1,000 rows the 95th percentile of noise is about 0.017 only against an effectively unlimited reference, about 0.02 against a reference of several thousand rows and about 0.034 when the reference window is also 1,000 rows. The 0.25 rollback gate is not immune either: about 1 percent of unshifted 100-row batches cross it per PSI-scored feature, which adds up across many features and batches and is one reason the false-alarm budget is set per model.
KS has a second trap on top of the same one. Its statistic is noisy on small batches too, crossing 0.10 on roughly a quarter to a third of unshifted 100-row batches, and on a large batch a KS p-value flags differences too small to matter, so the gate belongs on the statistic, calibrated per batch size. Detection power falls as batches shrink, and detectors differ in how fast: in Failing Loudly, Rabanser, Günnemann and Lipton report that the domain classifier "performs badly in the low-sample regime (≤ 100 samples), but catches up as more samples are obtained." Replaying stable history is the practical method. Run each detector over several weeks of known-good traffic at the production batch size, count how often each candidate cutoff fires and keep the one inside the false-alarm budget, which Pharos sets below 2 alerts per model per month. Repeat it when batch size or binning changes.
Slow tier: matured labels
Ground truth arrives late. AWS documentation on merging ground truth with predictions states the precondition: "To match Ground Truth labels with captured prediction data, there must be a unique identifier for each record in the dataset." Store a prediction ID and the manifest ID with every prediction, or the slow tier cannot tie an accuracy drop to a release.
Each use case also needs a declared label-maturity window. As illustrative examples only, a click label can mature within hours, a chargeback in weeks and a loan default in months. A slow-tier comparison uses matured cohorts served by each version, never recent ones whose negative outcomes are still pending.
Once labels mature, Hidden Technical Debt's prediction-bias check applies: "In a system that is working as intended, it should usually be the case that the distribution of predicted labels is equal to the distribution of observed labels." The authors note that a null model predicting average label rates also passes it, so it complements accuracy rather than replacing it.
Canary and Shadow Deployments as Rollback Enablers
Canary and shadow releases limit what a bad release touches. The SRE Workbook explains: "The canary process risks only a small fragment of our error budget, which is limited by time and the size of the canary population." Its worked example shows that a canary's errors barely move the fleet-wide rate, so the comparison has to be per version. Serving platforms such as KServe, in its serverless deployment mode, keep the last revision that served full traffic as the target to return to.
Shadow mode sends the candidate a copy of live traffic and logs its predictions without serving them, so the whole manifest, container included, meets real load with no user impact. Pharos runs shadow evaluation for at least three days before promotion.
Automatic rollback is only as good as its alarm. The AWS deployment guardrails documentation says "If any of your alarms trip during the specified monitoring period, SageMaker AI initiates a complete rollback to the old endpoint to protect your application." With no alarms configured, it notes, auto-rollback does not work, and a canary without a per-version alarm is a slow full rollout.
The right path back depends on when the problem shows up and what the release changed. During a rollout, aborting the canary is cheapest, because the previous version still serves most traffic. After full promotion, switching the registry alias back is enough when the release changed only what the model version carries, its artifact and threshold. When it also changed the container, runtime or feature code, a blue-green swap to the previous environment, kept warm with its whole manifest, restores the container, runtime and feature code in one step; shared state outside the environment, such as a backfilled feature store, still needs its recorded snapshot.
How Do You Rehearse an ML Model Rollback?

An untested rollback path is an assumption. Our checklist, run in staging on every promotion and as a periodic game day:
- Resolve the previous version's manifest and confirm that every identifier in it still resolves, from the artifact hash and image digest to the feature pipeline version and lockfile.
- Start that version in staging from the manifest alone.
- Replay a recorded traffic slice and compare outputs, including post-threshold decisions, with its stored baseline.
- Check that served feature values match those logged when it was live.
- Time the switch against the 5-minute target.
- Re-score or invalidate what the newer version wrote downstream, from caches to batch predictions already persisted, finding the rows by the manifest ID stored with each prediction.
- Exercise the fallback path too, so the team knows how decisions continue if no model version is usable at all, including who approves serving the heuristic and how long it may run.
- Record the approved rollback target and the on-call owner in the release handover.
What Changes When the Model Is a Hosted LLM?
Providers retire dated model snapshots on published schedules, and an undated alias can move to a newer snapshot with no release on your side. Pin the dated version where one is offered. Treat the prompt template, tool definitions and retrieval index as pinned state too, since an index rebuild changes answers without any model change, as our RAG data pipeline guide explains.
Retirement turns rollback into a planned roll through. Keep a frozen evaluation set with expected outcomes, run it against the replacement well before the retirement date and promote it through the same shadow and canary stages. Our state of production AI engineering report covers snapshot retirement among LLM drift sources, and the LLM observability cost guide covers the tracing that version comparison relies on.
Does the EU AI Act Require a Rollback Plan?
Not in those words. The Act mentions neither rollback nor drift thresholds, but three provisions bear on providers of high-risk systems. Article 12(1) of the consolidated AI Act reads: "High-risk AI systems shall technically allow for the automatic recording of events (logs) over the lifetime of the system." Under Article 20(1), a provider with reason to consider such a system non-conforming "shall immediately take the necessary corrective actions to bring that system into conformity, to withdraw it, to disable it, or to recall it, as appropriate." Article 72(1) adds: "Providers shall establish and document a post-market monitoring system in a manner that is proportionate to the nature of the AI technologies and the risks of the high-risk AI system."
A rehearsed rollback is one engineering means of executing a corrective action quickly, and manifest IDs logged with each prediction tie an incident to a release. What the monitoring plan must contain is covered in our AI Act post-market monitoring plan guide.
How Pharos Production Helps
Our MLOps services team builds release paths with a manifest-pinning registry, shadow and canary stages, calibrated drift triggers and a rehearsed path back. When the fix is a better model rather than an older one, our machine learning development team builds the next version.
Sources: Google SRE Workbook, Canarying Releases; Breck et al., The ML Test Score; Sculley et al., Hidden Technical Debt in Machine Learning Systems; Google, Rules of Machine Learning; AWS Prescriptive Guidance, MLOps checklist; NIST, AI RMF 1.0; Yurdakul, Statistical Properties of Population Stability Index; Rabanser et al., Failing Loudly; Regulation (EU) 2024/1689, consolidated text. Engineering guidance, not legal advice.
FAQ
Quick answers to common questions about custom software development, pricing, process and technology.
Type to filter questions and answers. Use Topic to narrow the list.
Showing all 6
No matches
Try a different keyword, change the topic or clear filters
-
After rolling back a batch-scored model, which predictions still need fixing?
Every prediction the bad version persisted between promotion and rollback. Rolling back the scorer changes only future runs, so rows already written to tables and caches keep the bad scores.
Select them by the manifest ID stored with each row, re-score them with the restored version or mark them invalid so consumers fall back. Then tell downstream owners which tables and time range were affected.
-
How do you roll back a model that keeps learning online?
Stop the updates first, or the model keeps learning from the traffic that caused the incident. Then restore a parameter checkpoint taken before the onset, together with the input offset it was trained up to, and decide whether the data since then is replayed after filtering or discarded.
Without such checkpoints there is no earlier version to return to, and a fallback is the only safe path until a clean model is rebuilt.
-
Does a blue-green setup make the registry alias unnecessary?
No. A blue-green swap moves traffic between two running environments, while the alias records which model version is approved for serving. Keep the alias as the source of truth and have the swap follow it, or the two disagree after a few incidents.
Blue-green also costs a second warm environment for as long as the previous one is kept ready, so it pays off mainly for releases that change the container or feature code.
-
Who approves an automatic rollback?
For severe cases the approval is given in advance, not during the incident. In the Pharos drift flow, a score above 0.25 rolls back to the previous version, with reversion targeted within 5 minutes, and pages on-call, while a per-version canary alarm aborts a rollout, because the next action is set by severity rather than by engineer attention.
Between 0.10 and 0.25 a retrain is queued automatically and nothing is rolled back. Engineers come in afterwards, when the causal checks decide whether the previous version is kept long-term or a new one is promoted. Steps that reach beyond the serving pointer, such as serving a heuristic fallback or restoring a feature store snapshot, need a named approver recorded with the rollback target in the release handover.
-
What if the hosted model snapshot you would roll back to has been retired?
Then the previous manifest cannot be restored as written, and the rehearsal step that resolves every identifier should flag it before an incident does. What remains is a fallback or a roll through to the replacement snapshot, which is safe only if it has already passed the frozen evaluation set and the shadow stage.
Mark the retired manifest as unrestorable so on-call does not reach for it under pressure.
-
Where should the pin manifest be stored?
Next to the model version in the registry, as a write-once record keyed by its own ID and referenced from the release handover. Keep it outside the serving cluster so it survives the incident it exists to fix, and never edit it after promotion: a manifest changed later no longer describes the combination that was evaluated.
I work with startup founders who need a dedicated software development team but don’t want to gamble on hiring, random outsourcing, or opaque delivery.
Most founders face the same problem sooner or later.
Early technical and team decisions lock the product into tech debt, slow delivery, missed milestones and constant re-hiring. By the time this becomes visible, fixing it is already expensive.As a CTO and software architect, I help founders design, build and run dedicated development teams that work as a true extension of the startup. Not as a black-box vendor.
My focus is on complex products where mistakes are costly:
- Web3 and blockchain platforms
- FinTech and regulated products
- High-load startup systems
- MVP → scale transitions
We don’t do body-shopping.
We don’t sell generic outsourcing.Instead, we help founders:
- build the right team structure from day one
- keep technical ownership and transparency
- scale delivery without losing control
- avoid vendor lock-in and hidden risks
Teams are aligned with the product roadmap, business goals and long-term architecture. Not just short-term velocity.