MLOps: From Notebook to Production
Most ML projects die between the notebook and production. Feature stores, model registries, and serving infrastructure are what actually close that gap.
Most machine learning projects do not fail because the model is bad. They fail in the gap between a notebook that produces a promising result and a service that keeps producing correct results months later, on live data, without anyone watching it constantly. At Scutiger, our platform engineering practice exists specifically to close that gap: feature stores, model registries, and serving infrastructure that treat a model the way we treat any other production dependency — versioned, monitored, and built to fail safely.
Key Takeaways
- A notebook result is a hypothesis, not a deliverable. The engineering work of making a model reliable in production is usually larger than the data science work of training it.
- Training-serving skew is the most common silent failure. Feature logic implemented twice — once for training, once for serving — drifts apart in ways that degrade accuracy without throwing an error.
- Models need version control like code does. Without a registry, “which model is actually running in production” becomes a question nobody can answer confidently.
- Drift monitoring is not optional. A model that stops matching reality does not crash — it just gets quietly, increasingly wrong.
The Notebook-to-Production Gap
A data scientist trains a model in a notebook against a CSV export or a database snapshot. The features are computed with whatever pandas code was convenient at the time. The result looks good — accuracy, precision, recall all clear the bar.
Then someone asks: how does this run against live data, every day, without a human re-running the notebook? That question exposes everything the notebook did not have to solve: where do features come from in production, who retrains the model when performance degrades, what happens if the input schema changes upstream, and how does anyone know the model degraded before a business metric visibly suffers.
This mirrors a pattern we see across industries we serve for data science: the model is rarely the hard part. This is not a data science problem. It is a platform engineering problem, and it is the reason most ML initiatives stall between “we built a promising model” and “the model is running reliably in production.”
The Platform Layer We Build
Feature stores
A feature store computes features once and serves them consistently to both training pipelines and live inference — eliminating the two-implementations problem. When a data scientist trains against the feature store’s historical view and the production service queries the same store’s online view, the numbers a model sees in production match what it saw in training, by construction rather than by careful manual discipline.
Model registries
Every trained model gets a version, a record of what data and code produced it, and its evaluation metrics — before it is eligible to serve traffic. When a model in production starts underperforming, the registry answers “what changed” in minutes instead of a multi-day archaeology project through Slack history and old notebooks.
Serving infrastructure
Low-latency inference endpoints, typically behind the same API gateway and observability stack as the rest of a client’s services — a model is a production dependency, not a special case that lives outside normal engineering practice. Canary deployment for new model versions lets a new model take a small percentage of traffic before it takes all of it.
Monitoring for drift and staleness
Two kinds of monitoring matter, and they catch different failures. Input drift monitoring watches whether the live data distribution is diverging from the training distribution — a warning that the world has changed even before ground truth is available. Outcome monitoring compares predictions against ground truth once it arrives, to catch quality degradation directly. A model with neither is running blind.
A Worked Example: Retraining Without Drama
Consider a model that scores lead quality for a sales team. Set up correctly, the pipeline looks like this: the feature store computes lead features consistently for training and serving; a scheduled job retrains the model weekly against the latest labeled outcomes; the new model version is registered with its evaluation metrics attached; a canary deployment routes 5% of traffic to the new version; automated monitoring compares the canary’s outcome distribution against the previous version’s; if metrics hold, the rollout proceeds automatically — if they regress, it rolls back automatically, and an engineer is alerted rather than a customer being affected.
Without this pipeline, the same retraining event typically means: someone remembers it is time to retrain, manually re-runs a notebook, manually swaps a model file in production, and finds out something went wrong only when a downstream metric looks off — days later, after the bad model has already been making decisions.
Comparing the Two Approaches
| Ad hoc notebook workflow | MLOps pipeline | |
|---|---|---|
| Feature computation | Reimplemented separately for training and serving | Computed once, served consistently by a feature store |
| Model versioning | Informal — a filename or a date in a comment | Registry entry with data lineage and metrics |
| Deployment | Manual file swap | Canary rollout with automated rollback |
| Detecting degradation | Someone notices a downstream metric looks wrong | Drift and outcome monitoring alert automatically |
| Retraining cadence | Whenever someone remembers | Scheduled, with automated evaluation gating |
Limitations
Building this platform layer is a real investment, and it is not warranted for every model. A one-off analysis that runs once and informs a single decision does not need a feature store or a registry — that is over-engineering for a model with no ongoing serving requirement. The platform layer earns its cost when a model runs continuously against live data and a wrong prediction has a real cost, which is most models that actually inform a production decision rather than a single report. We also cannot promise a model will stay accurate forever: monitoring tells you when a model has drifted, but deciding whether to retrain, retire, or redesign the model is a judgment call, not something the platform makes for you.
Bottom Line
The distance between a working notebook and a reliable production model is mostly infrastructure, not modeling skill: consistent feature computation, versioned models, safe rollout, and monitoring that catches degradation before a downstream metric does. We build that infrastructure as part of our platform engineering practice — see internal developer platforms and golden paths for how this fits into the broader paved-road approach, or our engineering platform service for what a typical MLOps engagement covers.
Frequently Asked Questions
- Why do most machine learning projects never reach production?
- A model that performs well in a notebook was trained on a static snapshot of data using ad hoc feature computation. Reaching production requires the same features to be computed consistently on live data, the model to be versioned and rollback-capable, and monitoring to catch drift — none of which a data scientist's notebook environment provides by default.
- What is a feature store and why does it matter?
- A feature store is a system that computes and serves the input features a model needs, using the same logic for both training and live inference. Without one, teams often reimplement feature logic separately for training (in a notebook or batch job) and serving (in production code), and the two implementations quietly drift apart — a problem known as training-serving skew.
- How do you know if a production model has degraded?
- You monitor for drift: statistical changes in the input data distribution, and degradation in prediction quality once ground truth becomes available. Both require dashboards and alerting wired into the serving layer from day one, not added after a model has already been silently wrong for a month.