📋 KEY INSIGHTS
- A model that is not in production has zero business value. Deployment is not a post-modelling afterthought β it is the goal, and every modelling decision should be made with deployment constraints (latency, throughput, memory, interpretability) in mind from the start.
- REST APIs wrapped in Docker containers have become the standard unit of ML model deployment β they are cloud-agnostic, testable, scalable, and separable from the training infrastructure, making them the most portable and maintainable deployment pattern.
- Model serving latency has two components: the model inference time (which is a function of model complexity and hardware) and the API overhead (serialisation, network, authentication). For latency-sensitive applications, both must be profiled and optimised separately.
- Blue-green and canary deployment strategies allow new model versions to be rolled out gradually with automatic rollback β essential for production ML where a degraded model silently produces wrong predictions rather than throwing an error.
- Model monitoring is the most neglected component of the ML lifecycle: most teams monitor infrastructure (server uptime, API latency) but not the model itself (prediction distribution shift, feature drift, output calibration). A model can be “serving” while silently failing to deliver value.
- ONNX (Open Neural Network Exchange) is the standard format for exporting trained models from any framework (PyTorch, TensorFlow, scikit-learn) and serving them with optimised runtimes (ONNX Runtime, TensorRT) that can deliver 2β10Γ lower latency than serving from the original framework.
The gap between a trained model in a Jupyter notebook and a model delivering value in production is where most ML projects stall. The modelling work β data exploration, feature engineering, training, evaluation β feels like the substance of the project. But a model that lives in a notebook is an interesting experiment, not a product. Getting from experiment to production requires understanding REST API design, containerisation with Docker, orchestration with Kubernetes, deployment strategies, monitoring, and rollback. These are engineering skills that data scientists increasingly need to own, not delegate. This guide covers the full deployment stack for ML models β from wrapping a model in an API to monitoring it in production β with architecture diagrams in table form and practical decision guidance at each stage.
The ML Serving Stack β Components and Responsibilities
A production ML serving system has five layers, each with distinct responsibilities. Understanding what each layer does and where failures can occur is the foundation of reliable deployment.
Layer 1 β The Model Artefact: The serialised, versioned representation of the trained model. For scikit-learn and LightGBM models, this is a pickle file or joblib file. For PyTorch models, it is a state_dict checkpoint or a TorchScript-compiled module. For TensorFlow/Keras, it is a SavedModel directory. The best practice is to export to ONNX format, which decouples the model from the training framework and enables serving with optimised runtimes. Alongside the model artefact, the serving system needs the preprocessing pipeline β the exact transformations (scaling, encoding, imputation) applied to training data β serialised with the same library version used during training.
Layer 2 β The Prediction Service: A web service that receives prediction requests, applies the preprocessing pipeline, runs model inference, post-processes outputs, and returns predictions. FastAPI (Python) is the most popular framework for building prediction services β it provides automatic request validation via Pydantic schemas, async support for concurrent requests, and auto-generated OpenAPI documentation. The prediction service should validate inputs (reject requests with missing required features or out-of-range values), log every prediction with its inputs and outputs (essential for monitoring and debugging), and return structured responses with prediction confidence alongside the primary prediction.
Layer 3 β The Container: Docker packages the prediction service and all its dependencies (Python version, library versions, system libraries) into a portable image that runs identically across development, staging, and production environments. The key practices are: use a minimal base image (python:3.11-slim rather than a full OS image) to reduce image size and attack surface; pin all dependency versions in requirements.txt to ensure reproducibility; and implement a multi-stage build that separates the build environment from the runtime image, further reducing the final image size.
Layer 4 β The Orchestrator: Kubernetes manages the lifecycle of containerised services at scale β scheduling containers onto cluster nodes, restarting failed pods, scaling the number of replicas in response to traffic, performing rolling updates, and routing traffic. For most ML serving use cases, a Deployment resource (managing a set of identical prediction service pods) behind a Service resource (providing a stable internal IP) behind an Ingress resource (routing external traffic) is the standard pattern. Horizontal Pod Autoscaler (HPA) automatically adjusts replica count based on CPU utilisation or custom metrics (requests per second, queue depth), enabling the serving infrastructure to absorb traffic spikes without manual intervention.
Layer 5 β The Monitoring System: Prometheus (metrics collection) + Grafana (visualisation) is the standard open-source monitoring stack for containerised services. Metrics to track divide into two categories: infrastructure metrics (CPU utilisation, memory, request latency, error rate, pod restarts) and model metrics (prediction distribution, feature distribution, output confidence distribution, and business-level KPIs). Model metric monitoring is the more critical and more neglected layer β a model can be serving requests with low latency while its prediction distribution has drifted significantly from the distribution seen during validation, silently delivering degraded business value.
| Layer | Technology | Key Decision | Common Failure |
|---|---|---|---|
| Model Artefact | ONNX / TorchScript / SavedModel | Framework-native vs ONNX | Preprocessing pipeline not versioned alongside model |
| Prediction Service | FastAPI, Flask, BentoML, Triton | Sync vs async; batch vs single | Missing input validation; no prediction logging |
| Container | Docker | Base image; dependency pinning | Unpinned dependencies break on rebuild |
| Orchestrator | Kubernetes (EKS, GKE, AKS) | HPA thresholds; resource limits | No resource limits β one pod crashes node |
| Load Balancer | Nginx Ingress, AWS ALB, Istio | Path-based vs header-based routing | No rate limiting β spike takes down service |
| Monitoring | Prometheus + Grafana, Datadog, Evidently | Which model metrics to track | Only infra metrics; model drift invisible |
Deployment Strategies β Reducing Risk for New Model Versions
Releasing a new model version to 100% of production traffic immediately is the highest-risk deployment strategy. A degraded model does not throw an exception β it continues serving predictions silently, making it possible for a bad model to run in production for hours or days before the business impact becomes visible. Safer strategies manage the rollout progressively.
Blue-Green Deployment: Two identical production environments are maintained β blue (the current live version) and green (the new version). The new model is deployed and validated on the green environment with synthetic or replayed traffic before any live traffic is switched. The load balancer then switches all traffic from blue to green in a single step. Rollback is instantaneous β switch back to blue. The cost is maintaining two full production environments simultaneously, which doubles infrastructure cost during the transition period.
Canary Deployment: The new model version receives a small percentage of live traffic (typically 1β5%) while the current version continues serving the remainder. The canary is monitored against the same business metrics as the current version. If the canary performs equivalently or better over a predefined evaluation window, the traffic percentage is gradually increased (5% β 20% β 50% β 100%). If the canary underperforms, traffic is rolled back to zero. Canary deployment requires a traffic splitting mechanism (feature flags, A/B routing in the load balancer, or a service mesh like Istio) and automated monitoring with alerting rules that trigger rollback when metrics fall below thresholds.
Shadow Mode: The new model receives a copy of all live traffic and generates predictions, but those predictions are not served to users β they are logged and compared to the current model’s predictions. Shadow mode is invaluable for validating that a new model’s prediction distribution matches expectations on live data before exposing it to users. It is also the safest way to validate a model that produces consequential outputs (medical diagnoses, financial decisions) where even a brief period of degraded performance would be unacceptable.
| Strategy | Traffic Split | Rollback Speed | Infrastructure Cost | Best For |
|---|---|---|---|---|
| Full cutover | 0% β 100% instantly | Minutes (re-deploy old) | 1Γ | Low-risk updates only |
| Blue-Green | 0% β 100% at switch | Seconds (flip load balancer) | 2Γ during transition | Zero-downtime releases, fast rollback needed |
| Canary | 1% β 5% β … β 100% | Seconds (reduce to 0%) | ~1Γ (small canary) | Gradual validation on live traffic |
| Shadow Mode | 100% copy (not served) | N/A (not in production) | 2Γ (full shadow infra) | High-stakes validation before any live traffic |
| A/B Test | 50%/50% (or split by segment) | Seconds (remove split) | ~1Γ | Measuring business impact of model change |
✦ SUMMARIZE THIS ARTICLE WITH AI
The MLOps practices that surround deployment β CI/CD pipelines for models, feature stores, and experiment tracking β are covered in our MLOps Interview Q&A. ML system design patterns at the architecture level β online vs batch serving, feature pipelines, model registries β are in our ML System Design guide. Cloud infrastructure for running Kubernetes clusters (EKS, GKE, AKS) is covered in our Cloud for Data Scientists guide. Building lightweight prediction UIs on top of deployed model APIs is covered in our Streamlit Deployment guide.



