What is MLOps?
Google Cloud's MLOps architecture guidance defines MLOps as an ML engineering culture and practice that unifies ML system development (Dev) with ML system operation (Ops). It advocates automation and monitoring at every step, from integration and testing to release, deployment and infrastructure management.
Microsoft describes MLOps as DevOps principles, such as continuous integration and continuous delivery, applied to the machine learning lifecycle. The difference is that an ML system's behaviour depends on data as well as code, so datasets, schemas and trained models must be versioned, tested and released too.
Why do machine learning models fail in production without MLOps?
Google's MLOps guidance notes that only a small fraction of a real-world ML system is the ML code. Many production problems therefore come from everything around the model, not the algorithm.
- Drift and staleness: Azure Machine Learning documentation lists data distribution changes, training-serving skew, data quality problems, environment shifts and changes in consumer behaviour as causes of a model becoming stale.
- Poor reproducibility: if the data snapshot, code version, dependencies and hyperparameters were not recorded, nobody can rebuild or audit the model that is serving predictions.
- Manual handoffs: in Google's level 0 description, data scientists hand a trained model to engineers, deployment covers only the prediction service and releases are infrequent.
- Silent degradation: a model keeps returning answers even when they are wrong, so without active monitoring the first warning can be a business complaint.
What are the stages of the MLOps lifecycle?
A mature lifecycle connects eight stages, each feeding the next automatically.
- 1Data versioning: snapshot training data, for example with DVC metafiles in Git or versioned data assets in Azure Machine Learning, so every model points to the exact data it learned from.
- 2Experiment tracking: log parameters, code versions, metrics and artefacts for every run, as MLflow Tracking does.
- 3Automated training pipelines: replace notebook cells with containerised, parameterised steps for data preparation, validation, training and evaluation.
- 4Model registry: store each candidate as a named, versioned entry with lineage back to the run that produced it.
- 5CI/CD for models: test pipeline code, data schemas and model quality, then promote approved models between environments automatically.
- 6Deployment and serving: expose models through real-time endpoints or batch scoring, using canary or traffic-split rollouts.
- 7Monitoring: compare production inputs and predictions with a baseline to detect drift, data quality issues and falling performance.
- 8Retraining: trigger the pipeline on a schedule, on new data or on detected degradation, and send the new model back through validation.
What are Google Cloud's MLOps maturity levels 0, 1 and 2?
Level 0 is a manual process. Data preparation, training and validation are run by hand in scripts or notebooks, models are released only a few times a year, there is no CI or CD, and prediction quality in production is not actively monitored.
Level 1 automates the ML pipeline to achieve continuous training. The pipeline retrains on fresh data when triggered on demand, on a schedule, when new training data arrives, when performance degrades or when data distributions change significantly. The same pipeline runs in development and production, which Google calls experimental-operational symmetry, and the level adds data and model validation and metadata management, with a feature store as an optional addition.
Level 2 adds automated CI/CD for the pipeline itself. Pipeline code is built and tested automatically, with unit tests for feature logic, checks that training converges and integration tests between components, then deployed to the target environment.
Which open-source tools are used for MLOps?
MLflow covers tracking and model management. MLflow Tracking is an API and UI for logging parameters, code versions, metrics and output files, and the MLflow Model Registry is a centralised model store with versioning, aliases such as a champion model, tags and lineage to the experiment run that produced each model.
Kubeflow Pipelines is a platform for building and deploying portable, scalable ML workflows using containers on Kubernetes, with pipelines authored in Python and caching that avoids rerunning unchanged steps. DVC complements both by versioning data and models through small Git metafiles while the files themselves stay in cloud or on-premises storage.
Open source keeps workflows portable across clouds, but your team then runs and secures the tracking server and Kubernetes cluster.
How do Azure, AWS and Google Cloud support MLOps?
Each provider offers a managed platform covering most of the lifecycle: Azure Machine Learning, Amazon SageMaker AI and Gemini Enterprise Agent Platform on Google Cloud, formerly Vertex AI. The table uses current Agent Platform names.
Azure Machine Learning workspaces are MLflow-compatible and SageMaker AI offers fully managed MLflow, so MLflow tracking code moves between them with little change. Amazon SageMaker Model Monitor is no longer open to new customers; AWS documents a replacement that combines open-source SageMaker AI monitoring solutions, built on SageMaker AI MLflow Apps and Evidently AI, with Amazon QuickSight dashboards and Amazon CloudWatch.
| Capability | Azure | AWS | Google Cloud |
|---|---|---|---|
| Experiment tracking | MLflow tracking in an Azure Machine Learning workspace | Managed MLflow on Amazon SageMaker AI | Agent Platform Experiments (formerly Vertex AI Experiments) |
| Training pipelines | Azure Machine Learning pipelines | Amazon SageMaker Pipelines | Agent Platform Pipelines (formerly Vertex AI Pipelines) |
| Model registry | Azure Machine Learning model registry | SageMaker Model Registry with Model Groups and approval status | Agent Platform Model Registry (formerly Vertex AI Model Registry) |
| Lineage and metadata | Job history, versioned data assets and model metadata | Amazon SageMaker ML Lineage Tracking | Vertex ML Metadata |
| Serving | Online endpoints and batch endpoints | Real-time endpoints and batch transform | Online inference on endpoints and batch inference |
| Production monitoring | Azure Machine Learning model monitoring | Model Monitor for existing customers; open-source solutions with QuickSight and CloudWatch as its replacement | Model Monitoring on Agent Platform (formerly Vertex AI Model Monitoring) |
| CI/CD automation | Azure Pipelines, with Azure Event Grid for lifecycle events | SageMaker Projects templates with CodePipeline or Jenkins | Cloud Build with Agent Platform Pipelines |
What belongs on an MLOps production-readiness checklist?
Work through these in order before a model takes real traffic.
- 1Version code, data and the software environment together so any model can be rebuilt exactly.
- 2Log every training run's parameters, metrics and artefacts to a shared tracking server.
- 3Package training as a pipeline of containerised steps rather than a notebook.
- 4Validate incoming data against an expected schema and statistical profile before training.
- 5Evaluate each candidate against the current production model, overall and on important data segments.
- 6Register the approved model with its lineage, metrics and an explicit approval status.
- 7Deploy through CI/CD with a canary or traffic-split rollout and a tested rollback path.
- 8Capture production inputs and predictions so they can be compared with the training baseline.
- 9Alert on data drift, prediction drift, data quality and, once labels arrive, model performance.
- 10Define retraining triggers and who approves promotion of a retrained model.
How can WIEWAVE help?
WIEWAVE is a cloud, DevOps and AI company that builds MLOps foundations for clients worldwide on Azure, AWS and Google Cloud, from experiment tracking and training pipelines to model registries, deployment and drift monitoring, defined as infrastructure-as-code in your own cloud accounts. If your models still live in notebooks, placing yourself against the maturity levels above is a practical first step.
Sources
- Google Cloud: MLOps, continuous delivery and automation pipelines in machine learning
- Microsoft Learn: MLOps model management with Azure Machine Learning
- Microsoft Learn: Model monitoring in production with Azure Machine Learning
- AWS documentation: Amazon SageMaker Model Monitor availability change
- AWS documentation: Managed MLflow on Amazon SageMaker AI
- Google Cloud documentation: MLOps on Gemini Enterprise Agent Platform
