Course 2 — MLOps & Production AI Engineering (60-Class Career Track)
Goal: Take engineers from machine learning models to production scale. Master model serving APIs, experiment tracking, data versioning, feature stores, automated continuous training (CT), model monitoring, and scalable AI infrastructure on Kubernetes.
Core Engineering Competencies
Detailed Class-by-Class Syllabus
What is MLOps, ML technical debt, Python for MLOps, FastAPI model serving, Pydantic request validation, model serialization (ONNX/Joblib), and testing.
DevOps vs MLOps, why 85% of ML models fail to reach production, MLOps pillars (Code, Data, Model).
Hidden technical debt in ML systems, glue code, pipeline jungles, dead experimental code.
Architecture blueprint: Data Ingestion → Feature Store → Training → Registry → Serving → Monitoring.
Clean code principles, type hints (typing module), dataclasses, Pydantic validation schemas.
Building REST APIs with FastAPI, routing, path/query parameters, dependency injection for model loading.
Pydantic BaseModel for feature inputs, input boundary validation, HTTP error status handling.
Pickle vs Joblib vs ONNX (Open Neural Network Exchange) vs TorchScript for portable inference.
Async endpoints, batch prediction vectors, minimizing latency and CPU overhead.
Unit testing with Pytest, testing API endpoints with TestClient, data schema validation tests.
Containerizing and deploying a production-ready ML inference microservice with Docker and Swagger documentation.
Project: Production-Ready FastAPI Machine Learning Microservice
Trained ML Model → ONNX Export → FastAPI Asynchronous Endpoint → Pydantic Schema Validation → Multi-Stage Docker Container.
Reproducibility, MLflow Tracking, autologging, artifacts, metric logging, remote MLflow server on AWS, and Model Registry workflows.
Why model training must be version-controlled, tracking hyperparameters, metrics, and seeds.
MLflow Tracking, MLflow Projects, MLflow Models, MLflow Model Registry core components.
mlflow.log_param(), mlflow.log_metric(), mlflow.log_artifact(), tracking loss curves over training epochs.
Autologging with Scikit-learn, PyTorch, TensorFlow, and XGBoost with zero boilerplate code.
Setting up remote MLflow backend on AWS EC2, PostgreSQL backend store, and S3 artifact store.
Registering models, semantic versioning (v1, v2, v3), model stages: Staging, Production, Archived.
Automating model stage transitions based on validation metric thresholds (e.g. F1-score > 0.90).
Overview of W&B dashboards, hyperparameter sweeps, and collaborative model evaluation.
Packaging models with MLmodel flavor definitions for universal deployment targets.
Setting up production MLflow server on AWS, running hyperparameter experiments, and promoting best model to Production.
Project: Centralized Enterprise Experiment Tracking & Model Registry
Remote MLflow on AWS EC2 + S3 Artifacts + RDS PostgreSQL → Hyperparameter Experimentation → Automated Validation Gate → Production Promotion.
Data Version Control (DVC), S3 remote storage, reproducible pipelines, Great Expectations data validation, and Feast feature store.
Why Git cannot handle large datasets, pointer files (.dvc), content-addressable storage.
Configuring S3 remote storage in DVC, dvc add, dvc push, dvc pull across developer workstations.
Writing dvc.yaml pipelines, stage dependencies, output caching, executing end-to-end with dvc repro.
Data testing principles, detecting silent data corruption and unexpected schema shifts.
Defining Expectation Suites, automated data profiling, generating HTML Data Docs.
What is a feature store? Eliminating training-serving skew, feature reusability across models.
Feast definitions, Entities, FeatureViews, source datasets, feature repository structure.
Offline store (Parquet/Snowflake for batch training) vs Online store (Redis for low-latency inference).
Materializing features to Redis, fetching real-time feature vectors during live model inference.
Building an end-to-end data pipeline with DVC, Great Expectations validation, and Feast feature serving.
Project: Versioned Data Pipeline & Feast Feature Store
Raw Data → DVC + S3 Versioning → Great Expectations Data Validation → Feast Feature Store → Real-time Redis Serving.
MLOps maturity levels, Docker for ML, GitHub Actions CI for ML, Continuous Machine Learning (CML), automated model training, and evaluation gates.
Level 0: Manual process, Level 1: ML Pipeline automation, Level 2: CI/CD/CT automated pipeline.
Optimizing Docker images for ML, CUDA GPU support, lightweight slim base images.
Building production container images with separate build and runtime environments.
Automated linting (Flake8, Black), Pytest test suites, and data validation in GitHub Actions.
Automating pull-request model reports, markdown metrics tables, ROC curve image generation in PR comments.
Triggering automated training runs on new data ingestion or schedule via GitHub Actions.
Automated metric comparison: comparing newly trained candidate model against active production champion model.
Building Dockerized inference service upon successful model training and pushing to AWS ECR.
Automating model artifact sync to production servers without service interruption.
Code/Data Commit → Automated Training → Quality Gate Check → Docker Build → Automated Staging Deployment.
Project: Automated CI/CD/CT Pipeline for Machine Learning
Git Commit → GitHub Actions CI → DVC Data Pull → Automated Training → Evaluation Quality Gate → Container Registry → Deployment.
Containerized ML serving on Kubernetes, TorchServe/Triton, KServe serverless ML, Kubeflow Pipelines (KFP), Canary rollouts, and Autoscaling.
Deploying FastAPI ML inference Pods, Services, and Nginx Ingress on Kubernetes.
NVIDIA Triton Inference Server architecture, dynamic batching, multi-model execution.
KServe (KFServing) CRDs, scale-to-zero serverless inference, inference graphs.
Kubeflow architecture on Kubernetes, Kubeflow dashboard, Central Dashboard components.
Writing Kubeflow pipeline components in Python with @component decorator, passing data artifacts.
Compiling and submitting Kubeflow training workflows with data preprocessing, training, and evaluation.
Routing 10% traffic to candidate model and 90% to champion model on Kubernetes.
Mirroring live production inference requests to test new models without impacting users.
Horizontal Pod Autoscaler configuration based on GPU utilization and inference request metrics.
Deploying a high-throughput multi-model inference fleet on AWS EKS with Canary rollout and Autoscaling.
Project: Scalable Multi-Model Serving Fleet on AWS EKS
AWS EKS Kubernetes Cluster → KServe / FastAPI Serving → Ingress Canary Traffic Splitting (90/10) → Horizontal Pod Autoscaling.
Production ML monitoring, data drift, concept drift, statistical drift tests (KS, PSI), Evidently AI, Prometheus/Grafana, LLMOps, and Final Capstone.
Operational metrics (latency, error rate) vs Data Science metrics (drift, feature stability, accuracy).
Why production models degrade: distribution changes, seasonal trends, covariate shift.
Kolmogorov-Smirnov (KS) test, Population Stability Index (PSI), Wasserstein Distance.
Automated data drift dashboards, target drift reports, data quality test suites.
Exporting drift scores and prediction metrics to Prometheus, real-time Grafana dashboards.
Triggering automated retraining pipelines in GitHub Actions/Kubeflow upon drift threshold breaches.
Monitoring Large Language Models, prompt drift, token usage, latency, hallucination detection.
Model explainability (SHAP), fairness, audit logging, model lineage tracking.
Building the end-to-end production MLOps system: Data Ingestion → DVC → MLflow → CI/CD/CT → EKS Serving → Evidently AI Drift Monitoring.
1-on-1 viva defense, live production architecture demonstration, and Course 2 MLOps Certification.
Course 2 Capstone: Production End-to-End MLOps & Continuous Training Platform on AWS & Kubernetes
Data Versioning (DVC + S3) → Experiment Tracking (MLflow) → Automated CI/CD/CT (GitHub Actions) → Kubernetes Deployment (AWS EKS) → Real-Time Monitoring & Drift Detection (Evidently AI + Prometheus).