FutureStackDev AI Agent Development

MLOps Deep Dive

Training is not the job. Pipelines, registries, promotion, monitoring, and a rollback when the new model is worse.

Created by Baljeet Dogra

Course Objectives

  • Promote a model because of evals, not because the loss looked nice.
  • Reproduce a run from a commit and a data pin.
  • Watch live quality, not only latency.
  • Roll back in a drill, not in a post-mortem.

What you will produce

Pipeline

Train-to-serve

Pinned data, registry, a gated deploy, a rollback switch.

Ops

Quality dashboard

Online proxy metrics and an alert you have actually fired.

6 weeks curriculum

Expand a part for the syllabus. Content stays searchable when closed.

01 Reproducibility Week 1

Seeds, data pins, environments, experiment tracking without theatre.

  • Pins
  • Tracking
  • Seeds
  • Artefacts
02 Registry and CI Weeks 2–3

Model cards, promotion rules, evals that can fail the build.

  • Registry
  • Cards
  • Gates
  • CI
03 Deploy patterns Week 4

Batch, online, shadow, canary. Feature freshness.

  • Batch
  • Online
  • Shadow
  • Canary
04 Monitor and roll back Weeks 5–6

Drift, data quality, quality proxies. Capstone drill.

  • Drift
  • Alerts
  • Rollback
  • Capstone

Who this is for

ML engineers in production

You train. Ops still happens in a notebook and a hope.

Platform engineers

You need ML-shaped CI, not a copy of the app pipeline.

Prerequisites

  • You have trained a model
  • Git and CI at a basic level
  • Containers at a basic level

Not a fit if

  • Kubeflow tourism
  • People who have never trained anything

Related: How to Become an AI Engineer · Kubernetes for AI Systems

Questions

Which stack?

One tracking tool, one registry, one deployer. Vendor-plural on purpose.

Is this Kubernetes?

Optional. Kubernetes for AI Systems is the dedicated course.

Ready to start?

Six weeks to promote, watch and roll back a model. Create an account to enrol.

Enrol now