FutureStackDev AI Agent Development

Kubernetes for AI Systems

GPUs, queues, autoscaling, and serving graphs on Kubernetes — without turning the cluster into the product.

Created by Baljeet Dogra

Course Objectives

  • Schedule a GPU job without hoarding the node.
  • Serve a model with a health check, a queue, and a scale rule you can explain.
  • Separate training and serving so a job cannot starve the API.
  • Read events and logs when the pod is not the bug.

What you will produce

Serve

Inference workload

Deployment, service, HPA or KEDA, a load test.

Train

Job with GPU

A finite job, node selectors, a cleanup you have practised.

4 weeks curriculum

Expand a part for the syllabus. Content stays searchable when closed.

01 Boring Kubernetes Week 1

Pods, deploys, services, probes, config, secrets. AI does not skip this.

  • Workloads
  • Probes
  • Config
  • RBAC
02 GPUs and jobs Week 2

Device plugins, sharing vs isolation, Jobs vs CronJobs.

  • GPUs
  • Jobs
  • Selectors
  • Quotas
03 Serve and scale Week 3

Queues, autoscaling on tokens or lag, streaming and timeouts.

  • Serving
  • HPA / KEDA
  • Queues
  • Timeouts
04 Operate Week 4

Network policies, cost, incident drill. Capstone: train job + serve path.

  • Network
  • Cost
  • Incidents
  • Capstone

Who this is for

Platform and MLOps engineers

You inherited a cluster. AI workloads are noisy neighbours.

Engineers moving off laptops

You can Docker. Production is a scheduler.

Prerequisites

  • Docker comfort
  • You have deployed a container
  • Basic YAML literacy

Not a fit if

  • CKA cram with no AI workloads
  • People who have never used a container

Related: MLOps Deep Dive · Enterprise AI Engineer

Questions

Cloud or local?

A local cluster for muscle memory, and a cloud path so GPUs are real.

Which serving project?

One mainstream option. You should be able to leave it.

Ready to start?

Four weeks to GPU jobs and a serving path you can operate. Create an account to enrol.

Enrol now