Ship #1
Instrumented LLM API
Schema-validated output, retries, streaming, live cost and latency dashboard.
From API calls to production systems. Organised by shipped increments, not topics. You deploy in week 4, and every four weeks after that — seven sprints, seven ships.
Created by Baljeet Dogra
The destination is the same as a topic-ordered GenAI course: you start calling LLM APIs and you leave able to build and deploy a RAG or agent system. The route is different. Engineers who deploy seven times learn things that engineers who deploy once at the end do not.
Seven deployed artefacts, not seven notebooks. Each sprint: week 1 concepts and a guided build, week 2 a harder requirement, week 3 you break it and fix it, week 4 you deploy, instrument, and defend in peer review. Nothing is complete until it runs off your laptop.
Ship #1
Schema-validated output, retries, streaming, live cost and latency dashboard.
Ship #2
Golden set on real data, gating deploys of Ship #1, with a baseline report.
Ship #3
Citations, abstention, retrieval metrics through the same harness.
Ship #4
Real multi-step task, guardrails, budgets, full run tracing.
Ship #5
Adapter behind an API, plus an honest report on whether it was worth it.
Ship #6
Latency and cost SLOs, tracing, alerting, runbook. Cut Ship #3 cost ~50%.
Capstone
Problem, ADR, evals, cost per task, runbook, known limits. Assessment: capstone 35%, six ships 30%, eval rigour 15%, peer review 10%, debugging 10%. No exams.
Seven sprints. About 3 hours live each week plus about 5 hours of lab. Each classroom lesson has a talk, a diagram, a ship, and a scenario quiz. Expand a sprint for the syllabus. Content stays searchable when closed.
Most first LLM code fails in production for reasons that have nothing to do with AI.
Ship #1: deployed LLM-backed API with schema-validated output, retries, streaming, and a live cost/latency dashboard.
Deliberately second, not last. Everything after this sprint gets measured.
Ship #2: eval harness on real data, gating deploys in CI, with a baseline report on Ship #1.
The unglamorous 60% is document processing. That is where most RAG projects actually fail.
Ship #3: deployed RAG over a genuinely messy corpus, with citations, abstention, and retrieval metrics in the Sprint 2 harness.
Build the loop yourself before a framework hides the bug.
Ship #4: deployed agent on a real multi-step task with tools, guardrails, budgets and full run tracing.
Fine-tuning is a decision to justify, not a trophy model.
Ship #5: fine-tuned model serving behind an endpoint, with an evidence-based report — including whether it was worth it.
Cut Ship #3 cost about in half without measurable quality loss. Then operate it.
Ship #6: hardened deployment with published latency and cost SLOs, tracing, alerting, and a runbook.
A real problem — ideally from your workplace — from statement to deployed system. The 20-minute defence is questions on trade-offs, failure modes and cost, not the demo.
Deliberately excluded: transformer internals beyond practical decisions, pre-training, RLHF, GPU kernel work, and tool surveys. You learn two or three tools properly. Anything that cannot be deployed in the sprint it is taught does not belong here. If you want that depth afterwards, take Becoming an AI Principal Engineer.
You are calling OpenAI or Anthropic APIs with prompt-and-hope. Target: owning an LLM feature end to end.
Strong systems skills, little LLM exposure. Target: running LLM services in production.
Pipelines and SQL, curious about GenAI. Target: retrieval systems over company data.
You review AI work you cannot fully assess. Target: making and defending architecture calls.
You do not need machine learning, maths beyond basic statistics, or prior model training. This is software engineering applied to LLMs.
Still choosing which architecture to ship? Start with LLM Apps: Architecture by Use Case. Prefer a topic-ordered path? Take Production Generative AI Systems. Already shipping and need judgment under incomplete information? Take Scenarios alongside this, not instead of it.
Same tier, same destination. That course is organised by topic; first deployment is late. This one is organised by shipped increments: first deploy in week 4, then every four weeks. Evaluation is its own sprint, second, gating everything after. You leave with seven deployed artefacts and more deployment scar tissue; they leave with more theoretical grounding. Choose based on what Monday morning actually asks of you.
28 weeks, 84 hours live: about 3 hours live each week plus about 5 hours of lab. Expect around 40% of contact time in labs and reviews. If it drifts back toward lectures, it becomes another topic-ordered course.
API credits for most sprints, plus a single GPU for Sprint 5. Materially lighter than an advanced principal-level track.
28 weeks, seven ships, a 20-minute defence on trade-offs. Create an account to enrol in the next cohort.
Enrol now