Evidence
Eval harness
Baseline and post-change results on a frozen set, including the improvement that made things worse.
For engineers who have shipped at least one LLM-backed feature and are ready to own a system end to end: architect it, prove it, cost it, operate it, and lead others to do the same.
Created by Baljeet Dogra
Most advanced AI courses stack more techniques. That is not the gap. The gap between senior and principal is judgment, systems, evaluation and leverage. Every part below has a technical track and a judgment track.
One constrained system, carried from week 12. Format each part: concepts, guided lab, open-ended build, written artefact, review. The capstone is weighted so a demo with no evidence cannot pass.
Evidence
Baseline and post-change results on a frozen set, including the improvement that made things worse.
Cost
Cost per request, per user, per resolved ticket — defensible to a CFO, with a 60% cost-cut lab behind it.
Decisions
Rejected options and reverse conditions. Plus FMEA, platform design, and a risk-and-control map.
Operations
Kill switch, incident comms, two-year roadmap that survives a leadership change.
Capstone
Working, instrumented, in a domain with real constraints. Assessment: system 30%, evidence 30%, writing 20%, defence 20%.
Twelve parts. Also offered as a 12-week intensive. Each classroom lesson has a talk, a diagram, a lab or artefact, and a scenario quiz. Expand a part for the syllabus. Content stays searchable when closed.
How the classroom works, then a map of the twelve parts, then Calibration: where you sit, three archetypes, a capability matrix.
Deliverable: personal gap analysis and a 6-month development plan.
Attention, KV cache, tokenisation, sampling, precision, the model landscape, embeddings. Reason from first principles, not vendor docs.
Lab: instrument a request end to end. Judgment: one-page memo on why you are (not) using open weights.
The highest-leverage part. Golden sets, offline metrics, LLM-as-judge, statistics, online evals, CI gates, task-specific harnesses.
Lab: prove an “obvious improvement” made things worse. Deliverable: eval strategy including what you will stop shipping without.
Chunking, hybrid retrieval, query understanding, structured sources, freshness, permission-aware retrieval, context economics, RAG diagnosis.
Lab: messy enterprise corpus to a measured target. Judgment: cost/accuracy frontier of three architectures.
When not to use an agent. Tools, memory, multi-agent honesty, reliability for non-determinism, guardrails, HITL, tracing.
Lab: build a consequential agent, then break it. Deliverable: FMEA.
The ladder from prompting to continued pretraining. Data curation, synthetic data, PEFT, preference optimisation, distillation, honest eval, maintenance burden. Capstone starts here.
Lab: fine-tune a small model and prove whether it beats prompting a larger one on quality, cost and latency together.
TTFT vs TPOT, batching, caching, routing, managed vs self-hosted, capacity, unit economics, load tests.
Lab: cut cost 60% with under 5% quality loss. Deliverable: unit economics model.
Gateway, registry, eval service, versioned prompts, CI for stochastic systems, tracing, drift, incidents, on-call.
Deliverable: platform design document with build-vs-buy decisions.
Attack surface, lethal trifecta, defence in depth, red teaming, privacy, EU AI Act / ISO 42001 / NIST AI RMF, model risk, fairness that is measurable.
Deliverable: risk assessment and control mapping for the capstone.
Model churn, build vs buy vs wait, legacy coexistence, multi-tenancy, reading the field, two-year roadmaps, reversible vs irreversible decisions.
Deliverable: five ADRs with rejected options and reverse conditions.
Writing, influence without authority, exec briefing, estimation under uncertainty, mentorship, hiring, killing work without burning capital.
Labs: design doc torn apart in review; 10-minute exec briefing; mock review as the reviewer.
Build, ship and defend one non-trivial system in a domain with real constraints. Eight required artefacts. Hostile panel, 45 minutes.
Running across all parts: a paper club (does this change a decision?), a failure journal, and a cost ledger from week 1. Lab stack is vendor-plural: at least one managed frontier API, one open-weights path, one framework and one raw-SDK path, one tracer, one eval tool.
You can wire an API and read a model card. You have not owned a system end to end or set direction for others.
You need evaluation, cost, security and writing as a single job — not another techniques course.
The three archetypes are named on purpose. Calibration makes you pick a narrative.
Need to choose an architecture before you own one at scale? Take LLM Apps: Architecture by Use Case. Have not yet shipped an LLM feature? Start with Production Generative AI Systems (topic-ordered) or Applied GenAI Engineering (ship every four weeks).
The gap between senior and principal is judgment under ambiguity, systems thinking, evaluation rigour, and leverage. Every part has a technical track and a judgment track. The judgment track is not filler.
24 weeks part-time at about 8–10 hours a week, or 12 weeks intensive. Live teaching plus labs in the week. The classroom holds talks, diagrams, briefs, notes and exercises between sessions.
A managed frontier API is enough for most weeks. The adaptation and serving labs need a GPU or a hosted equivalent. The stack stays vendor-plural on purpose: a principal must be able to move.
24 weeks from calibration to a defence a hostile panel cannot wave away. Create an account to enrol in the next cohort.
Enrol now