FutureStackDev AI Agent Development

Becoming an AI Principal Engineer

For engineers who have shipped at least one LLM-backed feature and are ready to own a system end to end: architect it, prove it, cost it, operate it, and lead others to do the same.

Created by Baljeet Dogra

Course Objectives

Most advanced AI courses stack more techniques. That is not the gap. The gap between senior and principal is judgment, systems, evaluation and leverage. Every part below has a technical track and a judgment track.

  • Choose what not to build, and say why in writing — judgment under ambiguity.
  • Own latency, cost, failure modes and blast radius, not just accuracy.
  • Prove whether the system works: golden sets, judges you validate, evals that can block a release.
  • Design retrieval, agents and adaptation as a ladder you can justify, not a pile of tools.
  • Pass a security review and a regulator’s questions without theatre.
  • Multiply other engineers, then defend a deployed system with evidence a hostile panel cannot wave away.

What you will produce

One constrained system, carried from week 12. Format each part: concepts, guided lab, open-ended build, written artefact, review. The capstone is weighted so a demo with no evidence cannot pass.

Evidence

Eval harness

Baseline and post-change results on a frozen set, including the improvement that made things worse.

Cost

Unit economics

Cost per request, per user, per resolved ticket — defensible to a CFO, with a 60% cost-cut lab behind it.

Decisions

Five ADRs

Rejected options and reverse conditions. Plus FMEA, platform design, and a risk-and-control map.

Operations

Runbook and roadmap

Kill switch, incident comms, two-year roadmap that survives a leadership change.

Capstone

Deployed system and 45-minute defence

Working, instrumented, in a domain with real constraints. Assessment: system 30%, evidence 30%, writing 20%, defence 20%.

24-week curriculum

Twelve parts. Also offered as a 12-week intensive. Each classroom lesson has a talk, a diagram, a lab or artefact, and a scenario quiz. Expand a part for the syllabus. Content stays searchable when closed.

00 Getting started Week 1 · 3 lessons

How the classroom works, then a map of the twelve parts, then Calibration: where you sit, three archetypes, a capability matrix.

Deliverable: personal gap analysis and a 6-month development plan.

01 Foundations you cannot fake Weeks 2–3 · 1 lesson

Attention, KV cache, tokenisation, sampling, precision, the model landscape, embeddings. Reason from first principles, not vendor docs.

Lab: instrument a request end to end. Judgment: one-page memo on why you are (not) using open weights.

02 Evaluation as the core discipline Weeks 4–6 · 1 lesson

The highest-leverage part. Golden sets, offline metrics, LLM-as-judge, statistics, online evals, CI gates, task-specific harnesses.

Lab: prove an “obvious improvement” made things worse. Deliverable: eval strategy including what you will stop shipping without.

03 Retrieval and context engineering Weeks 7–8 · 1 lesson

Chunking, hybrid retrieval, query understanding, structured sources, freshness, permission-aware retrieval, context economics, RAG diagnosis.

Lab: messy enterprise corpus to a measured target. Judgment: cost/accuracy frontier of three architectures.

04 Agentic systems Weeks 9–11 · 1 lesson

When not to use an agent. Tools, memory, multi-agent honesty, reliability for non-determinism, guardrails, HITL, tracing.

Lab: build a consequential agent, then break it. Deliverable: FMEA.

05 Model adaptation Weeks 12–13 · 1 lesson

The ladder from prompting to continued pretraining. Data curation, synthetic data, PEFT, preference optimisation, distillation, honest eval, maintenance burden. Capstone starts here.

Lab: fine-tune a small model and prove whether it beats prompting a larger one on quality, cost and latency together.

06 Inference, serving and cost Weeks 14–15 · 1 lesson

TTFT vs TPOT, batching, caching, routing, managed vs self-hosted, capacity, unit economics, load tests.

Lab: cut cost 60% with under 5% quality loss. Deliverable: unit economics model.

07 AI platform and operations Weeks 16–17 · 1 lesson

Gateway, registry, eval service, versioned prompts, CI for stochastic systems, tracing, drift, incidents, on-call.

Deliverable: platform design document with build-vs-buy decisions.

08 Safety, security and governance Weeks 18–19 · 1 lesson

Attack surface, lethal trifecta, defence in depth, red teaming, privacy, EU AI Act / ISO 42001 / NIST AI RMF, model risk, fairness that is measurable.

Deliverable: risk assessment and control mapping for the capstone.

09 Architecture and technical strategy Weeks 20–21 · 1 lesson

Model churn, build vs buy vs wait, legacy coexistence, multi-tenancy, reading the field, two-year roadmaps, reversible vs irreversible decisions.

Deliverable: five ADRs with rejected options and reverse conditions.

10 Leverage: the principal skill set Weeks 22–23 · 1 lesson

Writing, influence without authority, exec briefing, estimation under uncertainty, mentorship, hiring, killing work without burning capital.

Labs: design doc torn apart in review; 10-minute exec briefing; mock review as the reviewer.

11 Capstone Week 24 · running from week 12

Build, ship and defend one non-trivial system in a domain with real constraints. Eight required artefacts. Hostile panel, 45 minutes.

Running across all parts: a paper club (does this change a decision?), a failure journal, and a cost ledger from week 1. Lab stack is vendor-plural: at least one managed frontier API, one open-weights path, one framework and one raw-SDK path, one tracer, one eval tool.

Who this is for

Engineers past the first LLM feature

You can wire an API and read a model card. You have not owned a system end to end or set direction for others.

Staff engineers aiming at principal

You need evaluation, cost, security and writing as a single job — not another techniques course.

Platform and product-embedded ICs

The three archetypes are named on purpose. Calibration makes you pick a narrative.

Prerequisites

  • Confident Python: async, typing, testing
  • Has shipped something LLM-backed, even if small
  • Basic cloud and containers
  • Basic statistics: distributions, hypothesis testing

Not a fit if

  • You want a first course in Python or prompting
  • You only need a management overview of AI
  • You want notebooks instead of evals, ADRs and a defence

Need to choose an architecture before you own one at scale? Take LLM Apps: Architecture by Use Case. Have not yet shipped an LLM feature? Start with Production Generative AI Systems (topic-ordered) or Applied GenAI Engineering (ship every four weeks).

Questions

How is this different from stacking more techniques?

The gap between senior and principal is judgment under ambiguity, systems thinking, evaluation rigour, and leverage. Every part has a technical track and a judgment track. The judgment track is not filler.

Is this 12 weeks or 24?

24 weeks part-time at about 8–10 hours a week, or 12 weeks intensive. Live teaching plus labs in the week. The classroom holds talks, diagrams, briefs, notes and exercises between sessions.

Do I need GPUs?

A managed frontier API is enough for most weeks. The adaptation and serving labs need a GPU or a hosted equivalent. The stack stays vendor-plural on purpose: a principal must be able to move.

Ready to own the system end to end?

24 weeks from calibration to a defence a hostile panel cannot wave away. Create an account to enrol in the next cohort.

Enrol now