Block A
Cold scenario
The brief arrives live. Fifteen minutes to form a position, then facilitated cold-calling and cross-examination. Escalations drop in as you get comfortable. Discomfort is the mechanism.
Every session, you are put in a room where something has gone wrong and asked what you would do about it. The brief is incomplete. Someone senior is unhappy. There is no obviously correct answer.
Created by Baljeet Dogra
The other programmes teach you to build systems. This one teaches you to decide when the situation is ambiguous, the data is incomplete, and people are watching. That is what the job actually is, and it is the thing interviews try — usually badly — to test for.
There are no lectures. Two 90-minute blocks each week. Roughly 48 facilitated scenarios over the programme, drawn from a bank of 150+, plus asynchronous solo scenarios. Cohort capped at 16–20: cold-calling does not work at scale.
Block A
The brief arrives live. Fifteen minutes to form a position, then facilitated cold-calling and cross-examination. Escalations drop in as you get comfortable. Discomfort is the mechanism.
Block B
Hidden structure revealed, failure modes catalogued, the real outcome discussed. You rewrite your position as a one-page decision memo, peer-reviewed before the next session.
A scenario course fails the moment it becomes a quiz show with an answer key. Participants are never told which type they are facing.
~40%
There is a defensibly correct answer, and most people get it wrong. Builds diagnostic skill.
~40%
Competent engineers genuinely disagree. The reasoning is what is assessed.
~20%
The correct answer is “don’t build this,” “do nothing yet,” or “this isn’t an AI problem.” Saying no is the rarest senior skill.
Every scenario uses the same anatomy: a half-page brief as it would actually arrive; facts withheld unless you ask; a hidden structure revealed after the attempt; two or three escalations mid-discussion; a rubric of reasoning moves, not an answer key; common failure modes; and, where it is drawn from a real incident, what actually happened — including when the “correct” answer lost.
About 3 hours live each week. Expand a block for the syllabus. Content stays searchable when closed.
Before any domain content: how to attack a situation you do not understand.
Something is broken and nobody knows why.
Skill: isolate the failing stage with evidence before touching anything.
Design it, but you cannot have what you want.
Skill: choosing the least-bad option and articulating what you gave up.
It is broken now, and it is expensive. Timed exercises with a facilitator playing an increasingly agitated stakeholder.
Skill: triage, containment, communication, then root cause — in that order. Includes the postmortem and the customer communication.
Someone is attacking you, or your system is harming someone.
Skill: threat modelling, and escalating something you found that nobody asked you to look for.
The hardest scenarios in the course, and the ones interviews never test.
Skill: disagreeing upward with evidence, and knowing when to accept a decision you lost.
You choose one vertical and work a concentrated set of that sector’s pressures — regulatory, data, failure-cost and cultural. Each track pairs a domain practitioner with the facilitator.
Full mock loops under realistic conditions, recorded and reviewed. The panel includes at least one external interviewer you have never met. Written feedback against a published rubric.
Shortened, to show the format. Friday, 16:40.
Brief. A message from the Head of Customer Operations: “The assistant is confidently giving customers wrong refund amounts. Started sometime this week. We’ve had 6 complaints. Can you look before Monday? We can’t turn it off, it’s handling 70% of contacts.”
Withheld unless asked. Nothing was deployed this week — but the model provider silently updated the default endpoint alias on Tuesday. The refund policy was updated on Monday, in a table with merged cells. There is no eval suite; quality is monitored by CSAT, which lags four days. “Confidently wrong” means it is citing a real clause and misreading a number.
Escalations. At 17:30 legal asks whether the wrong amounts were honoured. The CEO asks for a public statement. Your first hypothesis is disproven at 18:15.
A strong response asks about the change window before proposing a fix, triages containment before root cause, separates the two candidate causes with a cheap test, recognises the merged-cell table as a parsing failure not a model failure, and treats the missing eval suite as the actual incident — this went undetected for four days.
Common failure modes: rewriting the prompt first · assuming a model regression because it is fashionable · promising a Monday fix with no diagnosis · treating the parsing bug as the whole story.
Real outcome: both causes were real and compounding. The team fixed the parsing in two hours and spent three weeks building the eval suite they should have had.
Scored on reasoning quality, not conclusion. A participant reaching a defensible position the facilitator disagrees with scores full marks. A participant reaching the “right” answer by luck does not.
You have shipped LLM features. Interviews and Friday outages still feel like a different sport. Target: deciding under incomplete information.
You are in the room when a VP has promised something, or a vendor demo cannot be reproduced. Target: disagreeing upward with evidence.
You do not need advanced depth. This runs well alongside a build programme rather than only after it.
Need to choose an architecture first? Take LLM Apps: Architecture by Use Case. Need to learn to ship first? Take Applied GenAI Engineering (ship-first) or Production Generative AI Systems (topic-ordered). This course is the overlay: can you decide when people are watching?
LLM Apps: Architecture by Use Case asks which architecture, and what will go wrong. Applied GenAI Engineering asks whether you can build it. Production Generative AI Systems asks whether you understand it deeply enough to be accountable. This one asks whether you can decide, under pressure, with incomplete information, when people are watching. It is the overlay that transfers directly into interview performance.
16 weeks, 48 hours live: two 90-minute blocks each week, plus asynchronous solo scenarios and a weekly decision memo. Cohort capped at 16–20.
Sometimes. About 40% of the bank is determinate, 40% contested, 20% traps where the right move is not to build. You are never told which. Asking the right withheld question is scored — it is the single most predictive seniority signal in the course.
16 weeks, 48 live scenarios, a mock interview loop with someone you have never met. Create an account to enrol in the next cohort.
Enrol now