FutureStackDev AI Agent Development

AI Testing & QA

Golden sets, judges you calibrate, property tests, and CI that can fail a prompt or a model — QA for systems that do not always say the same thing.

Created by Baljeet Dogra

Course Objectives

  • Build a frozen set that is not your demo script.
  • Put deterministic checks in front of a judge.
  • Calibrate a judge against humans.
  • Fail the build on a real drop, not on noise.

What you will produce

Harness

QA suite

Frozen set, checks, judge, CI, a written stop-ship rule.

Bug

Found regression

A change that “looked better” and failed the suite.

4 weeks curriculum

Expand a part for the syllabus. Content stays searchable when closed.

01 What good means Week 1

Specs, oracles, when there is no single right answer.

  • Oracles
  • Specs
  • Risk
  • Coverage
02 Sets and checks Week 2

Sourcing, labels, schema and constraint checks.

  • Golden sets
  • Labels
  • Checks
  • Leakage
03 Judges and humans Week 3

Bias, agreement, when a human is non-negotiable.

  • Judges
  • Agreement
  • Sampling
  • Noise
04 CI Week 4

Flake, sample size, the gate. Capstone: a caught regression.

  • CI
  • Flake
  • Stop-ship
  • Capstone

Who this is for

QA engineers

Your playbook assumes determinism. Generative systems do not.

AI engineers

You ship on vibes. This is the discipline that replaces that.

Prerequisites

  • You have tested software before
  • Python enough to run a harness
  • Access to an LLM API

Not a fit if

  • A selenium-only web QA path
  • People who will not maintain a frozen set

Related: Prompt Engineering Mastery · Applied GenAI Engineering

Questions

Is this MLOps?

Testing the behaviour. MLOps Deep Dive is pipelines and promotion. You will overlap on gates.

Will we write unit tests?

Yes, for the deterministic bits. The rest is evals with statistics.

Ready to start?

Four weeks to a suite that can fail the build. Create an account to enrol.

Enrol now