FutureStackDev AI Agent Development

AI Performance Optimisation

TTFT, TPOT, cache, batch, route — cut cost and latency with quality inside noise, and prove it on a frozen set.

Created by Baljeet Dogra

Course Objectives

  • Instrument a request until you know where the milliseconds go.
  • Cut cost about in half on a real workload without measurable quality loss.
  • Choose cache, batch or route with numbers, not slogans.
  • Load-test the path you will actually ship.

What you will produce

Cut

Cost/latency report

Baseline, levers, quality on the frozen set, £ per task.

Test

Load profile

p50/p95, saturation, the knob you will turn first in an incident.

4 weeks curriculum

Expand a part for the syllabus. Content stays searchable when closed.

01 Measure Week 1

TTFT, TPOT, queue, retrieval, network. A table before a lever.

  • Breakdown
  • Tracing
  • Tokens
  • Queues
02 Levers Weeks 2–3

Prompt cache, semantic cache, batch, trim, cascade, spec decode if it earns it.

  • Cache
  • Batch
  • Route
  • Trim
03 Prove it Week 4

Frozen eval, load test, capstone report.

  • Quality
  • Load
  • SLOs
  • Capstone

Who this is for

Engineers with a live LLM feature

The bill and the p95 are both too high. You need levers.

Performance engineers

You know web perf. Tokens are a different budget.

Prerequisites

  • You have an LLM-backed path to measure — or we provide a lab system
  • Comfortable with logs and dashboards

Not a fit if

  • Premature optimisation on a demo
  • Kernel-level CUDA work — out of scope

Related: Applied GenAI Engineering · LLM API Integration

Questions

GPUs required?

Only if you self-host. Managed APIs are enough for caching, routing and trimming labs.

Is 50% a promise?

It is the lab target on the provided workload. Your production number is evidence, not a slogan.

Ready to start?

Four weeks to a measured cut in cost and latency. Create an account to enrol.

Enrol now