Adapter
Served LoRA
Endpoint plus a report against a frozen baseline.
Quantisation, distillation, LoRA — and an honest verdict on whether the smaller or tuned model was worth it.
Created by Baljeet Dogra
Adapter
Endpoint plus a report against a frozen baseline.
Opt
Size, latency, quality — three numbers, not a blog claim.
Expand a part for the syllabus. Content stays searchable when closed.
Prompt, retrieve, tune. What each fixes. Data quality first.
A real run, hyperparameters, overfitting, evals including regressions.
INT8/INT4, distillation, routing cheap traffic.
Adapter behind an API. Capstone: worth it or not, in writing.
You can prompt and retrieve. Fine-tuning is still a trophy in your head.
You want PEFT and quantisation with production numbers.
Related: Applied GenAI Engineering · Custom LLM Training
No. An honest, evidenced no is a pass. Hiding regressions is not.
A small open model you can actually serve. Frontier fine-tunes are discussed, not required.