You cannot ship LLM-backed features without a way to know when you are getting worse. Most teams paste outputs into a spreadsheet for a week, then stop. This lab builds a 200-line evaluation harness that runs every prompt change against a fixed dataset, scores results with property-based assertions (no human grading), and emits a diff you can put in a pull-request comment.
Eval suites are one piece. The full course covers cost guardrails, fallback strategies, and the day-2 ops of running LLMs in a system you have to support.
The first time you ship an LLM-backed feature, you copy a few outputs into a Google Sheet, score them by hand, and feel responsible. The fifth time you change the prompt, the sheet has 400 rows, nobody remembers the column conventions, and you stop opening it. From that point on, every prompt change is a vibes-based decision.
The cure is not "review more carefully." The cure is a tiny harness that runs every prompt change against a fixed dataset, scores the results with code, and prints a diff. Two hundred lines. Five files. No human grading.
A dataset is a JSON file. The smallest useful one has 20 entries, each with three keys: id, input, and expected. The expected field is not the exact correct answer; it is a set of properties the answer must satisfy.
The discipline is to build the dataset from real production traces, not synthetic examples. Pull 20 real inputs from your logs, write down what makes the answer "good enough," and you have a baseline that catches regressions in the actual surface area users hit.
Bounded concurrency is a quiet decision: too low and the run takes 20 minutes, too high and you trip rate limits or the request budget. Four is a fine default for most APIs.
The language detection is intentionally crude. Property-based scoring is not about perfect classification; it is about catching the case where the model returns English when you asked for Arabic, which happens more than you would think.
A two-line GitHub Action turns the harness into the gate that prevents prompt regressions from shipping:
- run: npm run eval -- --prompt prompts/v3.txt --threshold 0.85- run: npm run eval -- --prompt prompts/v3.txt --diff prompts/v2.txt
The --threshold flag fails the job if the pass rate drops below 85%. The --diff flag runs both prompts and prints the per-id deltas. Together they prevent the most common LLM regression: the prompt looks better in three spot-checks, ships, and fails on the long tail.
This is the runner, the scorers, and the diff. It is not the cost guardrail (per-call token accounting and budget alerts), not the fallback strategy (when do you stop calling the model and serve a cached answer), and not the day-2 ops (how do you update the dataset weekly without breaking the baseline). The full AI-paired engineering course covers all three on top of the same harness.