Skip to content
ALI HAITHAM·TECH·responds in <1h
  • Homestart here
  • Workcase studies, projects
  • Serviceswhat I will build for you
  • Writingessays, series
  • Traincourses, labs, cohorts
  • Aboutthe engineer behind this site
Sign inStart a project
ALI HAITHAM · TECH
  • Home↗
  • Work↗
  • Services↗
  • Writing↗
  • Train↗
  • About↗
Sign inStart a project
Online
ALI HAITHAM·TECH

Engineering studio. Damascus, GMT+3.

aliyosef.online

Studio

  • Work
  • Writing
  • Training
  • About
  • Contact me

Portal

  • Sign in
  • Open a ticket
  • Track project
  • My account

Resources

  • Docs
  • Status
  • Changelog
  • Brand kit
  • Privacy
  • Terms

Newsletter

Field notes and tech news. Weekly. No fluff.

Free. Unsubscribe anytime.

Find me elsewhere
© 2026 Ali Haitham Yosef. All rights reserved.Hand-built in React 19. No frameworks of frameworks.Last deployed · 2026-05-08All systems operational
  1. Home/
  2. Train/
  3. Labs/
  4. Set up an LLM eval suite
LAB04

Set up an LLM eval suite

You cannot ship LLM-backed features without a way to know when you are getting worse. Most teams paste outputs into a spreadsheet for a week, then stop. This lab builds a 200-line evaluation harness that runs every prompt change against a fixed dataset, scores results with property-based assertions (no human grading), and emits a diff you can put in a pull-request comment.

Start the lab↓Open the repo↗
min
90
STACK
Node
LEVEL
Advanced
/02 — SETUP

What you will build

A 200-line evaluation harness for non-deterministic tools. Property-based scorers, no human grading.

Prerequisites

  1. 01Node 20+ and an OpenAI- or Anthropic-compatible API key.
  2. 02You have made at least one production API call to an LLM.
  3. 03A handful of real prompt-input pairs from your app to seed the dataset.
/03 — WALKTHROUGH
/04 — RESOURCES

Take it with you.

  • ⌥
    RepositoryStarter code on GitHub.
  • ↓
    Zip downloadSame code, no git required.
/USED IN

Courses that pair with this lab.

  • AI-paired engineering→
/05 — NEXT
Try this next01Build a queue in 60 minutesAnother short, free lab.→
Or go deep

AI-paired engineering

Eval suites are one piece. The full course covers cost guardrails, fallback strategies, and the day-2 ops of running LLMs in a system you have to support.

See the course→

Why most teams stop measuring after week one#

The first time you ship an LLM-backed feature, you copy a few outputs into a Google Sheet, score them by hand, and feel responsible. The fifth time you change the prompt, the sheet has 400 rows, nobody remembers the column conventions, and you stop opening it. From that point on, every prompt change is a vibes-based decision.

The cure is not "review more carefully." The cure is a tiny harness that runs every prompt change against a fixed dataset, scores the results with code, and prints a diff. Two hundred lines. Five files. No human grading.

Step 1: the dataset#

A dataset is a JSON file. The smallest useful one has 20 entries, each with three keys: id, input, and expected. The expected field is not the exact correct answer; it is a set of properties the answer must satisfy.

[
  {
    "id": "1",
    "input": "Summarise this support ticket: ...",
    "expected": {
      "must_contain_any": ["refund", "return", "exchange"],
      "max_words": 80,
      "must_be_in_language": "en"
    }
  },
  {
    "id": "2",
    "input": "Translate this Arabic invoice description: ...",
    "expected": {
      "must_be_in_language": "ar",
      "must_not_contain": ["TODO", "[INSERT", "<", ">"]
    }
  }
]

The discipline is to build the dataset from real production traces, not synthetic examples. Pull 20 real inputs from your logs, write down what makes the answer "good enough," and you have a baseline that catches regressions in the actual surface area users hit.

Step 2: the runner#

The runner is a small async loop that calls your LLM client for every input and yields a Result:

import { OpenAI } from 'openai';

interface DatasetEntry {
  id: string;
  input: string;
  expected: Record<string, unknown>;
}

interface Result {
  id: string;
  output: string;
  expected: DatasetEntry['expected'];
}

const client = new OpenAI({ apiKey: process.env.OPENAI_API_KEY });

async function runOne(entry: DatasetEntry, prompt: string): Promise<Result> {
  const resp = await client.chat.completions.create({
    model: 'gpt-4o-mini',
    messages: [
      { role: 'system', content: prompt },
      { role: 'user', content: entry.input },
    ],
  });
  return {
    id: entry.id,
    output: resp.choices[0].message.content ?? '',
    expected: entry.expected,
  };
}

export async function runAll(
  dataset: DatasetEntry[],
  prompt: string,
  concurrency = 4,
): Promise<Result[]> {
  const results: Result[] = [];
  let cursor = 0;
  await Promise.all(
    Array(concurrency).fill(0).map(async () => {
      while (cursor < dataset.length) {
        const idx = cursor++;
        results[idx] = await runOne(dataset[idx], prompt);
      }
    }),
  );
  return results;
}

Bounded concurrency is a quiet decision: too low and the run takes 20 minutes, too high and you trip rate limits or the request budget. Four is a fine default for most APIs.

Step 3: the property-based scorers#

A scorer is a pure function from (output, expected) to { ok: boolean; reason?: string }. Five of them cover most of what you actually care about:

type Score = { ok: boolean; reason?: string };

const scorers = {
  must_contain_any: (out: string, expected: string[]): Score =>
    expected.some((s) => out.toLowerCase().includes(s.toLowerCase()))
      ? { ok: true }
      : { ok: false, reason: `none of: ${expected.join(', ')}` },

  must_not_contain: (out: string, expected: string[]): Score => {
    const hit = expected.find((s) => out.includes(s));
    return hit ? { ok: false, reason: `contains: ${hit}` } : { ok: true };
  },

  max_words: (out: string, expected: number): Score => {
    const n = out.trim().split(/\s+/).length;
    return n <= expected
      ? { ok: true }
      : { ok: false, reason: `${n} > ${expected} words` };
  },

  min_chars: (out: string, expected: number): Score =>
    out.length >= expected
      ? { ok: true }
      : { ok: false, reason: `${out.length} < ${expected} chars` },

  must_be_in_language: (out: string, expected: 'ar' | 'en'): Score => {
    const arabic = /[؀-ۿ]/.test(out);
    const isAr = arabic;
    return (expected === 'ar') === isAr
      ? { ok: true }
      : { ok: false, reason: `expected ${expected}, looks ${isAr ? 'ar' : 'en'}` };
  },
};

The language detection is intentionally crude. Property-based scoring is not about perfect classification; it is about catching the case where the model returns English when you asked for Arabic, which happens more than you would think.

Step 4: scoring + the diff#

Apply every scorer that matches a key in expected, collect failures, and emit a structured result. Below is the orchestrator:

function scoreOne(result: Result): { id: string; pass: boolean; failures: string[] } {
  const failures: string[] = [];
  for (const [key, expected] of Object.entries(result.expected)) {
    const fn = scorers[key as keyof typeof scorers];
    if (!fn) continue;
    const s = fn(result.output as never, expected as never);
    if (!s.ok) failures.push(`${key}: ${s.reason}`);
  }
  return { id: result.id, pass: failures.length === 0, failures };
}

export function summarise(results: Result[]) {
  const scored = results.map(scoreOne);
  const passed = scored.filter((r) => r.pass).length;
  const total = scored.length;
  return {
    passRate: passed / total,
    failed: scored.filter((r) => !r.pass),
  };
}

Run it twice (once with the old prompt, once with the new), then JSON-diff the two summaries. That diff is what goes in the pull-request comment.

Step 5: the CI hook#

A two-line GitHub Action turns the harness into the gate that prevents prompt regressions from shipping:

- run: npm run eval -- --prompt prompts/v3.txt --threshold 0.85
- run: npm run eval -- --prompt prompts/v3.txt --diff prompts/v2.txt

The --threshold flag fails the job if the pass rate drops below 85%. The --diff flag runs both prompts and prints the per-id deltas. Together they prevent the most common LLM regression: the prompt looks better in three spot-checks, ships, and fails on the long tail.

What this lab is not#

This is the runner, the scorers, and the diff. It is not the cost guardrail (per-call token accounting and budget alerts), not the fallback strategy (when do you stop calling the model and serve a cached answer), and not the day-2 ops (how do you update the dataset weekly without breaking the baseline). The full AI-paired engineering course covers all three on top of the same harness.