Engineering

Eval-driven development: the only way we ship agents

We write the test suite before the agent. Here is how a golden set, a grader and a CI gate keep our AI systems honest.

Author
Azed AI
Published
18 Sept 2026
Reading time
7 min

Most AI projects fail quietly. The demo works, the pilot looks promising, and then real users arrive with inputs nobody imagined. The system starts making confident mistakes, and nobody can say whether the latest prompt change made things better or worse.

Our answer is boring on purpose: before we write an agent, we write the tests it has to pass.

Start with a golden set

A golden set is a few hundred real examples with the answer you expect. We build it with the client in the first week, from their actual tickets, documents or transactions. It includes the easy cases, the weird ones, and the ones that cost money when they go wrong.

Grade automatically, review by hand

Every example gets a grader: an exact match where possible, a rubric scored by a second model where not. We review a sample of the model-graded results by hand each week so the grader itself stays honest.

If we cannot prove it works under adversarial inputs, we will not ship it.

Gate every change

The eval suite runs in CI. A prompt tweak, a model swap or a new tool only merges if the score holds. That single rule is what lets us move fast in month six without breaking what worked in month one.