Retrieval that holds up in production
Chunking, hybrid search and citations: the unglamorous details that decide whether a chatbot is trusted.
We write the test suite before the agent. Here is how a golden set, a grader and a CI gate keep our AI systems honest.
Most AI projects fail quietly. The demo works, the pilot looks promising, and then real users arrive with inputs nobody imagined. The system starts making confident mistakes, and nobody can say whether the latest prompt change made things better or worse.
Our answer is boring on purpose: before we write an agent, we write the tests it has to pass.
A golden set is a few hundred real examples with the answer you expect. We build it with the client in the first week, from their actual tickets, documents or transactions. It includes the easy cases, the weird ones, and the ones that cost money when they go wrong.
Every example gets a grader: an exact match where possible, a rubric scored by a second model where not. We review a sample of the model-graded results by hand each week so the grader itself stays honest.
If we cannot prove it works under adversarial inputs, we will not ship it.
The eval suite runs in CI. A prompt tweak, a model swap or a new tool only merges if the score holds. That single rule is what lets us move fast in month six without breaking what worked in month one.
Chunking, hybrid search and citations: the unglamorous details that decide whether a chatbot is trusted.
About one in four strategy engagements ends with us recommending something more boring. Three questions we ask first.
Speed, privacy and offline support make on-device models compelling. Here is where they fit — and where the cloud still wins.