Eval-driven development: the only way we ship agents
We write the test suite before the agent. Here is how a golden set, a grader and a CI gate keep our AI systems honest.
Chunking, hybrid search and citations: the unglamorous details that decide whether a chatbot is trusted.
Retrieval-augmented generation looks simple in a tutorial: split documents, embed them, search, answer. In production, every one of those steps hides a failure mode.
Fixed-size chunks cut tables in half and separate questions from answers. We split on document structure — headings, list items, table rows — and keep a link back to the source section.
Vector search is great at meaning and poor at exact terms like product codes. Combining it with keyword search fixes most “it couldn’t find the obvious thing” complaints.
Every answer links to the passage it came from. Citations make wrong answers easy to spot and right answers easy to trust.
We write the test suite before the agent. Here is how a golden set, a grader and a CI gate keep our AI systems honest.
About one in four strategy engagements ends with us recommending something more boring. Three questions we ask first.
Speed, privacy and offline support make on-device models compelling. Here is where they fit — and where the cloud still wins.