Eval-driven development: the only way we ship agents
We write the test suite before the agent. Here is how a golden set, a grader and a CI gate keep our AI systems honest.
Engineering notes, strategy and honest opinions from inside real builds. No hype, no listicles.
We write the test suite before the agent. Here is how a golden set, a grader and a CI gate keep our AI systems honest.
About one in four strategy engagements ends with us recommending something more boring. Three questions we ask first.
Speed, privacy and offline support make on-device models compelling. Here is where they fit — and where the cloud still wins.
Chunking, hybrid search and citations: the unglamorous details that decide whether a chatbot is trusted.
What building AZ POS taught us about where AI coding tools speed you up — and where they quietly slow you down.
Model prices fall every quarter, yet bills keep rising. Four habits that keep AI features profitable.