Eval-driven development: the only way we ship agents
We write the test suite before the agent. Here is how a golden set, a grader and a CI gate keep our AI systems honest.
Speed, privacy and offline support make on-device models compelling. Here is where they fit — and where the cloud still wins.
Phones now ship with neural engines that can run vision and small language models in real time. For many features, that changes the architecture of a mobile app.
Anything that needs to react instantly — pose tracking, camera effects, live transcription — belongs on the device. So does anything sensitive: if video never leaves the phone, there is far less to secure.
Long reasoning, large knowledge bases and anything that must stay consistent across devices still runs best on a server. Our usual pattern is a hybrid: fast perception on the phone, planning and memory in the cloud.
The hardest part is not the model, it is the experience when the network drops. We design every AI feature with an offline state from day one.
We write the test suite before the agent. Here is how a golden set, a grader and a CI gate keep our AI systems honest.
About one in four strategy engagements ends with us recommending something more boring. Three questions we ask first.
Chunking, hybrid search and citations: the unglamorous details that decide whether a chatbot is trusted.