When the answer is not an LLM (and how to tell)
About one in four strategy engagements ends with us recommending something more boring. Three questions we ask first.
Model prices fall every quarter, yet bills keep rising. Four habits that keep AI features profitable.
A feature that costs two cents a call is fine at a thousand users and painful at a million. Cost has to be designed in, not discovered on the invoice.
Send easy requests to a small, cheap model and only escalate hard ones. In most systems we run, 70–80% of traffic never needs the largest model.
Many questions repeat. Caching answers and prompt prefixes is the cheapest optimisation there is.
Tag every call with the feature that made it. You cannot fix the expensive feature if all you see is one monthly total.
About one in four strategy engagements ends with us recommending something more boring. Three questions we ask first.
We write the test suite before the agent. Here is how a golden set, a grader and a CI gate keep our AI systems honest.
Speed, privacy and offline support make on-device models compelling. Here is where they fit — and where the cloud still wins.