Notes from the build
Practical writing on building AI products that survive production — how to scope them, how to evaluate them, what breaks, and what the work actually looks like inside each industry we ship in.
What is an AI-native product studio?
The term gets used loosely. Here is a precise definition, the four things that make a studio genuinely AI-native rather than an agency with a model behind a chat box, and how to tell the difference before you sign anything.
Buying AI
AI product studio vs software agency vs IT consultancy
A studio, an agency, and a consultancy will all quote you for the same AI project — and produce three different things. Here is what each is structurally good at, and the failure mode each one carries.
How to choose an AI development partner: an enterprise checklist
Every vendor deck looks the same. These are the questions that actually discriminate between a team that has run AI in production and a team that has run AI in a demo.
Engineering
AI agents in production: what actually breaks
An agent that works in a demo and an agent that runs unattended against real systems are different engineering problems. These are the failures that show up in week three.
LLM evaluation in practice: how to know your AI feature actually works
Without evaluation, every prompt change is a guess and every regression is invisible until a customer finds it. Here is how to build a harness that is small enough to actually maintain.
Private and on-premises LLM deployment for regulated industries
Not every compliance requirement means self-hosting a model. Here are the four real options, ordered by cost and control, and how to work out which one your constraint actually demands.
RAG vs fine-tuning: which one does your use case actually need?
The question is usually posed as a choice between two techniques. It is better understood as a diagnosis: is the model missing knowledge, or missing behaviour?
Playbooks
From AI proof of concept to production: what the gap actually contains
The prototype impressed everyone in the room and then nothing happened for eight months. That gap has a predictable content, and almost none of it is model quality.
How to scope an AI project so it actually ships
AI projects rarely die from a hard engineering problem. They die from a scope that was never testable in the first place. Here is a method for writing one that is.
Industry
AI for customer support: deflection without wrecking satisfaction
The fastest way to ruin a support operation is to optimise a bot for deflection. The interventions that actually work start behind the scenes, with the agent rather than the customer.
AI for healthcare operations: what is actually deployable today
The AI that reaches production in healthcare is rarely the AI that makes the news. It sits in the administrative layer, where the error cost is recoverable and the workload is crushing.
AI in logistics: the use cases that survive contact with the yard
Logistics is unusually well suited to AI and unusually unforgiving of it. The difference between the systems that stick and the dashboards that get ignored is fairly predictable.
Computer vision for industrial safety: what it takes to run on a real site
Detecting a hard hat in a photograph is a solved problem. Running that detection across twenty cameras on a working site, at night, without generating alerts nobody reads, is not.
Have a problem worth writing about?
Most of these came out of real builds. Tell us what you're working on and we'll tell you honestly whether we're the right team.
Start the conversation