How we know an AI system works before it ever touches your business.
The demo is the easy part. The real engineering is proving a system is reliable before it makes a single decision that matters.
Anyone can build an AI demo that works once. Type a question, get an impressive answer, everyone in the room nods. The distance between that and a system you would trust to run part of your business is enormous, and it is almost entirely about one thing that never makes it into the demo: evaluation. Knowing, with evidence, that the system works before it is allowed anywhere near a real decision.
Why a single good answer proves nothing
Language models do not behave like ordinary software. The same system can answer well nine times and badly the tenth. A demo shows you one of the nine. What matters in production is the tenth, because that is the one that quotes the wrong price, misreads a customer, or approves something it should not. A system that is right ninety per cent of the time sounds fine until you count how many decisions the other ten per cent represents across a full year of real volume.
In a demo you are shown the best case. In production you live with the worst case. Serious AI work is the discipline of measuring the worst case before it costs you anything.
What evaluation actually involves
Evaluating an AI system means building a test set of real cases from your operation, including the hard and unusual ones, and measuring how the system performs across all of them rather than only the easy ones. It means defining what correct means for your specific business before you start, so success is a number you can point at rather than a feeling in the room. And it means running those tests again every time the model or the system changes, so an upgrade that helps in general but happens to hurt your particular case gets caught before it ever ships.
- A test set drawn from your real operation, including the edge cases that quietly break naive systems.
- A clear definition of correct for your business, so performance is measured rather than guessed.
- Repeatable tests that run on every change, so nothing gets worse without someone noticing.
- Honest reporting of where the system fails, because a system with no known failures has simply not been tested hard enough.
This is the part that separates a lab from a shop
A software shop builds the feature, shows the demo, and ships. A research lab assumes the system is wrong until the evidence says otherwise, and treats the measurement as the actual product. That is the standard we hold our own work to, because it is the only standard under which you can responsibly let an AI system touch real money and real customers. The demo earns attention. The evaluation earns trust, and trust is the thing you are actually buying.
Before we would put a system into your operation, we can tell you how often it is right, where it fails, and what happens when it does. If a supplier cannot tell you that, they have shown you a demo, not a system, and the difference is your risk to carry.

Marc O'Brien
Co-founder & Managing Director, ACMR

