insights

What keeps AI agents in production

The gap between a demo and a system is not intelligence. It is infrastructure.

category >

ai & agents

read >

6 min

updated >

Anyone can demo an AI agent now; that stopped being impressive a while ago. The interesting question is what the systems still running a year later have in common. Across the ones we operate, the answer is consistent — and none of it is glamorous.

An eval suite instead of a vibe check

Every production system we run has a set of real cases with known-good outputs, and every change — prompt, model, retrieval, schema — runs against it before it ships. Without that, "we improved the prompt" means "it felt better on the three examples we tried". Model upgrades are the sneaky one: a newer, smarter model can quietly break formatting or tone that downstream steps depend on. The eval suite catches it; vibes do not.

Guardrails, and a human on the edges

Structured outputs validated against schemas. Hard bounds on anything that touches money. And a designed lane for the cases the system should not decide: our document pipeline clears the backlog precisely because the odd malformed document goes to a person instead of jamming the machine or — worse — being guessed at. The human review lane is not a temporary crutch to remove later; it is a permanent load-bearing part of the design.

You cannot fix what you cannot replay

When someone says "the assistant gave a weird answer yesterday", the difference between a five-minute fix and a shrug is whether you can replay exactly what happened: the input, the retrieved context, the model call, the output, the validation result. Every step logged, every decision traceable. This is ordinary software engineering discipline applied to systems people keep treating as magic.

Inputs drift. Someone has to notice.

The world the system was built against changes: suppliers redesign their invoices, product catalogs grow, users start asking new kinds of questions. Production systems need the boring feedback loop — corrections from the human lane flowing back into evals, metrics that make drift visible before users complain, and an owner whose job includes looking at them.

The boring stack wins

Notice what is missing from this list: exotic frameworks, autonomous swarms, anything from this week’s launch videos. The systems that survive are narrow, observable, validated, and owned. That is also why estimating them is possible at all — infrastructure is predictable in a way that research is not.

faq

Questions

A company is AI-native when AI is part of how the product works and how the work gets done — not a feature added at the end. It changes what you build and how your team operates. We work that way ourselves, which is why we can tell you what it costs.

If you already know what to build, start with software. If you do not, start with consulting. Most companies start with a two-week assessment and move straight into a build.

One call. We look at your business, product and operations, then send a short written plan with scope, timeline and price. No questionnaire, no discovery deck.

An assessment takes two weeks. A first system usually runs in production within four to eight weeks. Larger programmes run as a monthly engagement.

Companies with a real operational or product problem — funded startups through to established mid-size companies. We are based in Portugal and work remotely across Europe.

rúben martins, founder

everything you need to know before the call

Got another

question?