Integration programme
AI in your firm, without your clients' data ever leaving it
Your team could be using AI today. They should not, because they handle information that cannot circulate. We build and run the environment that resolves that contradiction, inside your own infrastructure: we choose the open model that fits your work, deploy it and serve it with the configuration we measured it under, until it reaches the criterion you set. With the same benchmark, the same documents and the same scoring we use on the provider models.
A real client, before the argument
A real client, their own infrastructure, and the same result as Claude Opus 5 and Sonnet 5.
A 20-person practice that could not use AI on its clients' case files. We analysed which open model fitted their real use cases and measured it on three of their case files, anonymised: the division of an estate, an income tax return and a severance settlement for unfair dismissal, five attempts each, end to end on their own node. The open model scores 100.0 out of 100. Claude Opus 5 and Claude Sonnet 5, measured with the same benchmark, the same documents and the same scoring, also score 100.0. That work is now resolved inside their perimeter, with the information never leaving their building and with no per-token bill.
- 15 de 15
- Answers with the exact amount and every piece of evidence tied to its source document
- 0,00 €
- Mean distance to the correct amount, across the three tasks
- 105 s
- Median latency, on the client’s own infrastructure
Measured on 20 September 2026 · Qwen3.8-Flash-Next at Q5_K quantisation on an own node
What it measures and what it does not: three extraction-and-calculation tasks over business documentation in Spanish, one request at a time. It does not measure concurrency, code generation or long-text summarisation. Every model in the table was measured with the same method.
| Model | Score out of 100 | Median latency |
|---|---|---|
| claude-opus-5Provider model, as a reference | 100,0 | — |
| claude-sonnet-5Provider model, as a reference | 100,0 | — |
| qwen3.8-flash-next-q5The one we deployed at the client | 100,0 | 105 s |
| ornith-1.5-35b-a3b-q8Did not measure all three tasks: its score does not compare | 97,0 | — |
| qwen3.6-35b-a3bDid not measure all three tasks: its score does not compare | 91,0 | — |
| glm-5.3-flash-aj-iq2The previous measurement, with a different open model | 75,0 | 172 s |
| claude-haiku-4-5Lightweight provider model. It does not reach the threshold | 68,7 | — |
Three of the measurements were run from a desktop tool rather than the automated harness, so they have no recorded latency and appear without it. The documents, the questions and the scoring are identical: the quality comparison is valid, the timing one only among those that do record it.
The model took a while to get there. The previous measurement, with another open model on the same benchmark, stopped at 75.0 and returned a tax return with a 148.00 € deviation, identical on all five attempts. What changed the result was not raising the cap or softening the test: it was changing the model and serving it with the sampling from its own model card. The report tells it in full, with the complete table of models and the actions split by scope.
The problem
Cost, compliance and reliability. Today you pick two.
When an organisation processes information with AI, three tensions appear at once and the usual answers resolve only two.
Cost
Repetitive volume on commercial models scales badly. What was affordable at small scale stops being so as volume multiplies. And the cost is unpredictable: you cannot foresee how many tokens your future activity will need, nor what they will cost on commercial models.
Compliance
GDPR, the EU AI Act, your ISO certification and your own clients' contracts all constrain where the information you handle may go. Sending it to an API outside the EU is a risk, or simply not permitted.
Reliability
Local models let you cut cost without affecting the outcome, when applied where they belong. Because degrading the quality you get is not an option when an error carries legal consequences.
What we do
The bulk of the work stays with you. The hard cases go where they need to.
We build and operate an AI environment inside your infrastructure, with open models resolving most of your volume without your information leaving your perimeter.
Cases the local model cannot resolve with sufficient confidence escalate to a commercial model or to human review, according to what each of your own clients' contracts permits. Every decision is recorded with its reason, which is what makes the process auditable: what you show an auditor is not a log of what happened, but why a given document did or did not leave.
You decide what may cross the threshold. We measure whether the system is still getting it right and adjust the rules when it stops. We do not hand you a system to maintain: we operate it, and the quality stays our responsibility.
Who it is for
It is not for everyone, and we would rather say so upfront.
This makes sense when both conditions hold. If only one does, we will tell you in the first conversation rather than sell you a project.
High, repetitive volume
Classifying, extracting from or indexing information on a recurring basis, not occasionally.
A real data constraint
GDPR or the EU AI Act. A certification, a sector regulation or, most often, contracts with your own clients that do not all say the same thing: one forbids their information leaving and another does not. Today that forces you to run two separate processes.
Who you work with
Two partners, no management layer, and training for the people who have to defend it.
Kairos Tek is two partners. One has spent more than twenty-five years creating digital product at Telefónica and, in parallel, spent a decade advising Europol, the European Commission and the Bank of Spain. The other has spent twenty-five writing code and has been CTO at eight companies; seven of those years as CTO of Del Romero Abogados, a law firm where everything crossing the systems is covered by legal professional privilege. There is no management layer between them and you, because there is nobody else.
That combination is exactly what this requires. Bringing AI into a regulated process handling sensitive information is not only a technical problem: it is deciding what may leave your perimeter, being able to defend it to your clients, and getting your team to actually use it. When you tell us your documents cannot leave your servers, you do not have to explain it to us: we have worked that way for seven years.
And we train your team, which is where most of these projects fall over. Not an operations manual: what changes for the person who used to classify everything and now reviews exceptions, how to defend to an auditor why a document did or did not leave, and how consistent judgement from your own people is what lets us measure whether the system is getting it right. We have been doing this for years with Europol, the Guardia Civil, the Bank of Spain, IE Business School, IEB, ESIC and AFI, among more than twelve institutions.
We are integrating this line alongside a limited number of organisations that have the need today, working with their teams and on their real processes. If that is your case, let us talk.
Not sure whether the AI Act applies to you? Take the free diagnostic →
Do you process sensitive information at volume?
A one-hour technical conversation is enough to know whether this fits. Tell us what information you process, at what volume, and what constraint you have.
It is the report the client receives, untouched: the full table of models measured on the same benchmark, negative control included, the result task by task with its evidence, the node’s verified configuration, the eleven pre-flight checks and the limits of what the measurement covers. It downloads without leaving any data.