Skip to main content

Model evaluation on your own infrastructure

You have your own AI node, but is it the one you need?

That question has gone unanswered for months while the node has been working. And if you have not built it yet, it is the one to answer before the purchase order, not after. We build the tests on your real work, agree the criterion with you, measure end to end on your infrastructure, and deliver a verdict with the evidence and the actions split by who executes them: the ones in your node and the ones in ours. And if you would rather not build it yourself, we choose it, build it and hand it over with the measurement already done on your own cases.

The problem

"It seems to work fine" is not a state.

When an answer comes out wrong there are four possible causes and none of them is visible from outside. Without separating them, every meeting about the node is an exchange of impressions, and the impression that wins is the loudest one.

The model

It is not good enough for that task, and no configuration will fix it.

The quantisation

The model would be good enough, but not at the bits per weight it is loaded with.

The node

A start-up flag, a context size, or a wrongly declared tokeniser.

The access layer

A timeout, a declared limit, or an output budget the node cannot meet.

If you have not built it yet, the problem is the same one step earlier: you are about to choose all four at once, committing money and a deadline, with no number telling you which of them decides the outcome.

And there is a fifth possibility, the worst of all: that it gets it right almost always and fails in one specific place, always the same one, with a figure credible enough that nobody looks at it twice.

The evidence, before the argument

A pass mark and a miscalculated settlement, in the same measurement.

This is our own node: an open-weights MoE model, on our hardware. Measured on 12 September 2026: three case files, five attempts each, everything through the same access layer the real work goes through, and 50 minutes of dedicated node from start to finish. The model, the node and the date are recorded together on purpose: without all three, the score cannot be produced again.

Open-weights MoE model · own node · 12 September 2026 · 5 attempts per case file · 50-minute measurement
Case fileScoreFigure returnedCorrect figureDeviationIdentical attempts
Inheritance · division of an estate100,0078.731,75 €78.731,75 €0,00 €5 of 5
Personal tax return25,003.872,79 €3.724,79 €148,00 €5 of 5
Severance for unfair dismissal100,0064.506,70 €64.506,70 €0,00 €5 of 5
Aggregate75,00 · passthreshold 70mean error 49,33 €complete and stable run

148.00 euros on 3,724.79. Four per cent. It is not an absurd figure: it is a credible one, which is why nobody would have caught it by reading the answer. The model returned it identically on all five attempts, so nobody would have caught it by asking again either. And the aggregate score (75.00 against a threshold of 70) would have passed it.

The failure is located, not guessed at. Of the six traceability checks on that case file, five pass: it gets the pension cap right, the net gain on the share sale right and the withholdings right, and it explicitly declares the non-deductible expense outside the calculation. Exactly one fails, the reduced rental income, which is precisely where the case hides its three expense traps. It goes wrong in one place, and always the same one.

A demonstration would have shown the inheritance and the severance. It would have taken ten minutes, it would have gone well, and it would not have been false: it would have been incomplete, which in a decision is worse. A pilot with no acceptance criterion agreed beforehand produces exactly that result, which is why it always goes well.

We can show this measurement in full because the workload is ours. It is no client's case file, so there is nothing to anonymise and no figure to trim.

How it works

Five steps. The first two need no hardware.

That is the part that tends to surprise: the criterion the node will be judged by can be built before the node exists. That is where the three ways in of the next section come from.

  1. The tests, on your work

    The cases are written from real case files of yours (the general work and the specific work of your teams) and the correct answer is calculated by hand, before any model sees them.

    No hardware
  2. The criterion, set by you

    What score is good enough, and why that one. With the floor measured (what nothing scores) and the ceiling measured against a provider model, so the number means something. And before the purchase order.

    No hardware
  3. The measurement, on your infrastructure

    End to end, through the same access layer the real work will go through: on the infrastructure you already have, or on the one we build. Five attempts per case, and the complete record of every run.

  4. The verdict

    One sentence that closes the question, the evidence beneath it, and the actions split by scope: the ones your team executes on the node and the ones we execute on the access layer.

  5. The iteration

    The measurement is repeated after every change, until the criterion is met or until we can tell you, with the record in front of us, that it will not be met with that model on that hardware.

Where you come in

Three ways through it. In the third one you do not walk it.

What changes between the second and the third is not convenience: it is who chooses the variables. If you choose the model, the quantisation and the machine, we measure what is there and tell you what comes out. If we choose, we answer for the result, and that can only be promised when you control what you are promising.

Building an environment like this turnkey does not set us apart: others do it and publish their phases and their prices. What is different is what travels with the handover: a score on your own case files, produced by an instrument whose floor is measured and whose defects are published.

And one honest point about sizing: choosing the machine is an estimate until it is measured. That is why the order is what it is, the tests before the purchase, and why the handover is conditional on the measurement rather than on the delivery van.

After handover, you operate it or we do. That is a separate decision, taken when it comes up.

The method

A number with a date, a version and a procedure.

The question "is it any good?" has no answer, because it is missing the term of comparison. The one that does have an answer is: does this model, in here, resolve this work as well as the one you would use if you did not have it? These are the six rules that turn that question into a figure.

The uncomfortable part

An instrument nobody has tried to break is not an instrument: it is an opinion with decimal places.

It is the reasonable objection, and it is worth saying out loud: a test designed by whoever is going to sell you the fix can give whatever result suits them. It is not answered with arguments. It is answered by showing what we have done to the instrument.

That is the whole discipline: separating at every moment what measures the model from what measures the instrument, and accepting no fix that has not been checked in both directions.

The commitment

Three things get committed. One of them does not.

If you are choosing who to build this with, you are going to be told that the result, the schedule and the cost are all guaranteed. It is worth separating them, because only two of the three can be sustained and the third is the most reliable signal you have for ruling suppliers out.

And one consequence worth stating in full: if the measurement contradicts the decision, the report will say so. With the record beside it, and in the first deliverable, not the third.

This is the iteration, with figures. Between the two measurements the node's configuration changed and so did the access layer's timeouts. Not the model, not the hardware, not the tests.

What happens between two measurements · same node, same model, same tests
Measure11 September 202612 September 2026
Case files measured1 de 33 de 3
Duration of the measurementmás de tres horas50 minutos
Median latency573 s172 s
Internal reasoning (median)4.657 tokens de 6.212638 tokens de 2.701
State of the runincompletacompleta y estable

The case file that could not be measured became measurable, and that is where the 148.00 € error appeared. That is the real order of things: first the node has to be able to answer, and only then can you talk about whether it is right. Confusing the two questions is what makes a team spend months tuning a configuration to fix a problem that was in the model, or the other way round.

With one warning that goes in the report and not in the small print: trimming a model's internal reasoning lowers its latency and may also lower its score. If it does, the decision stops being technical and becomes yours: a model that needs nine minutes for a 3,500-token case file may have no place in your process, and that is a legitimate result of the measurement, not a fault to fix.

The deliverable

A verdict, the evidence beneath it, and the work split by who executes it.

The report opens with a sentence that closes the question. Ours on our own node, dated 11 September 2026, opened with this one: "it fails on time, not on size". Beneath it, the arithmetic holding it up: the input of all three case files fitted thirty-five times inside the declared limit, so size was not involved; the node writes at 10.88 tokens per second and the access layer waited 600 seconds at most, so no answer longer than 6,528 tokens could ever complete. One case file made it with 29 seconds to spare; the other two did not. The node’s speed is not a defect (it is a physical fact) and the 600-second cap is not one either (it is a reasonable decision for an access layer). The defect was that nobody had compared the two numbers.

And then the thing that makes the report usable instead of arguable: the split. Thirteen findings, six on the node and seven on the access layer, each with its evidence, its estimated effort and its success criterion. Three actions came out of the node; the rest was context explaining the numbers, not outstanding work. A verdict that assigns blame does not get executed; one that assigns work does.

The report also includes the fixes you should not make, which tend to be the first three anyone thinks of. Raising the declared input limit breaks a protection that works: the node does not error when you exceed it, it silently truncates and answers anyway, and the result is a high score built on mutilated documents (credible, presentable and false). Trimming the output budget until the test passes turns it green without fixing anything and breaks comparability with everything else. And enlarging the server's context reduces write speed, which was the problem in the first place.

  • Verdict · one sentence, and the quantity that holds it up
  • Evidence · the record of every run, with the model behind it, the declared window, the version of the cases and the machine
  • Node actions · the ones your team executes, with effort and success criterion
  • Access-layer actions · the ones we execute, in the same detail
  • Rejected fixes · what looks like the solution and is not, with the reason

Who it is not for

Four cases where we will tell you on the first call.

The limits

Saying this is part of the method.

  • It does not measure the model's knowledge of real regulation. It measures its reading of a case file. Every rate, band and cap travels with the prompt and is fictitious on purpose, so the test stays comparable across environments and across dates.
  • It does not execute the code it generates. A requirement can be met in the file and the page still behave badly in a browser. The file is read, not least because executing model-generated code inside your perimeter is precisely what you do not want to do.
  • It does not yet say at what context size a model stops getting it right. The cases are a fixed size. The quality-against-context curve is separate work and it is outstanding.
  • It does not rank two models that already resolve the cases in full. It certifies that the reference level is reached, not by how much. For a perimeter decision that is enough; for a ranking it is not.
  • It does not replace validation on your real workload. The published cases are simulated case files: they resemble the work, they are not the work. A number from them steers a decision; cases written on your own case files close it.
  • We are not ISO 27001 certified. If your tender or your audit requires it of the supplier, say so on the first call and we will save you the cycle.

Who you work with

Whoever signs the verdict is whoever has broken the instrument four times.

Kairos Tek is two partners. One has spent more than twenty-five years leading technology and product, and has advised Europol, the European Commission and the Bank of Spain. The other has spent twenty-five writing code and has been CTO at more than eight companies, building and operating production systems. There is no management layer between them and you, because there is nobody else.

That matters here for one specific reason and not the usual one: the verdict you receive splits work between your team and ours, and whoever writes it has to have operated both halves. A report that only understands one of them ends up assigning to the other everything it cannot explain.

Meet the team →

Start by knowing what score it gets today.

And if you have not built it yet: what score are you going to require of it? An hour is enough to know whether there is a case. We bring a concrete agenda: what case files you process, what figure has to come out of them, and what score you would consider good enough. If at the end of that hour the answer is that there is no case, we say so there.

The method document carries the six tests one by one, the traps each case hunts with the figure each one produces, the measured floor, the negative control and the instrument's four defects with their dates. It downloads without leaving any data.