Model evaluation on your own infrastructure
You have your own AI node, but is it the one you need?
That question has gone unanswered for months while the node has been working. And if you have not built it yet, it is the one to answer before the purchase order, not after. We build the tests on your real work, agree the criterion with you, measure end to end on your infrastructure, and deliver a verdict with the evidence and the actions split by who executes them: the ones in your node and the ones in ours. And if you would rather not build it yourself, we choose it, build it and hand it over with the measurement already done on your own cases.
The problem
"It seems to work fine" is not a state.
When an answer comes out wrong there are four possible causes and none of them is visible from outside. Without separating them, every meeting about the node is an exchange of impressions, and the impression that wins is the loudest one.
The model
It is not good enough for that task, and no configuration will fix it.
The quantisation
The model would be good enough, but not at the bits per weight it is loaded with.
The node
A start-up flag, a context size, or a wrongly declared tokeniser.
The access layer
A timeout, a declared limit, or an output budget the node cannot meet.
If you have not built it yet, the problem is the same one step earlier: you are about to choose all four at once, committing money and a deadline, with no number telling you which of them decides the outcome.
And there is a fifth possibility, the worst of all: that it gets it right almost always and fails in one specific place, always the same one, with a figure credible enough that nobody looks at it twice.
The evidence, before the argument
A pass mark and a miscalculated settlement, in the same measurement.
This is our own node: an open-weights MoE model, on our hardware. Measured on 12 September 2026: three case files, five attempts each, everything through the same access layer the real work goes through, and 50 minutes of dedicated node from start to finish. The model, the node and the date are recorded together on purpose: without all three, the score cannot be produced again.
| Case file | Score | Figure returned | Correct figure | Deviation | Identical attempts |
|---|---|---|---|---|---|
| Inheritance · division of an estate | 100,00 | 78.731,75 € | 78.731,75 € | 0,00 € | 5 of 5 |
| Personal tax return | 25,00 | 3.872,79 € | 3.724,79 € | 148,00 € | 5 of 5 |
| Severance for unfair dismissal | 100,00 | 64.506,70 € | 64.506,70 € | 0,00 € | 5 of 5 |
| Aggregate | 75,00 · pass | threshold 70 | mean error 49,33 € | complete and stable run |
148.00 euros on 3,724.79. Four per cent. It is not an absurd figure: it is a credible one, which is why nobody would have caught it by reading the answer. The model returned it identically on all five attempts, so nobody would have caught it by asking again either. And the aggregate score (75.00 against a threshold of 70) would have passed it.
The failure is located, not guessed at. Of the six traceability checks on that case file, five pass: it gets the pension cap right, the net gain on the share sale right and the withholdings right, and it explicitly declares the non-deductible expense outside the calculation. Exactly one fails, the reduced rental income, which is precisely where the case hides its three expense traps. It goes wrong in one place, and always the same one.
A demonstration would have shown the inheritance and the severance. It would have taken ten minutes, it would have gone well, and it would not have been false: it would have been incomplete, which in a decision is worse. A pilot with no acceptance criterion agreed beforehand produces exactly that result, which is why it always goes well.
We can show this measurement in full because the workload is ours. It is no client's case file, so there is nothing to anonymise and no figure to trim.
How it works
Five steps. The first two need no hardware.
That is the part that tends to surprise: the criterion the node will be judged by can be built before the node exists. That is where the three ways in of the next section come from.
The tests, on your work
The cases are written from real case files of yours (the general work and the specific work of your teams) and the correct answer is calculated by hand, before any model sees them.
No hardwareThe criterion, set by you
What score is good enough, and why that one. With the floor measured (what nothing scores) and the ceiling measured against a provider model, so the number means something. And before the purchase order.
No hardwareThe measurement, on your infrastructure
End to end, through the same access layer the real work will go through: on the infrastructure you already have, or on the one we build. Five attempts per case, and the complete record of every run.
The verdict
One sentence that closes the question, the evidence beneath it, and the actions split by scope: the ones your team executes on the node and the ones we execute on the access layer.
The iteration
The measurement is repeated after every change, until the criterion is met or until we can tell you, with the record in front of us, that it will not be met with that model on that hardware.
Where you come in
Three ways through it. In the third one you do not walk it.
You already have your AI node running
You come in at step 3
Steps 1 and 2 still happen, but they are shorter: the case files already exist and the criterion is set against a node that can be measured that same week.
Your team is going to build it
You start at step 1
Steps 1 and 2 close before you sign for the hardware. What they deliver (your cases, your criterion and the measured ceiling) is worth having on its own even if you decide not to continue.
Turnkey
All five steps, and the hardware
We analyse the work, write the tests on your case files, choose and size the infrastructure, build it, and we do not hand it over until the measurement reaches the criterion you set. And if no viable configuration reaches it, there is no handover and you buy no hardware. The day it arrives it has already been tested on your own use cases, so it gets used rather than being the day you start finding out whether it works.
What changes between the second and the third is not convenience: it is who chooses the variables. If you choose the model, the quantisation and the machine, we measure what is there and tell you what comes out. If we choose, we answer for the result, and that can only be promised when you control what you are promising.
Building an environment like this turnkey does not set us apart: others do it and publish their phases and their prices. What is different is what travels with the handover: a score on your own case files, produced by an instrument whose floor is measured and whose defects are published.
And one honest point about sizing: choosing the machine is an estimate until it is measured. That is why the order is what it is, the tests before the purchase, and why the handover is conditional on the measurement rather than on the delivery van.
After handover, you operate it or we do. That is a separate decision, taken when it comes up.
The method
A number with a date, a version and a procedure.
The question "is it any good?" has no answer, because it is missing the term of comparison. The one that does have an answer is: does this model, in here, resolve this work as well as the one you would use if you did not have it? These are the six rules that turn that question into a figure.
The correct answer is calculated by hand, first
Every case carries its step-by-step calculation done by a person, and the expected value comes from that. Never the other way round. Writing the expected value from what a model answered is the quietest way to build a test that passes whoever designed it.
Nothing the model reads contains the answer
An automatic check walks everything the model will see and fails if any value appears that only exists once someone does the calculation. It is there because it failed: a manual carried a worked example and a model returned the example's result, copying the label along with it.
Five attempts, and you are told if they disagree
Five runs per case, and the median is reported. A model that oscillates is no use even when it gets one right: five identical attempts are a datum, one on its own is not.
The score splits: 70 % the figure, 30 % traceability
The figure is binary: inside the tolerance or outside it, with no partial credit for getting close, because a settlement that gets close is no use. The tolerance is justified against the smallest error the case is hunting: in the tax return one of the traps is 26.64 € off, so a one-euro tolerance absorbs rounding and does not absorb misreadings. Traceability does not check that the model cites a document, that is free: it checks that it ties the right amount to the right document, and that it declares what it leaves out. A model can hit the final figure through two errors cancelling out, and this is what stops it.
Everything goes through the same place
Measurements cross the same access layer, with the same authentication and the same logging as the real work. Comparing your model against a provider one is changing one word in a command, not touching the test's code. No measurement taken on the side gets into the comparison.
It closes with a single structured block
And one formatting rule that is a measurement in itself: the model reasons freely and ends with a single structured block. If there is none, or it cannot be read, the case scores zero. A model that cannot close with a structure is no use in an automated process, whatever it gets right inside.
The uncomfortable part
An instrument nobody has tried to break is not an instrument: it is an opinion with decimal places.
It is the reasonable objection, and it is worth saying out loud: a test designed by whoever is going to sell you the fix can give whatever result suits them. It is not answered with arguments. It is answered by showing what we have done to the instrument.
The floor is measured from both sides
A threshold means nothing unless you know what doing nothing scores, so it is measured on every run: a submission that transcribes the list of requirements without understanding a single data point scores 11.67 out of 100 on extraction and 73.33 on construction, and if that floor moves, the run fails. And from the other side a 3-billion-parameter local model is measured on purpose, scoring 2.00 and 0.00 out of 100 with the same command and the same tests. If an instrument does not fail whoever ought to fail, its passes are worth nothing.
An incomplete measurement is not a low score: it is not a score
If a case file leaves not one valid measurement, it does not score zero: it does not score, and the whole run is marked incomplete. It is the rule that rules out the most comfortable result of all (a high score drawn from the cases that did come out) and it is the rule that stops us publishing two of our own ten measurements.
The instrument has broken four times, and all four are written down
Every defect carries its date and has become a permanent check. Two turned up on the same day, in opposite directions: a network failure was being counted as the model’s failure, and at the same time a guard was protecting a model that ran away generating text, erasing its worst failure from the record. A guard that protects the model from its own failures is not a guard: it is a bias. And one of the four corrections cost us: fixing the rule that tells "ran out of output budget" from "did not know when to stop" sent a whole run up a great deal and at the same time left it marked incomplete. Net result: a higher score that cannot be published.
That is the whole discipline: separating at every moment what measures the model from what measures the instrument, and accepting no fix that has not been checked in both directions.
The commitment
Three things get committed. One of them does not.
If you are choosing who to build this with, you are going to be told that the result, the schedule and the cost are all guaranteed. It is worth separating them, because only two of the three can be sustained and the third is the most reliable signal you have for ruling suppliers out.
Depends on who chooses
The result
If you choose the model, the quantisation and the machine, nobody can commit to the score: whoever guarantees it is guaranteeing a score they have not measured, or a test soft enough to pass for certain. There, what we commit to is that the score exists, is reproducible and gets said, whether it is reached or not.
Turnkey is different, because we do the choosing: the criterion stops being an expectation and becomes a condition of handover. It is not handed over until the measurement reaches it. And if no viable configuration reaches it, there is no handover and you buy no hardware: you keep the tests, the criterion and the reason, in writing.
Committed
The time
Because the duration of a measurement is a measured quantity, not an estimate. Ours on 12 September (three case files, five attempts each) took 50 minutes of dedicated node, and that figure sits inside the run's own file, next to every attempt's latency. The one the day before, on the same node without the verdict's actions applied, ran past three hours and left one measurable case file out of three.
Bounded
The cost
We do not commit to a price per request: this is not about the cost of a token. What gets bounded is the cost of being wrong. The first two steps are fixed-price, short, and what they deliver stands on its own: if the candidate model does not reach your criterion, you have a written, dated, reproducible reason not to sign. That is cheaper than finding out with the machine already in the rack.
And one consequence worth stating in full: if the measurement contradicts the decision, the report will say so. With the record beside it, and in the first deliverable, not the third.
This is the iteration, with figures. Between the two measurements the node's configuration changed and so did the access layer's timeouts. Not the model, not the hardware, not the tests.
| Measure | 11 September 2026 | 12 September 2026 |
|---|---|---|
| Case files measured | 1 de 3 | 3 de 3 |
| Duration of the measurement | más de tres horas | 50 minutos |
| Median latency | 573 s | 172 s |
| Internal reasoning (median) | 4.657 tokens de 6.212 | 638 tokens de 2.701 |
| State of the run | incompleta | completa y estable |
The case file that could not be measured became measurable, and that is where the 148.00 € error appeared. That is the real order of things: first the node has to be able to answer, and only then can you talk about whether it is right. Confusing the two questions is what makes a team spend months tuning a configuration to fix a problem that was in the model, or the other way round.
With one warning that goes in the report and not in the small print: trimming a model's internal reasoning lowers its latency and may also lower its score. If it does, the decision stops being technical and becomes yours: a model that needs nine minutes for a 3,500-token case file may have no place in your process, and that is a legitimate result of the measurement, not a fault to fix.
The deliverable
A verdict, the evidence beneath it, and the work split by who executes it.
The report opens with a sentence that closes the question. Ours on our own node, dated 11 September 2026, opened with this one: "it fails on time, not on size". Beneath it, the arithmetic holding it up: the input of all three case files fitted thirty-five times inside the declared limit, so size was not involved; the node writes at 10.88 tokens per second and the access layer waited 600 seconds at most, so no answer longer than 6,528 tokens could ever complete. One case file made it with 29 seconds to spare; the other two did not. The node’s speed is not a defect (it is a physical fact) and the 600-second cap is not one either (it is a reasonable decision for an access layer). The defect was that nobody had compared the two numbers.
And then the thing that makes the report usable instead of arguable: the split. Thirteen findings, six on the node and seven on the access layer, each with its evidence, its estimated effort and its success criterion. Three actions came out of the node; the rest was context explaining the numbers, not outstanding work. A verdict that assigns blame does not get executed; one that assigns work does.
The report also includes the fixes you should not make, which tend to be the first three anyone thinks of. Raising the declared input limit breaks a protection that works: the node does not error when you exceed it, it silently truncates and answers anyway, and the result is a high score built on mutilated documents (credible, presentable and false). Trimming the output budget until the test passes turns it green without fixing anything and breaks comparability with everything else. And enlarging the server's context reduces write speed, which was the problem in the first place.
- Verdict · one sentence, and the quantity that holds it up
- Evidence · the record of every run, with the model behind it, the declared window, the version of the cases and the machine
- Node actions · the ones your team executes, with effort and success criterion
- Access-layer actions · the ones we execute, in the same detail
- Rejected fixes · what looks like the solution and is not, with the reason
Who it is not for
Four cases where we will tell you on the first call.
You are still deciding whether to do it
This page is for those who have already decided: whether you have built it, are about to, or want it handed over built. If the decision is still open, the conversation is a different one and it is on the AI infrastructure page.
What you want is for us to confirm a decision already taken
We do not do that, and it is said above in plain terms. A report that only confirms what you had already decided does not defend the decision: it repeats it.
The work does not admit a checkable answer
This measures well what has a single correct answer: an amount, a classification, a list of requirements met. Whether a summary "reads well" takes judgement, and one judgement cannot be set beside another to compare them. If that is your case, we say so before starting.
You want to do it yourself
You can, and the method is published in full so that you can. Running the test is the cheap part. The expensive part is everything that comes first: calculating the correct answer by hand before the model sees it, and justifying the tolerance against the smallest error you are hunting. Without that, whatever score comes out means nothing.
The limits
Saying this is part of the method.
- It does not measure the model's knowledge of real regulation. It measures its reading of a case file. Every rate, band and cap travels with the prompt and is fictitious on purpose, so the test stays comparable across environments and across dates.
- It does not execute the code it generates. A requirement can be met in the file and the page still behave badly in a browser. The file is read, not least because executing model-generated code inside your perimeter is precisely what you do not want to do.
- It does not yet say at what context size a model stops getting it right. The cases are a fixed size. The quality-against-context curve is separate work and it is outstanding.
- It does not rank two models that already resolve the cases in full. It certifies that the reference level is reached, not by how much. For a perimeter decision that is enough; for a ranking it is not.
- It does not replace validation on your real workload. The published cases are simulated case files: they resemble the work, they are not the work. A number from them steers a decision; cases written on your own case files close it.
- We are not ISO 27001 certified. If your tender or your audit requires it of the supplier, say so on the first call and we will save you the cycle.
Who you work with
Whoever signs the verdict is whoever has broken the instrument four times.
Kairos Tek is two partners. One has spent more than twenty-five years leading technology and product, and has advised Europol, the European Commission and the Bank of Spain. The other has spent twenty-five writing code and has been CTO at more than eight companies, building and operating production systems. There is no management layer between them and you, because there is nobody else.
That matters here for one specific reason and not the usual one: the verdict you receive splits work between your team and ours, and whoever writes it has to have operated both halves. A report that only understands one of them ends up assigning to the other everything it cannot explain.
Start by knowing what score it gets today.
And if you have not built it yet: what score are you going to require of it? An hour is enough to know whether there is a case. We bring a concrete agenda: what case files you process, what figure has to come out of them, and what score you would consider good enough. If at the end of that hour the answer is that there is no case, we say so there.
The method document carries the six tests one by one, the traps each case hunts with the figure each one produces, the measured floor, the negative control and the instrument's four defects with their dates. It downloads without leaving any data.