AI compliance · advisory
We measure how well your model follows your rules — on which cases, with which numbers — and deliver the documentation that shows it. We start with one bounded measurement.
Bounded · fixed price · no lock-in
In an inspection it is the numbers that count: how well the AI follows your rules, and on which cases. A measurement makes that visible.
This matters for two reasons. First, you want the rules you set for the system to actually be followed — that was the point of setting them. Second, you need to be able to document it: the EU AI Act requires you to declare your model's accuracy and robustness (Art. 15) and to show continuous risk management (Art. 9).
And systems often follow their own rules worse than people assume. Including the ones built to follow them. You do not see it until it is measured.
Our job is to make it measurable: how well your AI follows your rules, and the paperwork that proves it.
Do not do this
The more rules compete for the model's attention at once, the fewer are actually followed. "Covering yourself" by pasting in the entire handbook is the very thing that makes compliance fall. We fix it with two moves: show only the rules relevant to the question, and choose the right model size.
We measured it on our own service søgfonde.dk: with 56 rules in play at once, average adherence fell to 33 %, and 32 of the 52 relevant rules failed — several all the way down to 0 %. When the system instead looks up only the rule relevant to the question, adherence rises towards 100 %.
The EU AI Act is in force and the requirements for high-risk uses are phasing in. The question is not whether you will need to show the numbers, but when an authority asks for them. That is what we deliver.
From the article “large or small model?” (in Danish)
Every recommendation comes with a visual explanation you can use and verify. One example: a model that is too small cannot solve the task safely, and a larger one quickly becomes unnecessary — it cannot run on your own servers, and it costs more than the value it adds. The right choice is the smallest model that handles your task with good margin.
See all the articles → (in Danish)
Concrete deliverables. You know what you are paying for the whole way, and you can prove it afterwards.
We find out how advanced a model your task actually requires, so you neither overpay nor pick too little.
We tune the model to your risk profile, so it is neither too loose nor refuses too much.
The second net: if the model is unsure about something the consequence gate let through, it refers on instead of inventing an answer. Built in, not an add-on.
Documentation showing your auditor that you meet the AI Act's requirements on accuracy (Art. 15) and risk management (Art. 9).
The smallest model that solves the task safely. Cheaper, can run on your own servers, and uses less energy.
The method works on the model you prefer, and we install the risk controls in your own setup. The safety net is not universal: change the base model and we recalibrate — a new model does not inherit the old one's protection.
| Declared metric | Danish | English |
|---|---|---|
| Abstains where the source does not cover it (does not invent an answer) | 100 % | 100 % |
| Correct refusal on near-miss cases | 100 % | 93 % |
| Wrongful refusal of legitimate questions | 42 % | 0 % |
| Coverage on clear look-ups | 93 % | 100 % |
Our own demo instance: a gated handbook assistant on Qwen2.5-7B, measured on 100 cases in Danish and English. The 42 % is not a typo. In Danish it refused four out of ten legitimate questions; in English it refused none — same system, same week. That is why calibration has to happen per language, and it is the kind of thing the pack contains.
The regulation does not require the system never to be wrong. It requires you to know how often, to have taken a position on whether that is acceptable, and to be able to show both. What Articles 9 and 15 actually ask for →
No large contract to get started. First we measure. Then you know what needs doing, and whether you want us to do it.
A bounded review: how complex the task is, where the floor and the ceiling are, where the risk sits. Fixed scope, fixed price.
You get a concrete recommendation: which model, which safety nets, which settings. Do it yourselves, or let us do it. No lock-in.
Calibration, gate architecture and audit pack installed in your model, auditable from day one.
The whole method is open — you can verify it: the full measurement protocol → · AI Act & GDPR, the short version → (in Danish)
About
You can verify everything we measure.
Behind the advisory is Tomas Lund. We research how language models make decisions — when they commit to an answer, and when they should have abstained — and turn that research into measurable, documented AI compliance for companies. The research is publicly published as Behavioral Friction Theory — the method behind everything on this site.
The method is public, so you can verify everything we recommend: the measurement protocol and the sources come with it, and the measurement can be run on your own rule-set. You know exactly what you are buying before you buy it.
A bounded analysis with no lock-in. Afterwards you know exactly where your AI compliance stands, and what it takes.
Tell us briefly what you use AI for — we will come back with a bounded measurement offer.