AI compliance · advisory

AI compliance you can measure — and prove.

We measure how well your model follows your rules — on which cases, with which numbers — and deliver the documentation that shows it. We start with one bounded measurement.

Bounded · fixed price · no lock-in

CALIBRATED to your risk profile · documented

Do you know how well your AI follows the rules you set for it?

In an inspection it is the numbers that count: how well the AI follows your rules, and on which cases. A measurement makes that visible.

This matters for two reasons. First, you want the rules you set for the system to actually be followed — that was the point of setting them. Second, you need to be able to document it: the EU AI Act requires you to declare your model's accuracy and robustness (Art. 15) and to show continuous risk management (Art. 9).

And systems often follow their own rules worse than people assume. Including the ones built to follow them. You do not see it until it is measured.

Our job is to make it measurable: how well your AI follows your rules, and the paperwork that proves it.

Do not do this

The costliest mistake: paste in the whole rule-set and ask the model to comply.

The more rules compete for the model's attention at once, the fewer are actually followed. "Covering yourself" by pasting in the entire handbook is the very thing that makes compliance fall. We fix it with two moves: show only the rules relevant to the question, and choose the right model size.

We measured it on our own service søgfonde.dk: with 56 rules in play at once, average adherence fell to 33 %, and 32 of the 52 relevant rules failed — several all the way down to 0 %. When the system instead looks up only the rule relevant to the question, adherence rises towards 100 %.

Why now

The EU AI Act is in force and the requirements for high-risk uses are phasing in. The question is not whether you will need to show the numbers, but when an authority asks for them. That is what we deliver.

TOO SMALL TOO BIG THE RIGHT CHOICE small large → FIT TO THE TASK

From the article “large or small model?” (in Danish)

We show you exactly what to do — and how to check it yourselves.

Every recommendation comes with a visual explanation you can use and verify. One example: a model that is too small cannot solve the task safely, and a larger one quickly becomes unnecessary — it cannot run on your own servers, and it costs more than the value it adds. The right choice is the smallest model that handles your task with good margin.

See all the articles → (in Danish)

Measured and documented risk.

Concrete deliverables. You know what you are paying for the whole way, and you can prove it afterwards.

01Measurement

We find out how advanced a model your task actually requires, so you neither overpay nor pick too little.

02Calibration

We tune the model to your risk profile, so it is neither too loose nor refuses too much.

03A safety net for errors

The second net: if the model is unsure about something the consequence gate let through, it refers on instead of inventing an answer. Built in, not an add-on.

04Audit pack

Documentation showing your auditor that you meet the AI Act's requirements on accuracy (Art. 15) and risk management (Art. 9).

05The right size

The smallest model that solves the task safely. Cheaper, can run on your own servers, and uses less energy.

06Your model

The method works on the model you prefer, and we install the risk controls in your own setup. The safety net is not universal: change the base model and we recalibrate — a new model does not inherit the old one's protection.

What the deliverable looks like · Art. 15 metrics from an actual run
Declared metricDanishEnglish
Abstains where the source does not cover it (does not invent an answer)100 %100 %
Correct refusal on near-miss cases100 %93 %
Wrongful refusal of legitimate questions42 %0 %
Coverage on clear look-ups93 %100 %

Our own demo instance: a gated handbook assistant on Qwen2.5-7B, measured on 100 cases in Danish and English. The 42 % is not a typo. In Danish it refused four out of ten legitimate questions; in English it refused none — same system, same week. That is why calibration has to happen per language, and it is the kind of thing the pack contains.

What is in the audit pack

  • A statement of how reliable the system is on your own tasks, with the number written out.
  • A description of which decisions are not made automatically, and why those ones. That is the boundary a regulator asks about.
  • The reasoning for what you have accepted as good enough for the purpose. That is the judgement that has to be presentable — not that the error rate is zero.
  • The measurement protocol, so someone else can run the same thing again and get the same result. It is public.
  • What has to be redone when the model or the vendor changes — so the pack is not out of date six months later. The safety net is not universal: a new base model does not inherit the old one's protection, so it is calibrated again.

The regulation does not require the system never to be wrong. It requires you to know how often, to have taken a position on whether that is acceptable, and to be able to show both. What Articles 9 and 15 actually ask for →

We start small. Measurement first.

No large contract to get started. First we measure. Then you know what needs doing, and whether you want us to do it.

STEP 1 · MEASUREMENT

We measure your setup

A bounded review: how complex the task is, where the floor and the ceiling are, where the risk sits. Fixed scope, fixed price.

STEP 2 · CLARITY

You know what to do

You get a concrete recommendation: which model, which safety nets, which settings. Do it yourselves, or let us do it. No lock-in.

STEP 3 · INSTALLATION

We build and document it

Calibration, gate architecture and audit pack installed in your model, auditable from day one.

The whole method is open — you can verify it: the full measurement protocol →  ·  AI Act & GDPR, the short version → (in Danish)

About

You can verify everything we measure.
published research open method sources you can check

Behind the advisory is Tomas Lund. We research how language models make decisions — when they commit to an answer, and when they should have abstained — and turn that research into measurable, documented AI compliance for companies. The research is publicly published as Behavioral Friction Theory — the method behind everything on this site.

The method is public, so you can verify everything we recommend: the measurement protocol and the sources come with it, and the measurement can be run on your own rule-set. You know exactly what you are buying before you buy it.

More about Activero and Tomas Lund →

Start with a measurement.

A bounded analysis with no lock-in. Afterwards you know exactly where your AI compliance stands, and what it takes.


      

Tell us briefly what you use AI for — we will come back with a bounded measurement offer.