I measure where there is something to measure. There is not always. A workflow, a
text, a course — there the evidence is usually something else: that the problem can be
described precisely enough that you recognise it yourself, and that the fix turns out
cheaper than what you ordered. Our cases carry no measured effect,
and each one says so.
The method is open. The measurement protocol is public, so you can run it again on your own
material. The value is not in keeping the procedure secret; it is in running it properly and handing
over the evidence.
Build an AI system, though, and the uncertainty can be measured. That is where the
numbers actually are. Nobody can guarantee that a model never errs — but it can be
calibrated: measure how often it gets it wrong, and make sure the decisions that really
matter pass a human regardless of how confident the model sounds. You get the numbers,
including the ones that do not look good. That is also the version that holds up in
front of an auditor.
See the method and the articles → (in Danish)