A model that predicts perfectly does not price insurance more fairly; it stops insuring anyone at all. A fairness verdict published without its detection limit invites the reader to assume the limit is zero. Two different failures, one shared omission, and a borrowed pair of words for telling them apart.
The actuary’s objection usually arrives about twenty minutes into a first meeting, once the pleasantries are done and someone feels safe enough to say what they actually think.
Fairness, he says, is a soft word. Risk pricing is not a moral failure. The whole point of insurance is to tell good risks from bad ones, and if we could not do that there would be no product to sell. Protected attributes are off the table, fine, nobody is arguing. But past that line there is a legitimate right to charge more where more is expected to be lost. So sharpen the model. A sharper model gives a fairer price, not a less fair one.
The other objection comes from the people who buy audits, and it gets asked far less often than it should. Your report says the model is fair. What was that verdict made of? Did you see inside the thing, or did you send it a few thousand prompts and read what came back?
One treats accuracy as a virtue. The other treats it as a claim. Put the two next to each other and they stop being two arguments.
Who this is for
We state the audience before the argument, on principle, because the reader changes how a finding should be written, and pretending otherwise is how reports get fudged.
This one is for people who govern AI systems and people who review them from outside: risk and compliance, internal audit, model risk, external assessors, and whoever on the board has to put their name on something. Engineers may want to skip to parts three and four, where the method is. Part one needs no technical background at all and is written for anyone who has ever been quoted a premium and wondered how they arrived at the number.
We are not neutral here. We sell assessments, so we have an interest in the standard this article proposes. It also costs us something, which we come to in part six.
Executive summary
Two arguments that turn out to be one.
The pooling argument. A risk model that predicts perfectly does not price insurance fairly. It abolishes insurance. Risk transfer needs residual uncertainty to work, and as granularity climbs the pool breaks up into a set of individual prepayments. Cover then becomes unavailable to exactly the people the model has identified as needing it. Some coarseness is the mechanism, not a flaw in it. What makes the argument hard to dismiss is that it runs on actuarial logic, not ethics.
The seeing argument. A fairness or explainability finding is worth whatever the access, the method and the freshness behind it are worth. We grade all three: an access tier, an evidence grade and a validity grade, written A2 · E2 · V2.
And a correction. The orbit article set out three dials, body, pathway and audience, that locate an assessment. The obvious next move is to bolt assurance on as a fourth dial. We drafted it that way, and it was wrong. A dial says where you aimed. It says nothing about how well you could see once you got there, and those are different disclosures. Observational astronomy separated them centuries ago and has words for both: the pointing and the seeing. We are borrowing them.
What links the two halves is the title. Precision is not proof. In insurance, a perfectly precise model is not a fair one; pushed to the limit it is not insurance at all. In assessment, a perfectly confident verdict is not a proven one. Both mistake a precise number for the thing it only gestures at: the insurer reads a sharper model as a fairer price, the board reads “no disparity detected” as “there is no disparity,” and each time what would have qualified the number is left off the page. For a verdict, the two things left off are the ones this article is named for: the pointing, what we looked at, and the seeing, how well we could see it. An assessment is precise only when both are on the record.
Part one · Why a perfect model stops insuring anyone
Insurance is a bet on ignorance, and the ignorance has to be shared.
A group pays into a pot. Nobody knows in advance who will put the car through a hedge, who will find a lump, who will come downstairs to a foot of river water in the hallway. The pot pays whoever it happens to. Premiums reflect the group’s expected loss, so the people who stay lucky end up paying for the people who do not, and everyone buys the same thing: release from a variance they could not have absorbed on their own. That release is the product. Not the payout, which most members never see.
Now improve the model.
Every improvement splits the pool into finer compartments, and the early splits are obviously good. Separating a haulage professional from someone with three speeding convictions is the exact discrimination the product is built on. Refusing to do it would just be a subsidy running from the careful to the careless, with nothing to say for itself. Nobody sensible objects.
Keep going, though. Telematics resolves the driver instead of the demographic. Wearables resolve the body instead of the age band. Aerial imagery resolves your actual roof instead of your postcode. Each step is defensible on its own terms, and here is the awkward part: nobody has to decide to take it. An insurer that declines to segment finds its good risks quoted away by a competitor who does not decline, is left holding the residue, prices up to cover it, loses the next tier of good risks to the same competitor, and discovers a year later that the choice it thought it was making had been made for it somewhere around the second quarter. Competition does the walking. It is a ratchet.
Follow it to the end. Each person is quoted their own predicted loss plus a loading. No variance moves between anybody, because in the insurer’s view there is no longer any variance to move. What you have then is a prepayment plan with administrative overhead bolted on, and calling it insurance is a courtesy.
The damage is not spread evenly either. It lands on the person the model has flagged as high risk, who is of course the person the product existed for. She gets quoted a premium equal to the loss she is expected to suffer. She cannot pay it. That inability is the entire reason she wanted insurance. Meanwhile the man the model likes is offered something cheap, which he could most easily have gone without. The product becomes least available where it is most needed, at the exact moment it becomes most accurate.
You can watch this happening now, in daylight, without any modelling breakthrough at all. Flood and wildfire cover in exposed parts of California, Florida and increasingly southern Europe is priced past what residents can pay, or has simply been withdrawn. The models involved are nowhere near perfect. Better prediction still produced unavailability rather than fairer premiums.
None of this is a new observation about markets. Rothschild and Stiglitz built the formal apparatus for asymmetric information in insurance fifty years ago (1976). The modern case just runs their asymmetry backwards: the insurer holds the informational advantage now, and the pool erodes from the underwriting side rather than from applicants selecting themselves in (Barry and Charpentier, 2020).
Two things this argument is not, because it gets misread in both directions.
It is not a plea for ignorance. Granularity has an optimum, not a maximum, and where the optimum sits depends on the line of business, on how essential the cover is, and on what the law says. Compulsory motor liability, catastrophe cover on a primary residence, and pet insurance do not get the same answer. What the argument rules out is treating maximum resolution as the goal and the loss ratio as the only witness. Granularity becomes a design decision that has to be justified, like every other design decision.
It also is not the whole fairness argument, and most of the rest is settled law rather than open philosophy. Gender-based pricing was closed off in the EU by Test-Achats (C-236/09, 2011). Swiss compulsory basic health insurance may not risk-rate on health status at all, and nobody in Switzerland walks around calling that unfair, which settles the question of whether “actuarially fair” and “fair” are the same word. The genuinely live technical problem is proxy discrimination: does a variable earn its predictive power through some causal link to the insured risk, or by quietly rebuilding an attribute the insurer is not allowed to use (Prince and Schwarcz, 2020)? Price optimisation on propensity to shop is the clean case. It is not risk-based at all, so it fails by the actuary’s own standard, not ours.
“Some coarseness is the mechanism, not a defect in it. Perfect prediction does not give you a fair premium. It gives you a savings account with extra steps.”
Part one · The pooling argument
That is the first objection dealt with, and before leaving the insurer it is worth taking what the argument actually proved, because the two dimensions the rest of this article is built on both come straight out of it.
The insurance story had two edges. The first was about scope: how finely you carve the pool decides whether a premium is fair, so the same accurate model is fair for a haulier and ruinous for the person it prices out. Change what the number is about, and its meaning flips. The second was about measurement: a maximally accurate model is not a maximally fair one, and precision pushed to the limit abolishes the product. Accuracy was never the goal; it was mistaken for it. Precision was not proof.
Those two edges are the two things any number needs before you can lean on it: what it is about, and how it was arrived at. For a premium you would demand both without thinking. For a verdict about an AI system, higher in stakes and far lower in legibility, the industry ships the number alone. The rest of this article is those two disclosures, named and graded: the pointing, what the verdict is about, and the seeing, how well it could be seen. The second objection is the harder one, because from here on it is about our work rather than the insurer’s.
Part two · The pointing: what an assessment is about
Someone hands you a document. An AI system has been assessed, it says, and found trustworthy.
Before asking whether that is true, ask what it is a claim about. Most trustworthiness claims in circulation cannot survive the question, and usually not because the underlying work was bad. The work was fine. It was never located.
Three questions locate it. We call them dials because they turn independently: move one and the other two stay put. A real assessment is a setting on each, often a small range rather than a single notch.
| Dial | The question it answers | Its settings |
|---|---|---|
| Body | Which body, which offering? | Digital Trust · AI Fairness and Explainability · Self-Sovereign Identity (planned) |
| Pathway | What kind of AI system? | Predictive · Generative / LLM · Agentic (planned) · Multi-agent (planned) |
| Audience | Who is the report for? | Those who build it · govern it · are subject to it · review it externally |
A row in a table undersells each of them, because each one changes what the word “fair” is even asking.
Body
Any deployment divides into three actors with nothing left over. The model that decides. The person the decision lands on. The organisation that put the system into the world and has to answer for it. Trust is not any one of them; it is the shape they trace when they are in balance, which is the argument of the orbit piece and the reason no single offering can hand you trust on its own.
For an assessment the consequence is blunt. A clean bill of health on the model tells you nothing about governance, and an immaculate governance programme tells you nothing about whether the model discriminates. These are not stronger and weaker versions of one measurement. A large bank can run a textbook model-risk framework over a model that fails equalised odds by a wide margin. Four people in a room can run a demonstrably fair model with no oversight structure whatsoever. Both happen. Both get reported as trustworthy AI.
Pathway
This dial carries the most weight and attracts the least attention. What changes with the kind of system is the nature of the trust problem, not its difficulty.
Predictive models output a score or a class, and the fairness questions are well posed. Does the score mean the same thing across groups? Are the error rates comparable? Is selection proportionate? You cannot have all of them at once when base rates differ, which is a theorem rather than a bug (Chouldechova, 2017; Kleinberg, Mullainathan and Raghavan, 2017). So the assessor’s job is to find out which trade-off was taken, whether anyone noticed they were taking it, and whether they wrote down why.
Generative systems produce text, and often there is no label to be right or wrong against. Fairness has to be constructed rather than measured. You vary a name, a pronoun, a dialect, and compare what comes back. You frame the open-ended task as a scoreable decision so there is something to count. Failure modes appear that have no predictive analogue, fabrication and provenance chief among them (Ji et al., 2023). And the subject squirms: temperature introduces variance, and the endpoint can be swapped underneath a stable name without anyone telling you.
Agentic systems act. Now the question is not what came out but what got done, on whose authority, and who answers when a chain of tool calls goes wrong on somebody’s behalf. Disparity can hide in tool selection, in retrieval, in how work gets delegated, or in nothing visible at any single step while being obvious across the whole trajectory.
Multi-agent systems add failures that belong to no individual agent. Amplification. Convergence. Drift across turns. Audit each agent, pass each agent, and miss the phenomenon entirely.
Which is why benchmarks travel so badly. A fairness score tuned on a predictive classifier says close to nothing about a multi-agent workflow, and quoting it as though it did is one of the most common pieces of trust theatre in the field.
Audience
The subtle one, and the most open to abuse, so it gets a hard rule. The audience changes the wording, never the numbers.
A board wants the decision and the exposure. An engineer wants the metric, the interval, and enough detail to reproduce it. A person on the receiving end of a decision wants to know what happened to her and what she can do about it, in the second person, with no jargon at all. An external reviewer wants the protocol and the raw counts so she can run it again and disagree with us. Same evidence, four documents.
The rule exists because this dial is where a soft finding gets laundered. If the engineering report says a disparity is significant and material, and by the time the same evidence has passed through a summary, a deck, and a paragraph written by someone in communications who was told to keep it positive, the board is reading that the overall picture is broadly reassuring, then nobody along that chain made a communication choice. Somebody misstated a result. Our answer is mechanical rather than cultural: every audience view generates from one shared assessment record, and the numbers are inherited, not retold.
One insurer, three assessments, three true answers
Take a European insurer running a generative model that drafts claim correspondence, sitting on top of a predictive model that scores claims for fraud review.
| Dial settings | Honest verdict |
|---|---|
| Model · Predictive · External reviewer | The fraud score is not calibrated equally across two language groups. The gap is significant and material. |
| Model · Generative / LLM · Those who build it | The drafting model’s tone shifts measurably with claimant surname. No effect on the decision itself. |
| Organisation · Predictive · Those who govern it | Governance is sound: documented trade-off decisions, a working appeals route, quarterly re-testing. |
All three hold at once. Any one of them, reported as “we assessed the AI system,” is a slogan. Publish only the third and you have described a company with excellent oversight of a miscalibrated model, and called it trustworthy.
So an assessment is a location: a point in body, pathway and audience space, growing into a region as scope widens. That is the whole of the first dimension, the pointing.
This is not only a metaphor. It is the actual control on the live validant.ai platform. When you position an assessment, the three dials are three labelled axes; you pick one or several values on each, and a draggable cube shows the region the assessment occupies. This is how a scope is defined, set and visualised in practice, not just how we describe it on the page.
Now here is the thing the three dials cannot do. Every verdict in that table could have come from a pre-registered controlled study on a pinned model version, or from three afternoons of prompting an endpoint whose weights have since been replaced. The dial settings read identically either way.
Part three · The seeing: how well it was seen
That was the first dimension: the pointing, where we looked. It is necessary, and it is not enough, because it says nothing about how well we could see once we got there. That is the second dimension, the seeing, and it is the one the industry leaves out. The tempting fix is to bolt it on as a fourth dial for depth. We drafted it that way, and it was wrong, and the reason is worth a paragraph, because it is the whole design.
A dial locates. Calling assurance a fourth dial would say that a shallow finding is a different sort of finding from a deep one, sitting somewhere else in the same space. It is not somewhere else. It is the same finding, held less firmly. Where we aimed and how well we could see are answers to two different questions, and letting them share a coordinate system is how a narrow, shallow, expired result gets to wear the same clothes as a deep one.
Observational astronomy sorted this out a long time ago, and it is worth looking at how.
An astronomer records two things and never mixes them up. First the pointing: right ascension and declination, where the instrument was aimed. Then the seeing: which telescope, what aperture, which filter, how long the exposure ran, how steady the air was, and the epoch, because the sky moves and a position without a date is not a position. Pointing says where she looked. Seeing says what could possibly have been resolved from there, and it goes into the log as a number, in arcseconds, every night, without anyone treating it as an admission of weakness.
Strictly, seeing refers to atmospheric steadiness on its own. We use it the way working observers do, as shorthand for everything that together set the limit of what a given night could show.
We take both words. They keep the two disclosures apart without borrowing “object,” which in astronomy already means a body, and we have three of those.
| Disclosure | The question it answers | Reported as |
|---|---|---|
| The pointing | Where did we look? | Body · pathway · audience |
| The seeing | How well could we see it? | Access · evidence · validity |
The pointing can be immaculate while the seeing is hopeless. That combination describes a great deal of the published trustworthiness literature.
There is one more convention worth stealing, and it does more work than the rest put together. A non-detection is never published without a limiting magnitude. “We did not see it” is not a result anyone would accept. “We would have detected anything brighter than magnitude 21.3, and saw nothing” is. The first quietly tells the reader that nothing is there. The second states what could have been there unnoticed. Two centuries of precedent for our third rule.
Seeing has three components. They are independent, and a single grade would bury the trade-offs between them: fifty thousand controlled counterfactual pairs at arm’s length beat forty samples with the weights in hand, and no average can say so.
Aperture · the access tier
What the assessor could observe. Mostly not our choice.
| Tier | Name | What the assessor holds | Highest claim it supports |
|---|---|---|---|
| A0 | Attested | Vendor documentation, model card, published evaluations. No probing. | The subject’s own claims, recorded and checked for internal consistency. |
| A1 | Behavioural | Query access. Terminal output only: text, label, decision. | Disparity in observed outcomes, present or absent at a stated sensitivity. |
| A2 | Scored | A1 plus per-output scores: log probabilities, class probabilities, ranked alternatives, confidence. | Calibration and threshold behaviour by group. Ranking and margin disparity. |
| A3 | Internal | Weights, activations, gradients. The ability to intervene on the computation. | Which internal structures carry the disparity. Mechanistic attribution. |
| A4 | Provenance | A3 plus training-corpus lineage, fine-tuning and preference-data history, evaluation history. | Where in the lifecycle the disparity came from. |
Commercial reality mostly sits at A1 and A2. A deployer assessing a third-party frontier model cannot reach further, and contractually never will. That is not a reason to walk away from the engagement, and here the two capabilities separate sharply.
Fairness measurement is an input-output discipline and does fine at A1. Group metrics come out of predictions, labels and group membership; weights never enter it. Closed frontier models are genuinely assessable for fairness. The constraints are practical: query cost at the sample sizes power demands, non-determinism without seed control, safety filters swallowing probes, provider terms that sometimes forbid adversarial testing, and silent endpoint updates that quietly expire the work.
Explainability does not travel as well. At A1 you have perturbation surrogates and self-report, and a model’s stated reasons are evidence about its output rather than about its computation (Turpin et al., 2023). At A2 sensitivity analysis and contrastive attribution open up properly. Mechanism claims start at A3 and not before. An explainability report that will not state its tier is making a claim it may have had no instrument for.
Method · the evidence grade
How strong the inference is, given what actually got run.
| Grade | Name | Standard met |
|---|---|---|
| E0 | Asserted | Rests on a statement by the subject, or a self-report by the model. No independent measurement. |
| E1 | Observational | Measured on found or convenience data. Confounds uncontrolled. Descriptive only. |
| E2 | Controlled | Pre-registered protocol, powered sample, controlled counterfactual perturbation, interval estimates, multiple-testing correction. |
| E3 | Interventional | The mechanism gets manipulated rather than watched: randomised assignment of the varied attribute, ablation, or activation-level intervention with mediation analysis. |
Explainability findings carry a faithfulness modifier as well, since attribution methods are not interchangeable in evidential weight: F0 self-report, F1 perturbation surrogate (Ribeiro et al., 2016; Lundberg and Lee, 2017), F2 gradient or locally exact, F3 causal intervention.
Epoch · the validity grade
When it was true, and whether it still is.
| Grade | Name | Standard met |
|---|---|---|
| V0 | Unpinned | No version identifier, no configuration record, no expiry. |
| V1 | Pinned | Snapshot or version identifier, configuration and timestamp recorded. One point in time, with a stated expiry. |
| V2 | Resampled | V1 plus scheduled re-testing at a declared interval, with change detection between runs. |
| V3 | Continuous | Production monitoring with drift detection and alert thresholds. The claim is live rather than historical. |
V0 is everywhere and should be treated as disqualifying. Hosted endpoints get updated behind stable names, temperature and system prompt shift the subject under test, a safety layer can be rewritten on a Tuesday afternoon with no announcement anywhere, and the cumulative effect is that a finding filed in March about a model reachable at a given address may, by June, describe nothing that still exists at that address. Which is also why continuous assurance is the honest form of the product rather than an upsell. The orbit will not hold still, so the measurement cannot either.
The vocabulary here is borrowed on purpose. Auditors have distinguished limited from reasonable assurance for years (ISAE 3000 Revised), and identity frameworks graded levels of assurance long before anyone was auditing models (NIST SP 800-63-3; eIDAS low, substantial, high). Something assembled from both is harder to wave off as vendor coinage.
Part four · Five rules that make it load-bearing
A grading scheme that only ever flatters the grader is decoration. These five keep it honest, and each one costs us something.
R1 · Weakest link
The headline claim class is capped by the lowest component. min(A, E, V) maps to Indicative, Limited or Reasonable, using those words as auditors use them. One weak component caps the verdict. No averaging, because averaging is how a deep method on a stale snapshot turns into a strong claim.
R2 · Declared ceiling
One sentence per finding, stating what this configuration could not have detected. Not what was not found, but what was undetectable by construction. At A1 that sentence says, roughly, that no claim is made about the internal mechanism producing these outputs.
R3 · Limiting magnitude on every null
Absence of finding is not finding of absence. Every non-finding carries a minimum detectable effect with its power, alpha and sample size:
“This assessment would have detected a disparity of 3.0 percentage points or greater in selection rate at 80 percent power, α = 0.05, across 12,400 matched probes.”
R3 · A null with its limiting magnitude
It turns a marketing claim into a quantified one. It is also the first sentence anyone wanting a clean result will try to cut, which is a good argument for making it structural rather than editorial.
R4 · Separation
The assurance profile is metadata about our work, and never enters the subject’s fairness result. A vendor must not score better on fairness for having granted deeper access, or the measurement is contaminated by its own conditions of observation. Openness gets rewarded in a separate auditability grade, which is where it belongs. Two subjects with identical findings at A1 and A3 are not equally trustworthy. That difference is real. It is just not a difference in fairness.
R5 · Machine-readable
The profile ships as structured metadata on every finding, not prose in a footer, because it has to survive being pasted into somebody else’s compliance system. The metaphor stays out of the schema: pointing and seeing are how we explain this; scope and assurance are how we serialise it. A field name should not need an article to decode it.
{
"finding_id": "f_8c41e0",
"scope": {
"body": "model",
"pathway": "generative_llm",
"audience": "external_reviewer"
},
"assurance": {
"access": "A1",
"evidence": "E2",
"faithfulness": null,
"validity": "V2",
"claim_class": "limited"
},
"subject": {
"identifier": "vendor/model@2026-07-14",
"configuration_hash": "sha256:1f9a",
"assessed_at": "2026-07-22T09:14:00Z",
"expires_at": "2026-10-20T00:00:00Z"
},
"result": "no_disparity_detected",
"detection_sensitivity": {
"metric": "selection_rate_difference",
"minimum_detectable_effect": 0.030,
"power": 0.80,
"alpha": 0.05,
"n_probes": 12400
},
"declared_ceiling": "Behavioural access only. No claim is made about the internal mechanism producing these outputs, and no mechanistic attribution was possible at this tier."
}Part five · How to read one of these
None of it is worth anything if a reader cannot use it in ten seconds. So here is the record above, in the order a sceptical reader should take it.
Start at the end. The result says no disparity detected. Go straight past it to the sensitivity: three percentage points. A two-point disparity could be sitting right there, unseen. Whether two points matters is a business question rather than a statistical one, and it is now yours to answer instead of ours to blur.
Check the epoch. V2, pinned to a version from 14 July, expiring 20 October. If you are reading this in November, it is history and should be treated as such.
Check the aperture. A1, behavioural. So this is a finding about what the system does, not why it does it. Anyone quoting it as proof that the model contains no discriminatory mechanism has gone past the instrument.
Read the claim class. Limited rather than reasonable, because the weakest component was A1. That word is doing precise work borrowed from audit practice: the finding supports a conclusion without being conclusive.
Now read the pointing. Model, generative, external reviewer. Which tells you it says nothing about the deployer’s governance and nothing about the predictive model underneath it.
Five checks, and what comes out the other side is a claim you can lean on to a stated degree, about a stated thing, until a stated date.
What the five checks destroy is the sentence everybody actually wants, which is that the AI is fair. That sentence is not available at any assurance level. Anyone offering it is selling the omission.
Part six · From a reading to a record
A measurement is a moment. You pointed somewhere, you saw as well as your access, evidence and validity allowed, and you wrote down a grade. But a moment is easy to walk back. The quiet failure of most AI assurance is not that the first reading is wrong. It is that the target moves afterwards, and nobody notices. The metric that looked bad in March is quietly swapped for a kinder one in April. The scope narrows. The population that showed the gap is no longer in scope. The number improves and nothing real did. This is how precision becomes theatre: not by lying about the reading, but by moving what the reading was of.
So the reading has to become a record, and three things have to hold.
The objective must hold still. The moment you take a first honest reading, two of the pointing coordinates stop being editable: the body you are assessing and the pathway of harm you are watching for. Together they are the objective, and an objective you can rewrite is not an objective, it is a mood. From that moment the objective is frozen. You can still assess something else, but that is a new objective with its own record, not a quiet edit to this one. The third coordinate, audience, stays free, because audience never changes the numbers. The same frozen result is simply narrated to a builder, a governor, or the person it affects, in the register each needs. Who you are talking to is a lens, not a lever.
Assurance is earned, not declared. With the objective pinned, the seeing grades become the honest place for progress to happen. A first iteration might be Indicative: you had data but not the model, a single sample, no live monitoring. That is not a failure to hide, it is a starting line to record. The next iteration earns more: you gain the model, so access rises; you run proper confidence intervals with correction, so evidence rises; you turn on continuous monitoring, so validity rises. Each iteration seals its own grade and cannot un-earn it later. What the reader watches, across a stack of sealed iterations, is not a number bouncing around but a floor rising under a fixed goal. That is the difference between scoring well once and getting genuinely better at seeing the same thing. Only the second one is trust.
The headline for each iteration stays the strict rule from Part four: the assurance is the minimum of access, evidence and validity, never the average. A brilliant statistical test on data you were only allowed to glance at is still a glance. You cannot buy back access with cleverness, and the seal refuses to let you pretend otherwise.
| Coordinate | Rule | Why |
|---|---|---|
| Body and pathway, the objective | One value each, frozen when the first iteration closes | An objective you can rewrite is not an objective. If the target can move, improvement is meaningless. |
| Audience | All of them, always | Audience is a lens, not a lever. The same frozen numbers are narrated to a builder, a governor or the affected person. Who you tell never changes what is true. |
| Seeing: access, evidence, validity | Graded per iteration, sealed, can rise but never un-earn | A first pass may be Indicative. Gaining the model, real confidence intervals and live monitoring each raise a grade, so a floor rises under a fixed goal. |
From one seal to a programme
One seal covers one body. But a real deployment is never one body. When a company puts an AI system into the world, three bodies are in play at once: the model that makes the decision, the person the decision lands on, and the organisation that built and governs it. Trust in the system is not any one of them. It is the shape the three trace together. We have written before that digital trust is an orbit, not a pillar. This is where that stops being a metaphor.
A programme is the container for one deployment’s three assessments. Each body keeps its own frozen objective and its own rising stack of sealed iterations. The programme draws them together into a single, publicly verifiable object: a programme seal. And the programme seal inherits the same honesty rule, one level up. Its assurance is not the flattering average of its three bodies. It is the minimum. If the model earns Reasonable assurance but the organisation’s governance is still only Limited, the whole programme is Limited, and the seal says so out loud. You cannot launder a weak body inside a strong average. The orbit is only as trustworthy as its least-seen body, which is exactly the truth a decision-maker needs and almost never gets.
Why this is the whole point
Point honestly, see honestly, and then keep it honestly. Each of those, on its own, is a good habit. Together, and made into a signed, tamper-evident seal that anyone can check without asking us, they change what an assurance claim is. The pointing coordinates and the assurance grades travel inside the seal as structured, always-disclosed fields, kept separate from the fairness verdict itself, so a stranger can check how well it was seen without ever touching protected-group detail. It stops being a snapshot you take on faith and becomes a record you can audit: here is exactly where we aimed, here is exactly how well we could see, here is proof it has not moved, and here is the floor rising across every iteration. That is what it means to increase trust in a digital world. Not a louder claim, but a claim that holds still long enough, and shows its own limits clearly enough, that a stranger can rely on it. The seal is not a badge. It is the receipt.
Part seven · What we would like, and might not get
The tempting story goes like this. Assurance grading becomes procurement language. Buyers write minimum A2, evidence grade E2, sensitivity below three points, validity window under ninety days into their contracts. Providers who want to clear an A2 bar start exposing scores. The market pulls itself toward transparency without anyone having to appeal to conscience. It is a good story, and we would like it to be true.
We do not know that it will be, and there are decent reasons it might not.
Disclosing a detection limit is a competitive disadvantage for whoever goes first, because in a boardroom an unqualified clean bill reads better than a calibrated one. Plenty of buyers want a green light rather than a graded one, and a scheme that abolishes green lights is not self-evidently what the market is asking for. Regulators tend to specify process rather than statistical power, so an attestation regime can be satisfied end to end without anyone ever stating a minimum detectable effect. And the top two access tiers depend entirely on provider willingness, which for frontier models is trending closed rather than open. A scheme whose upper reaches are unreachable in practice might end up documenting nothing but its own ceiling, permanently.
What is actually true today is narrower, and worth stating precisely. Colorado’s Regulation 10-1-1 requires quantitative testing, documentation and annual attestation from life insurers using external consumer data, with auto and health brought into the statute’s scope and their sector rules still in rulemaking. The NAIC model bulletin expects a documented AI-system programme. The EU AI Act puts risk assessment and pricing in life and health insurance in its high-risk annex. That is a real and accelerating move toward documented, testable claims. None of those regimes requires anyone to say how small a disparity their testing could have found. The direction is right. The specific thing this article argues for is on nobody’s statute book, and we should not imply otherwise.
There is precedent for the pattern arriving anyway. Graded assurance became ordinary in identity and in financial audit, in both cases through some mixture of professional convention and eventual codification, and in neither case quickly.
So: a proposal, not a forecast. We are adopting it unilaterally because we think a verdict without a stated ceiling is not really a verdict, and we would rather publish a modest claim that holds than a confident one that does not. It costs us the cleaner headline. We think that is the right trade, and we would be glad to be argued with by anyone who thinks the grades are wrong, the thresholds arbitrary, or the whole apparatus more precise than the underlying science can support. That last objection is the one we take most seriously.
Trust is assessed, not asserted. If your organisation deploys AI and wants its trust posture measured by an independent party, across fairness, explainability and governance, and reported with the ceiling stated rather than hidden, our closed beta opens to a small group in Q3 2026. Write to hello@validant.ai with the subject “Closed Beta,” or request a demo. Seats are limited and assigned in order of fit, not order of arrival.
In one line
Precision is not proof. A model that predicts perfectly is not fairer, it has stopped insuring anyone; a verdict that sounds certain is not proven, it has only hidden how little was seen. A number becomes something you can lean on when two things travel with it: the pointing, what it is about, and the seeing, how well it was seen. Publish the number alone, and you are the insurer again, reading precision as proof.
“Trust is assessed, not asserted. An assessment that will not state its own ceiling has asserted something after all.”
Precision Is Not Proof
Sources and further reading
- 01Glinz, D. (2026). Digital Trust Is an Orbit, Not a Pillar. validant.ai Signal.
- 02Glinz, D. (2026). The Architecture of Digital Trust: A Multi-Level Framework for Bridging the AI Value Gap. 2026 IEEE Swiss Conference on Data Science and AI (SDS), Zurich, pp. 60-67.
- 03Rothschild, M. and Stiglitz, J. (1976). Equilibrium in Competitive Insurance Markets. Quarterly Journal of Economics, 90(4), 629-649.
- 04Barry, L. and Charpentier, A. (2020). Personalization as a Promise: Can Big Data Change the Practice of Insurance? Big Data & Society, 7(1).
- 05Prince, A. E. R. and Schwarcz, D. (2020). Proxy Discrimination in the Age of Artificial Intelligence and Big Data. Iowa Law Review, 105(3).
- 06Barocas, S. and Selbst, A. D. (2016). Big Data’s Disparate Impact. California Law Review, 104, 671-732.
- 07Chouldechova, A. (2017). Fair Prediction with Disparate Impact. Big Data, 5(2), 153-163.
- 08Kleinberg, J., Mullainathan, S. and Raghavan, M. (2017). Inherent Trade-Offs in the Fair Determination of Risk Scores. ITCS 2017.
- 09Lee, J. D. and See, K. A. (2004). Trust in Automation: Designing for Appropriate Reliance. Human Factors, 46(1), 50-80.
- 10Ji, Z., et al. (2023). Survey of Hallucination in Natural Language Generation. ACM Computing Surveys, 55(12).
- 11Ribeiro, M. T., Singh, S. and Guestrin, C. (2016). “Why Should I Trust You?”: Explaining the Predictions of Any Classifier. KDD 2016.
- 12Lundberg, S. M. and Lee, S.-I. (2017). A Unified Approach to Interpreting Model Predictions. NeurIPS 2017.
- 13Turpin, M., Michael, J., Perez, E. and Bowman, S. R. (2023). Language Models Don’t Always Say What They Think. NeurIPS 2023.
- 14International Auditing and Assurance Standards Board. ISAE 3000 (Revised).
- 15National Institute of Standards and Technology. (2017). Digital Identity Guidelines. NIST SP 800-63-3.
- 16Court of Justice of the European Union. (2011). Test-Achats, C-236/09.
- 17Colorado Division of Insurance. SB21-169 and Regulation 10-1-1.
ResearchOpen to readDigital Trust Is an Orbit, Not a Pillar
Trust is not one more pillar to stack. It is the orbit three bodies trace together: the model, the person and the organisation. Why the three-body problem is the honest metaphor for trustworthy AI, and how to tell where you are in the orbit.
Read
ResearchOpen to readThe Trust Problem Nobody Wants to Name
Everybody is spending on AI; almost nobody can say what they got back. A new paper, awarded the Best Paper Award at the IEEE Swiss Conference on Data Science and AI, names that gap and shows it is a trust problem wearing a technical costume: seven mechanisms, a four-layer architecture, an iceberg beneath it, and five design principles for closing it.
Read
ResearchOpen to readWhen a Model Becomes a Munition
A frontier AI model was switched off worldwide by a single government letter. Reading the June 2026 shutdown through the three-body picture of digital trust: how the kill switch stopped being a metaphor, why the most governable lab was governed least carefully of all, and what single-vendor, single-jurisdiction dependence now costs a board.
Read