All updates
Research

Precision Is Not Proof: The Trap in Every AI Fairness Verdict

Fine grey-blue technical drawing on warm paper of an astronomer at a large telescope inside a single bounded circular aperture ring: a few stars rendered sharp and resolved within the ring, many fainter stars beyond it washed in pale coral, the aperture ring picked out in teal, and a small human figure at the eyepiece for scale.

Precision is not proof, and mistaking one for the other is a single trap wearing two disguises. Push a risk model to perfect precision and it does not price insurance more fairly; it stops insuring anyone at all, because once each person is resolved to their own predicted loss there is no shared uncertainty left to pool. Publish a fairness verdict that looks every bit as precise, “no disparity detected,” and it can still prove nothing, because what it pointed at may be ill-defined and what it managed to see may be shallow. Granularity is not accuracy, and a confident number is not a true one. This is about assessing AI systems for fairness and explainability, and the two disclosures that separate a precise verdict from a proven one, the two an assessment must make and almost never does.

The actuary’s objection usually arrives about twenty minutes into a first meeting, once the pleasantries are done and someone feels safe enough to say what they actually think.

Fairness, he says, is a soft word. Risk pricing is not a moral failure. The whole point of insurance is to tell good risks from bad ones, and if we could not do that there would be no product to sell. Protected attributes are off the table, fine, nobody is arguing. But past that line there is a legitimate right to charge more where more is expected to be lost. So sharpen the model. A sharper model gives a fairer price, not a less fair one.

The other objection comes from the people who buy audits, and it gets asked far less often than it should. Your report says the model is fair. What was that verdict made of? Did you see inside the thing, or did you send it a few thousand prompts and read what came back?

One treats accuracy as a virtue. The other treats it as a claim. Put the two next to each other and they stop being two arguments.

Who this is for

We state the audience before the argument, on principle, because the reader changes how a finding should be written, and pretending otherwise is how reports get fudged.

This one is for people who govern AI systems and people who review them from outside: risk and compliance, internal audit, model risk, external assessors, and whoever on the board has to put their name on something. Engineers may want to skip to parts three and four, where the method is. Part one needs no technical background at all and is written for anyone who has ever been quoted a premium and wondered how they arrived at the number.

We are not neutral here. We sell assessments, so we have an interest in the standard this article proposes. It also costs us something, which we come to in part seven.

Executive summary

Two arguments that turn out to be one.

The pooling argument. A risk model that predicts perfectly does not price insurance fairly. It abolishes insurance. Risk transfer needs residual uncertainty to work, and as granularity climbs the pool breaks up into a set of individual prepayments. Cover then becomes unavailable to exactly the people the model has identified as needing it. Some coarseness is the mechanism, not a flaw in it. What makes the argument hard to dismiss is that it runs on actuarial logic, not ethics.

The seeing argument. A fairness or explainability finding is worth whatever the access, the method and the freshness behind it are worth. We grade all three: an access tier, an evidence grade and a validity grade, written A2 · E2 · V2.

And a correction. The orbit article set out three dials, body, pathway and audience, that locate an assessment. The obvious next move is to bolt assurance on as a fourth dial. We drafted it that way, and it was wrong. A dial says where you aimed. It says nothing about how well you could see once you got there, and those are different disclosures. Observational astronomy separated them centuries ago and has words for both: the pointing and the seeing. We are borrowing them.

What links the two halves is the title. Precision is not proof. In insurance, a perfectly precise model is not a fair one; pushed to the limit it is not insurance at all. In assessment, a perfectly confident verdict is not a proven one. Both mistake a precise number for the thing it only gestures at: the insurer reads a sharper model as a fairer price, the board reads “no disparity detected” as “there is no disparity,” and each time what would have qualified the number is left off the page. For a verdict, the two things left off are the ones this article is named for: the pointing, what we looked at, and the seeing, how well we could see it. An assessment is precise only when both are on the record.

Part one · Why a perfect model stops insuring anyone

Insurance is a bet on ignorance, and the ignorance has to be shared.

A group pays into a pot. Nobody knows in advance who will put the car through a hedge, who will find a lump, who will come downstairs to a foot of river water in the hallway. The pot pays whoever it happens to. Premiums reflect the group’s expected loss, so the people who stay lucky end up paying for the people who do not, and everyone buys the same thing: release from a variance they could not have absorbed on their own. That release is the product. Not the payout, which most members never see.

Now improve the model.

Every improvement splits the pool into finer compartments, and the early splits are obviously good. Separating a haulage professional from someone with three speeding convictions is the exact discrimination the product is built on. Refusing to do it would just be a subsidy running from the careful to the careless, with nothing to say for itself. Nobody sensible objects.

Keep going, though. Telematics resolves the driver instead of the demographic. Wearables resolve the body instead of the age band. Aerial imagery resolves your actual roof instead of your postcode. Each step is defensible on its own terms, and here is the awkward part: nobody has to decide to take it. An insurer that declines to segment finds its good risks quoted away by a competitor who does not decline, is left holding the residue, prices up to cover it, loses the next tier of good risks to the same competitor, and discovers a year later that the choice it thought it was making had been made for it somewhere around the second quarter. Competition does the walking. It is a ratchet.

Follow it to the end. Each person is quoted their own predicted loss plus a loading. No variance moves between anybody, because in the insurer’s view there is no longer any variance to move. What you have then is a prepayment plan with administrative overhead bolted on, and calling it insurance is a courtesy.

The damage is not spread evenly either. It lands on the person the model has flagged as high risk, who is of course the person the product existed for. She gets quoted a premium equal to the loss she is expected to suffer. She cannot pay it. That inability is the entire reason she wanted insurance. Meanwhile the man the model likes is offered something cheap, which he could most easily have gone without. The product becomes least available where it is most needed, at the exact moment it becomes most accurate.

As the model sharpens, the shared pool fragments into individual prepayments, and cover recedes from the person who needed it most.

You can watch this happening now, in daylight, without any modelling breakthrough at all. Flood and wildfire cover in exposed parts of California, Florida and increasingly southern Europe is priced past what residents can pay, or has simply been withdrawn. The models involved are nowhere near perfect. Better prediction still produced unavailability rather than fairer premiums.

None of this is a new observation about markets. Rothschild and Stiglitz built the formal apparatus for asymmetric information in insurance fifty years ago (1976). The modern case just runs their asymmetry backwards: the insurer holds the informational advantage now, and the pool erodes from the underwriting side rather than from applicants selecting themselves in (Barry and Charpentier, 2020).

Two things this argument is not, because it gets misread in both directions.

It is not a plea for ignorance. Granularity has an optimum, not a maximum, and where the optimum sits depends on the line of business, on how essential the cover is, and on what the law says. Compulsory motor liability, catastrophe cover on a primary residence, and pet insurance do not get the same answer. What the argument rules out is treating maximum resolution as the goal and the loss ratio as the only witness. Granularity becomes a design decision that has to be justified, like every other design decision.

It also is not the whole fairness argument, and most of the rest is settled law rather than open philosophy. Gender-based pricing was closed off in the EU by Test-Achats (C-236/09, 2011). Swiss compulsory basic health insurance may not risk-rate on health status at all, and nobody in Switzerland walks around calling that unfair, which settles the question of whether “actuarially fair” and “fair” are the same word. The genuinely live technical problem is proxy discrimination: does a variable earn its predictive power through some causal link to the insured risk, or by quietly rebuilding an attribute the insurer is not allowed to use (Prince and Schwarcz, 2020)? Price optimisation on propensity to shop is the clean case. It is not risk-based at all, so it fails by the actuary’s own standard, not ours.

“Some coarseness is the mechanism, not a flaw in it. Perfect prediction does not give you a fair premium. It gives you a savings account with extra steps.”

Part one · The pooling argument

That is the first objection dealt with, and before leaving the insurer it is worth taking what the argument actually proved, because the two dimensions the rest of this article is built on both come straight out of it.

The insurance story had two edges. The first was about scope: how finely you carve the pool decides whether a premium is fair, so the same accurate model is fair for a haulier and ruinous for the person it prices out. Change what the number is about, and its meaning flips. The second was about measurement: a maximally accurate model is not a maximally fair one, and precision pushed to the limit abolishes the product. Accuracy was never the goal; it was mistaken for it. Precision was not proof.

Those two edges are the two things any number needs before you can lean on it: what it is about, and how it was arrived at. For a premium you would demand both without thinking. For a verdict about an AI system, higher in stakes and far lower in legibility, the industry ships the number alone. The rest of this article is those two disclosures, named and graded: the pointing, what the verdict is about, and the seeing, how well it could be seen. The second objection is the harder one, because from here on it is about our work rather than the insurer’s.

A trade-off, and one that only looks like it

The two edges fail in different ways, and it helps to name the difference before building on it. The insurer’s is a genuine trade-off: accuracy and the pool cannot both be maximised, so fairness sits at an optimum in the middle rather than at the sharp end. This is measured, not asserted. The share of a population’s losses that insurance actually covers is highest under moderate risk classification and falls as prices approach perfect risk-differentiation (Thomas, 2017; Crocker and Snow, 1986), which is the formal reason actuarially fair was never the same word as fair (Landes, 2015).

The assessment edge looks like a trade-off and is not one. A precise verdict and a proven verdict are not two ends of a dial you slide between; you want both, and precision on its own is simply not enough. A number earns reliance only when it also says what it is about and how well it was seen, so publishing it bare is not a balance struck but a disclosure withheld. Two older disciplines are blunt about this. Metrology treats a measurement with no stated uncertainty as no usable result at all, and keeps precision and trueness as separate properties (the GUM and ISO 5725); the study of judgment calls the bare confident number overprecision, the most stubborn way we are certain of things we have not earned (Moore and Healy, 2008); and medical statistics insists that a null is not a finding of nothing until it states how small an effect it could have caught (Altman and Bland, 1995).

Enlarge
Two lanes, two kinds of failure. On the left a real trade-off: fairness peaks at an optimum, and a maximally accurate model is not a maximally fair one. On the right not a trade-off at all: the trap is a precise number with nothing disclosed, and the way out is up, by adding what it is about and how well it was seen, not by trading one good for the other. Hover any element for detail.

Part two · The pointing: what an assessment is about

Someone hands you a document. An AI system has been assessed, it says, and found trustworthy.

Before asking whether that is true, ask what it is a claim about. Most trustworthiness claims in circulation cannot survive the question, and usually not because the underlying work was bad. The work was fine. It was never located.

Three questions locate it. We call them dials because they turn independently: move one and the other two stay put. A real assessment is a setting on each, often a small range rather than a single notch.

DialThe question it answersIts settings
BodyWhich body, which offering?Digital Trust · AI Fairness and Explainability · Self-Sovereign Identity (planned)
PathwayWhat kind of AI system?Predictive · Generative / LLM · Agentic (planned) · Multi-agent (planned)
AudienceWho is the report for?Those who build it · govern it · are subject to it · review it externally
Three dials that turn independently. A real assessment is a setting on each.

A row in a table undersells each of them, because each one changes what the word “fair” is even asking.

Body

Any deployment divides into three actors with nothing left over. The model that decides. The person the decision lands on. The organisation that put the system into the world and has to answer for it. Trust is not any one of them; it is the shape they trace when they are in balance, which is the argument of the orbit piece and the reason no single offering can hand you trust on its own.

Balanced: three bodies hold one shape, and trust is the orbit they trace.The three bodies of an assessment: the model, the person and the organisation, in the balance the orbit article set out.

For an assessment the consequence is blunt. A clean bill of health on the model tells you nothing about governance, and an immaculate governance programme tells you nothing about whether the model discriminates. These are not stronger and weaker versions of one measurement. A large bank can run a textbook model-risk framework over a model that fails equalised odds by a wide margin. Four people in a room can run a demonstrably fair model with no oversight structure whatsoever. Both happen. Both get reported as trustworthy AI.

Pathway

This dial carries the most weight and attracts the least attention. What changes with the kind of system is the nature of the trust problem, not its difficulty.

Predictive models output a score or a class, and the fairness questions are well posed. Does the score mean the same thing across groups? Are the error rates comparable? Is selection proportionate? You cannot have all of them at once when base rates differ, which is a theorem rather than a bug (Chouldechova, 2017; Kleinberg, Mullainathan and Raghavan, 2017). So the assessor’s job is to find out which trade-off was taken, whether anyone noticed they were taking it, and whether they wrote down why.

Generative systems produce text, and often there is no label to be right or wrong against. Fairness has to be constructed rather than measured. You vary a name, a pronoun, a dialect, and compare what comes back. You frame the open-ended task as a scoreable decision so there is something to count. Failure modes appear that have no predictive analogue, fabrication and provenance chief among them (Ji et al., 2023). And the subject squirms: temperature introduces variance, and the endpoint can be swapped underneath a stable name without anyone telling you.

Agentic systems act. Now the question is not what came out but what got done, on whose authority, and who answers when a chain of tool calls goes wrong on somebody’s behalf. Disparity can hide in tool selection, in retrieval, in how work gets delegated, or in nothing visible at any single step while being obvious across the whole trajectory.

Multi-agent systems add failures that belong to no individual agent. Amplification. Convergence. Drift across turns. Audit each agent, pass each agent, and miss the phenomenon entirely.

Which is why benchmarks travel so badly. A fairness score tuned on a predictive classifier says close to nothing about a multi-agent workflow, and quoting it as though it did is one of the most common pieces of trust theatre in the field.

Audience

The subtle one, and the most open to abuse, so it gets a hard rule. The audience changes the wording, never the numbers.

A board wants the decision and the exposure. An engineer wants the metric, the interval, and enough detail to reproduce it. A person on the receiving end of a decision wants to know what happened to her and what she can do about it, in the second person, with no jargon at all. An external reviewer wants the protocol and the raw counts so she can run it again and disagree with us. Same evidence, four documents.

The rule exists because this dial is where a soft finding gets laundered. If the engineering report says a disparity is significant and material, and by the time the same evidence has passed through a summary, a deck, and a paragraph written by someone in communications who was told to keep it positive, the board is reading that the overall picture is broadly reassuring, then nobody along that chain made a communication choice. Somebody misstated a result. Our answer is mechanical rather than cultural: every audience view generates from one shared assessment record, and the numbers are inherited, not retold.

One insurer, three assessments, three true answers

Take a European insurer running a generative model that drafts claim correspondence, sitting on top of a predictive model that scores claims for fraud review.

Dial settingsHonest verdict
Model · Predictive · External reviewerThe fraud score is not calibrated equally across two language groups. The gap is significant and material.
Model · Generative / LLM · Those who build itThe drafting model’s tone shifts measurably with claimant surname. No effect on the decision itself.
Organisation · Predictive · Those who govern itGovernance is sound: documented trade-off decisions, a working appeals route, quarterly re-testing.
Three settings, three true answers about one company. Publish only the third and you have called a miscalibrated model trustworthy.

All three hold at once. Any one of them, reported as “we assessed the AI system,” is a slogan. Publish only the third and you have described a company with excellent oversight of a miscalibrated model, and called it trustworthy.

So an assessment is a location: a point in body, pathway and audience space, growing into a region as scope widens. That is the whole of the first dimension, the pointing.

An assessment is a location in body, pathway and audience space; the mark has volume because scope is a region, not a point.

This is not only a metaphor. It is the actual control on the live validant.ai platform. When you position an assessment, the three dials are three labelled axes; you pick one or several values on each, and a draggable cube shows the region the assessment occupies. This is how a scope is defined, set and visualised in practice, not just how we describe it on the page.

The same three dials, live on the validant.ai platform: positioning an assessment sets its body, pathway and audience, and the cube shows the region it occupies. This is how it is actually done.

Now here is the thing the three dials cannot do. Every verdict in that table could have come from a pre-registered controlled study on a pinned model version, or from three afternoons of prompting an endpoint whose weights have since been replaced. The dial settings read identically either way.

Part three · The seeing: how well it was seen

That was the first dimension: the pointing, where we looked. It is necessary, and it is not enough, because it says nothing about how well we could see once we got there. That is the second dimension, the seeing, and it is the one the industry leaves out. The tempting fix is to bolt it on as a fourth dial for depth. We drafted it that way, and it was wrong, and the reason is worth a paragraph, because it is the whole design.

A dial locates. Calling assurance a fourth dial would say that a shallow finding is a different sort of finding from a deep one, sitting somewhere else in the same space. It is not somewhere else. It is the same finding, held less firmly. Where we aimed and how well we could see are answers to two different questions, and letting them share a coordinate system is how a narrow, shallow, expired result gets to wear the same clothes as a deep one.

Observational astronomy sorted this out a long time ago, and it is worth looking at how.

An astronomer records two things and never mixes them up. First the pointing: right ascension and declination, where the instrument was aimed. Then the seeing: which telescope, what aperture, which filter, how long the exposure ran, how steady the air was, and the epoch, because the sky moves and a position without a date is not a position. Pointing says where she looked. Seeing says what could possibly have been resolved from there, and it goes into the log as a number, in arcseconds, every night, without anyone treating it as an admission of weakness.

Strictly, seeing refers to atmospheric steadiness on its own. We use it the way working observers do, as shorthand for everything that together set the limit of what a given night could show.

We take both words. They keep the two disclosures apart without borrowing “object,” which in astronomy already means a body, and we have three of those.

DisclosureThe question it answersReported as
The pointingWhere did we look?Body · pathway · audience
The seeingHow well could we see it?Access · evidence · validity
Two disclosures, kept apart. The pointing can be immaculate while the seeing is hopeless.

The pointing can be immaculate while the seeing is hopeless. That combination describes a great deal of the published trustworthiness literature.

There is one more convention worth stealing, and it does more work than the rest put together. A non-detection is never published without a limiting magnitude. “We did not see it” is not a result anyone would accept. “We would have detected anything brighter than magnitude 21.3, and saw nothing” is. The first quietly tells the reader that nothing is there. The second states what could have been there unnoticed. Two centuries of precedent for our third rule.

Same sky, same pointing, different seeing. “We saw nothing” is a finding only once you state the limit that makes it one.

Seeing has three components. They are independent, and a single grade would bury the trade-offs between them: fifty thousand controlled counterfactual pairs at arm’s length beat forty samples with the weights in hand, and no average can say so.

Aperture · the access tier

What the assessor could observe. Mostly not our choice.

TierNameWhat the assessor holdsHighest claim it supports
A0AttestedVendor documentation, model card, published evaluations. No probing.The subject’s own claims, recorded and checked for internal consistency.
A1BehaviouralQuery access. Terminal output only: text, label, decision.Disparity in observed outcomes, present or absent at a stated sensitivity.
A2ScoredA1 plus per-output scores: log probabilities, class probabilities, ranked alternatives, confidence.Calibration and threshold behaviour by group. Ranking and margin disparity.
A3InternalWeights, activations, gradients. The ability to intervene on the computation.Which internal structures carry the disparity. Mechanistic attribution.
A4ProvenanceA3 plus training-corpus lineage, fine-tuning and preference-data history, evaluation history.Where in the lifecycle the disparity came from.
The aperture is mostly not the assessor’s choice. Commercial reality sits at A1 and A2.

Commercial reality mostly sits at A1 and A2. A deployer assessing a third-party frontier model cannot reach further, and contractually never will. That is not a reason to walk away from the engagement, and here the two capabilities separate sharply.

Fairness measurement is an input-output discipline and does fine at A1. Group metrics come out of predictions, labels and group membership; weights never enter it. Closed frontier models are genuinely assessable for fairness. The constraints are practical: query cost at the sample sizes power demands, non-determinism without seed control, safety filters swallowing probes, provider terms that sometimes forbid adversarial testing, and silent endpoint updates that quietly expire the work.

Explainability does not travel as well. At A1 you have perturbation surrogates and self-report, and a model’s stated reasons are evidence about its output rather than about its computation (Turpin et al., 2023). At A2 sensitivity analysis and contrastive attribution open up properly. Mechanism claims start at A3 and not before. An explainability report that will not state its tier is making a claim it may have had no instrument for.

Method · the evidence grade

How strong the inference is, given what actually got run.

GradeNameStandard met
E0AssertedRests on a statement by the subject, or a self-report by the model. No independent measurement.
E1ObservationalMeasured on found or convenience data. Confounds uncontrolled. Descriptive only.
E2ControlledPre-registered protocol, powered sample, controlled counterfactual perturbation, interval estimates, multiple-testing correction.
E3InterventionalThe mechanism gets manipulated rather than watched: randomised assignment of the varied attribute, ablation, or activation-level intervention with mediation analysis.
Evidence grade: how strong the inference is, given what was actually run.

Explainability findings carry a faithfulness modifier as well, since attribution methods are not interchangeable in evidential weight: F0 self-report, F1 perturbation surrogate (Ribeiro et al., 2016; Lundberg and Lee, 2017), F2 gradient or locally exact, F3 causal intervention.

Epoch · the validity grade

When it was true, and whether it still is.

GradeNameStandard met
V0UnpinnedNo version identifier, no configuration record, no expiry.
V1PinnedSnapshot or version identifier, configuration and timestamp recorded. One point in time, with a stated expiry.
V2ResampledV1 plus scheduled re-testing at a declared interval, with change detection between runs.
V3ContinuousProduction monitoring with drift detection and alert thresholds. The claim is live rather than historical.
Validity grade: when it was true, and whether it still is. V0 should be treated as disqualifying.

V0 is everywhere and should be treated as disqualifying. Hosted endpoints get updated behind stable names, temperature and system prompt shift the subject under test, a safety layer can be rewritten on a Tuesday afternoon with no announcement anywhere, and the cumulative effect is that a finding filed in March about a model reachable at a given address may, by June, describe nothing that still exists at that address. Which is also why continuous assurance is the honest form of the product rather than an upsell. The orbit will not hold still, so the measurement cannot either.

The vocabulary here is borrowed on purpose. Auditors have distinguished limited from reasonable assurance for years (ISAE 3000 Revised), and identity frameworks graded levels of assurance long before anyone was auditing models (NIST SP 800-63-3; eIDAS low, substantial, high). Something assembled from both is harder to wave off as vendor coinage.

Part four · Five rules that make it load-bearing

A grading scheme that only ever flatters the grader is decoration. These five keep it honest, and each one costs us something.

R1 · Weakest link

The headline claim class is capped by the lowest component. min(A, E, V) maps to Indicative, Limited or Reasonable, using those words as auditors use them. One weak component caps the verdict. No averaging, because averaging is how a deep method on a stale snapshot turns into a strong claim.

One weak component caps the verdict; the beam rests on the shortest column, and no averaging can raise it.

R2 · Declared ceiling

One sentence per finding, stating what this configuration could not have detected. Not what was not found, but what was undetectable by construction. At A1 that sentence says, roughly, that no claim is made about the internal mechanism producing these outputs.

R3 · Limiting magnitude on every null

Absence of finding is not finding of absence. Every non-finding carries a minimum detectable effect with its power, alpha and sample size:

“This assessment would have detected a disparity of 3.0 percentage points or greater in selection rate at 80 percent power, α = 0.05, across 12,400 matched probes.”

R3 · A null with its limiting magnitude

It turns a marketing claim into a quantified one. It is also the first sentence anyone wanting a clean result will try to cut, which is a good argument for making it structural rather than editorial.

R4 · Separation

The assurance profile is metadata about our work, and never enters the subject’s fairness result. A vendor must not score better on fairness for having granted deeper access, or the measurement is contaminated by its own conditions of observation. Openness gets rewarded in a separate auditability grade, which is where it belongs. Two subjects with identical findings at A1 and A3 are not equally trustworthy. That difference is real. It is just not a difference in fairness.

R5 · Machine-readable, and verifiable

The profile ships as structured metadata on every finding, not prose in a footer, because it has to survive being pasted into somebody else’s compliance system. The metaphor stays out of the schema: pointing and seeing are how we explain this; scope and assurance are how we serialise it. A field name should not need an article to decode it.

{
  "finding_id": "f_8c41e0",
  "scope": {
    "body": "model",
    "pathway": "generative_llm",
    "audience": "external_reviewer"
  },
  "assurance": {
    "access": "A1",
    "evidence": "E2",
    "faithfulness": null,
    "validity": "V2",
    "claim_class": "limited"
  },
  "subject": {
    "identifier": "vendor/model@2026-07-14",
    "configuration_hash": "sha256:1f9a",
    "assessed_at": "2026-07-22T09:14:00Z",
    "expires_at": "2026-10-20T00:00:00Z"
  },
  "result": "no_disparity_detected",
  "detection_sensitivity": {
    "metric": "selection_rate_difference",
    "minimum_detectable_effect": 0.030,
    "power": 0.80,
    "alpha": 0.05,
    "n_probes": 12400
  },
  "declared_ceiling": "Behavioural access only. No claim is made about the internal mechanism producing these outputs, and no mechanistic attribution was possible at this tier."
}
One finding as machine-readable metadata: scope, assurance profile, pinned subject, and the detection sensitivity behind a null result.

Machine-readable is necessary, not sufficient. A JSON object anyone can edit in a text editor proves nothing; pasted into a compliance system it is still just a claim wearing a schema. So the profile does not travel as bare JSON. It travels as a verifiable credential: signed by the issuer, tamper-evident, and checkable by a stranger without ever calling us. We issue it as an SD-JWT VC, a selective-disclosure verifiable credential signed with an ES256 key we publish at did:web:validant.ai and withdraw through an IETF Token Status List, so the always-disclosed verdict verifies for anyone while the sensitive detail stays sealed until the holder chooses to reveal it. None of those are house formats. They are the building blocks of Switzerland’s official e-ID and trust infrastructure, swiyu, and of the EU digital-identity wallet, chosen so a seal slots into a national trust ecosystem instead of asking anyone to take our word for it. Machine-readable means a computer can parse the profile; verifiable means a computer can prove it is ours and unaltered. A finding you can rely on across organisations needs both.

Part five · How to read one of these

None of it is worth anything if a reader cannot use it in ten seconds. So here is the Part four record again, taken the way a sceptical reader should take it: from its foot upward, sensitivity first and pointing last.

The read, in orderWhat the record saysWhat it lets you claim, and what it kills
1 · Result, then its sensitivityno disparity detected, down to 3.0 ppNothing was found above three points. A two-point gap could still be sitting there, unseen. Whether two points matters is your business call, not our statistics to blur.
2 · EpochV2 · pinned 14 Jul · expires 20 OctTrue for one frozen version, and only until it lapses. Read it in November and it is history.
3 · ApertureA1 · behaviouralWhat the system does, never why. Quote it as proof of a clean internal mechanism and you have gone past the instrument.
4 · Claim classLimited, capped by the A1 apertureThe weakest link sets the ceiling. Borrowed from audit practice, the word means the finding supports a conclusion without being conclusive.
5 · PointingModel · Generative · External reviewerThis model alone. It says nothing about the deployer’s governance, and nothing about the predictive model underneath it.
The Part four record, read from its foot upward: sensitivity first, pointing last.

Five checks, and what comes out the other side is a claim you can lean on to a stated degree, about a stated thing, until a stated date.

The record read from its foot up: the lowest grade, Access A1, sets the ceiling on the claim above it, Limited.

What the five checks destroy is the sentence everybody actually wants, which is that the AI is fair. That sentence is not available at any assurance level. Anyone offering it is selling the omission.

Part six · From a reading to a record

A measurement is a moment. You pointed somewhere, you saw as well as your access, evidence and validity allowed, and you wrote down a grade. But a moment is easy to walk back. The quiet failure of most AI assurance is not that the first reading is wrong. It is that the target moves afterwards, and nobody notices. The metric that looked bad in March is quietly swapped for a kinder one in April. The scope narrows. The population that showed the gap is no longer in scope. The number improves and nothing real did. This is how precision becomes theatre: not by lying about the reading, but by moving what the reading was of.

So the reading has to become a record, and three things have to hold.

The objective must hold still. The moment you take a first honest reading, two of the pointing coordinates stop being editable: the body you are assessing and the pathway of harm you are watching for. Together they are the objective, and an objective you can rewrite is not an objective, it is a mood. From that moment the objective is frozen. You can still assess something else, but that is a new objective with its own record, not a quiet edit to this one. The third coordinate, audience, stays free, because audience never changes the numbers. The same frozen result is simply narrated to a builder, a governor, or the person it affects, in the register each needs. Who you are talking to is a lens, not a lever.

Assurance is earned, not declared. With the objective pinned, the seeing grades become the honest place for progress to happen. A first iteration might be Indicative: you had data but not the model, a single sample, no live monitoring. That is not a failure to hide, it is a starting line to record. The next iteration earns more: you gain the model, so access rises; you run proper confidence intervals with correction, so evidence rises; you turn on continuous monitoring, so validity rises. Each iteration seals its own grade and cannot un-earn it later. What the reader watches, across a stack of sealed iterations, is not a number bouncing around but a floor rising under a fixed goal. That is the difference between scoring well once and getting genuinely better at seeing the same thing. Only the second one is trust.

The headline for each iteration stays the strict rule from Part four: the assurance is the minimum of access, evidence and validity, never the average. A brilliant statistical test on data you were only allowed to glance at is still a glance. You cannot buy back access with cleverness, and the seal refuses to let you pretend otherwise.

CoordinateRuleWhy
Body and pathway, the objectiveOne value each, frozen when the first iteration closesAn objective you can rewrite is not an objective. If the target can move, improvement is meaningless.
AudienceAll of them, alwaysAudience is a lens, not a lever. The same frozen numbers are narrated to a builder, a governor or the affected person. Who you tell never changes what is true.
Seeing: access, evidence, validityGraded per iteration, sealed, can rise but never un-earnA first pass may be Indicative. Gaining the model, real confidence intervals and live monitoring each raise a grade, so a floor rises under a fixed goal.
What a record locks, what it lets vary, and what it lets rise.

From one seal to a programme

One seal covers one body. But a real deployment is never one body. When a company puts an AI system into the world, three bodies are in play at once: the model that makes the decision, the person the decision lands on, and the organisation that built and governs it. Trust in the system is not any one of them. It is the shape the three trace together. We have written before that digital trust is an orbit, not a pillar. This is where that stops being a metaphor.

One programme, three bodies, many iterations. Each iteration freezes its objective and grades how well it was seen; assurance is earned upward over time, and the programme seal binds each body’s latest sealed state into one verifiable posture, no stronger than its weakest body.

A programme is the container for one deployment’s three assessments. Each body keeps its own frozen objective and its own rising stack of sealed iterations. The programme draws them together into a single, publicly verifiable object: a programme seal. And the programme seal inherits the same honesty rule, one level up. Its assurance is not the flattering average of its three bodies. It is the minimum. If the model earns Reasonable assurance but the organisation’s governance is still only Limited, the whole programme is Limited, and the seal says so out loud. You cannot launder a weak body inside a strong average. The orbit is only as trustworthy as its least-seen body, which is exactly the truth a decision-maker needs and almost never gets.

One body is not a failure

The programme seal is the destination, not the toll gate. In principle a deployment has three bodies, and in principle you would seal all three. In practice you will often see one. You get query access to a model through an interface and nothing else: not the training data, not the governance minutes, not the population the decisions land on. It is tempting to read a one-body assessment as a weak one, and anything short of the full orbit as not worth publishing. That reading is wrong, and the framework already says why.

Start with the distinction the seeing dimension exists to protect: not being able to see something is not the same as seeing something bad. If you were granted only behavioural access to a model, your assurance is capped at the lower classes, and it should be. But that cap is a statement about your line of sight, not a verdict about the subject. It says “this is as far in as we were let.” It does not say “this system is unfair.” Rule R4 is exactly this wall: the assurance profile is metadata about our work, kept structurally apart from the fairness result, so that no reader can mistake a closed door for a finding. A low grade earned by a denied door is honest information of a specific kind. It tells a buyer or a regulator precisely how much of a claim rests on trust rather than on inspection, which is a thing they need to know and are almost never told.

A single body, honestly graded, is already a product. Financial assurance has lived here for a century. An auditor does not refuse to sign because they could not audit the whole economy; they state a scope, work to it, and issue assurance over exactly that scope and no further. “Limited assurance over this model’s predictive-fairness pathway, at behavioural access, on a pinned snapshot” is a real, defensible, sellable sentence. It is worth more than a glossy “our AI is fair” precisely because it says where it stops.

Which forces one honest correction to how the programme seal composes. The weakest-link rule is right for the bodies you assessed: you cannot average a weak governance reading away behind a strong model. But it must never run over bodies you did not assess, because “not seen” is not “seen and weak.” Treating an out-of-scope body as a zero would do two bad things at once. It would lie, by dressing an absence up as a low score. And it would wreck the one incentive the whole scheme runs on, by punishing an organisation for the parts of the chain it could not open, until the rational move is to attempt nothing. So the seal carries two axes, not one. Coverage, which bodies are in scope at all, is shown plainly and separately from assurance, how well each in-scope body was seen. The composite class is the minimum over the bodies actually assessed; every other body is marked “not asserted,” in the open, as a gap rather than a grade. You then climb both axes over time: a wider orbit as more of the chain opens to you, a higher floor as access, evidence and validity deepen on what you already hold.

Enlarge
The same idea as a working record. Three assessments nest inside one programme: each brick is a sealed iteration with its access, evidence and validity, each body’s verdict is its minimum, and the programme reads the minimum across the bodies it actually assessed. Here the model is seen well and earns Reasonable; the organisation is the weakest body and holds the whole programme to Limited; the person is drawn as planned, marked not asserted rather than scored as zero. Hover any tile for detail.

Why this is the whole point

Point honestly, see honestly, and then keep it honestly. Each of those, on its own, is a good habit. Together, and made into a signed, tamper-evident seal that anyone can check without asking us, they change what an assurance claim is. The pointing coordinates and the assurance grades travel inside the seal as structured, always-disclosed fields, kept separate from the fairness verdict itself, so a stranger can check how well it was seen without ever touching protected-group detail. It stops being a snapshot you take on faith and becomes a record you can audit: here is exactly where we aimed, here is exactly how well we could see, here is proof it has not moved, and here is the floor rising across every iteration. That is what it means to increase trust in a digital world. Not a louder claim, but a claim that holds still long enough, and shows its own limits clearly enough, that a stranger can rely on it. The seal is not a badge. It is the receipt.

Part seven · What we would like, and might not get

The tempting story goes like this. Assurance grading becomes procurement language. Buyers write minimum A2, evidence grade E2, sensitivity below three points, validity window under ninety days into their contracts. Providers who want to clear an A2 bar start exposing scores. The market pulls itself toward transparency without anyone having to appeal to conscience. It is a good story, and we would like it to be true.

We do not know that it will be, and there are decent reasons it might not.

Disclosing a detection limit is a competitive disadvantage for whoever goes first, because in a boardroom an unqualified clean bill reads better than a calibrated one. Plenty of buyers want a green light rather than a graded one, and a scheme that abolishes green lights is not self-evidently what the market is asking for. Regulators tend to specify process rather than statistical power, so an attestation regime can be satisfied end to end without anyone ever stating a minimum detectable effect. And the top two access tiers depend entirely on provider willingness, which for frontier models is trending closed rather than open. A scheme whose upper reaches are unreachable in practice might end up documenting nothing but its own ceiling, permanently.

Why anyone would open the door

All of this depends on access we cannot compel, which raises the only question that decides whether the scheme is a real product or a nice diagram: why would a company let us look, and look deeper over time?

Because access is the one dial the subject controls, and it sets the ceiling on the strongest thing the subject can walk away with. A costly, verifiable signal, in Spence’s sense, is one the honest can afford and the dishonest cannot cheaply fake, and access is exactly that at every rung. Ask to be taken at your word and you earn Indicative. Open your decisions to be measured and you earn Limited. Open your decisions and your confidence scores and you earn Reasonable, the auditor’s word for a powered, corrected study of your outcomes. Open the model itself and you reach the frontier above that: an interventional test that re-queries the model under a flipped attribute and confirms the verdict is not an artefact of the data, the strongest claim on offer. Each rung is a signal the dishonest cannot cheaply fake, and each is visible on the seal. The distance between “we were audited to the internals and it held” and “we asked you to take our word” stops being invisible. In a market where every vendor already claims to be fair, the claim is worth nothing and the proof is the whole asset. The seal is how you sell the proof.

And the pressure to buy it is no longer soft. The high-risk obligations of the EU AI Act apply from 2 August 2026, and breaching them carries fines of up to 15 million euro or three percent of worldwide annual turnover, with the heavier 35 million euro or seven percent ceiling reserved for the Act’s prohibited-practice tier. A Fundamental Rights Impact Assessment already lands on credit scoring and on life and health insurance pricing. New York City has required bias audits of hiring tools since 2023, and Colorado and others are close behind. In procurement, ISO/IEC 42001 has moved in barely a year from a differentiator to an entry ticket in regulated buying. In capital and insurance, one 2025 survey of corporate insurance buyers found more than nine in ten wanted cover for generative-AI risks and two-thirds would pay at least ten percent more for it, and you cannot underwrite a risk nobody has graded. And in the market for trust itself, a study across forty-seven countries found four in five people would be more willing to rely on an AI system when assurance mechanisms are attached to it. Each of these is a reason the door opens a little wider, because a stated, verifiable assurance profile is becoming the thing the regulator, the buyer, the insurer and the customer are all separately asking for. We would argue for it even if none of them were, because a verdict without a stated ceiling is not really a verdict. We just no longer have to argue alone.

Who gains, and what it costs to skip it

None of this is abstract, and it is not only the vendor’s problem. Every party around a deployed model has its own stake in whether a verifiable assurance record exists, and its own exposure when it does not. Here is what a stated seal is worth to each of them, and what each risks by leaving it out.

StakeholderWhat they gainWhat they risk by skipping it
Board and C-suite (the deployer)Duty of care discharged and documented, and a five-minute answer to “which fairness, measured how, audited by whom.” PwC modelled up to 4 percent higher valuation and 3.5 percent higher revenue for a robust responsible-AI programme over a compliance-only one; nearly six in ten executives say it lifts ROI and efficiency.Breaching the EU AI Act’s high-risk obligations costs up to 15 million euro or 3 percent of worldwide turnover, and a preventable Apple Card or Dutch childcare-style incident arrives as a bill in reputation and in human lives, years later.
Builders (engineering and data science)Fairness instrumented into delivery instead of a pre-launch scramble; disaggregation catches what aggregates hide (Gender Shades: over 90 percent overall accuracy masked a 0.8 percent versus 34.7 percent error gap between groups); a ready library of pre-, in- and post-processing interventions.Shipping that hidden disparity; expensive retraining after the fact; the quiet departure of the most principled engineers.
Risk, compliance and legalAudit-ready evidence produced as a by-product (EU AI Act Articles 10 to 15, the Fundamental Rights Impact Assessment, ISO 42001, NYC Local Law 144); a stated detection limit on every non-finding; a stated scope, so the claim cannot be over-read.Compliance treated as a research problem when it is now a calendar problem; an unfalsifiable “we use AI ethically” that collapses under one specific question; investigations that cost in fees and distraction even when they end in no fine.
AI vendor or model provider (the party asked for access)Openness turned into a priced, verifiable signal; Reasonable assurance as an RFP-clearing differentiator where ISO 42001 is now the entry ticket; the Mobley v. Workday agent-liability exposure blunted by an actual paper trail.Capped at Indicative, which reads in due diligence as “they asked us to take their word”; locked out of regulated procurement; co-defendant exposure with nothing to show, after the Workday collective was certified nationwide under age-discrimination law in May 2025.
Regulators, auditors and supervisorsA machine-readable, comparable assurance profile readable without ever touching protected-group data; coverage and detection limits stated, so oversight is possible at all; one grammar across many vendors.Continued opacity in which clean-sounding verdicts hide their own blind spots; nothing comparable across firms; enforcement that arrives years after the harm, as it did in the Netherlands, where the system ran for close to a decade.
Customers and affected individualsA plain-language reading of the decision and a real route to recourse (GDPR Article 22); the assurance that what went unseen is declared, not glossed; harms less likely to concentrate on those least able to absorb them.Wrongful denials and misidentifications that surface years later and cannot be undone; iTutorGroup’s software auto-rejected more than 200 older applicants and settled with the EEOC for 365,000 dollars in 2023.
Investors and insurers (capital)AI governance as a readable input to ESG and to the disclosure regimes arriving under CSRD and the SEC; a graded system is a different, underwritable object, and the demand is real (more than nine in ten corporate insurance buyers surveyed want cover for generative-AI risks, and two-thirds would pay at least 10 percent more).Capital and cover that will not sit behind an unverifiable posture. One study finds that overclaiming AI in a 10-K filing is followed by roughly 1.6 percent lower cumulative abnormal returns.
What a stated, verifiable seal is worth to each party around a deployed model, and what each risks by leaving it out.

“A seal converts an unprovable “trust us” into a costly, verifiable, priced signal. The party that does the work is the one that gets paid for it, and the party that refuses to be seen is the one left to explain why.”

Part seven · What is in it for whom

What is actually true today is narrower, and worth stating precisely. Colorado’s Regulation 10-1-1 requires quantitative testing, documentation and annual attestation from life insurers using external consumer data, with auto and health brought into the statute’s scope and their sector rules still in rulemaking. The NAIC model bulletin expects a documented AI-system programme. The EU AI Act puts risk assessment and pricing in life and health insurance in its high-risk annex. That is a real and accelerating move toward documented, testable claims. None of those regimes requires anyone to say how small a disparity their testing could have found. The direction is right. The specific thing this article argues for is on nobody’s statute book, and we should not imply otherwise.

There is precedent for the pattern arriving anyway. Graded assurance became ordinary in identity and in financial audit, in both cases through some mixture of professional convention and eventual codification, and in neither case quickly.

So: a proposal, not a forecast. We are adopting it unilaterally because we think a verdict without a stated ceiling is not really a verdict, and we would rather publish a modest claim that holds than a confident one that does not. It costs us the cleaner headline. We think that is the right trade, and we would be glad to be argued with by anyone who thinks the grades are wrong, the thresholds arbitrary, or the whole apparatus more precise than the underlying science can support. That last objection is the one we take most seriously.

Trust is assessed, not asserted. If your organisation deploys AI and wants its trust posture measured by an independent party, across fairness, explainability and governance, and reported with the ceiling stated rather than hidden, our closed beta opens to a small group in Q3 2026. Write to hello@validant.ai with the subject “Closed Beta,” or request a demo. Seats are limited and assigned in order of fit, not order of arrival.

Every finding ships with its access, evidence and validity grade.

In one line

Precision is not proof. A model that predicts perfectly is not fairer, it has stopped insuring anyone; a verdict that sounds certain is not proven, it has only hidden how little was seen. A number becomes something you can lean on when two things travel with it: the pointing, what it is about, and the seeing, how well it was seen. Publish the number on its own, without those two, and you have walked into the same trap as the insurer at the start of this piece: reading a precise number as a proven one.

“Trust is assessed, not asserted. An assessment that will not state its own ceiling has asserted something after all.”

Precision Is Not Proof

Sources and further reading

Share this post
Read next
Blanco-style technical line drawing of three spheres of different sizes held in balance by interlacing elliptical orbits, with soft coral washes, on a white ground.ResearchOpen to read
5 June 2026

Digital Trust Is an Orbit, Not a Pillar

Trust is not one more pillar to stack. It is the orbit three bodies trace together: the model, the person and the organisation. Why the three-body problem is the honest metaphor for trustworthy AI, and how to tell where you are in the orbit.

Read
Blanco-style line drawing of Daniel Glinz smiling, in a light beige blazer and open-collar shirt with a conference lanyard, holding a certificate that reads Best Paper Award, The Architecture of Digital Trust, on a clean white ground.ResearchOpen to read
24 July 2026

The Trust Problem Nobody Wants to Name

Everybody is spending on AI; almost nobody can say what they got back. A new paper, awarded the Best Paper Award at the IEEE Swiss Conference on Data Science and AI, names that gap and shows it is a trust problem wearing a technical costume: seven mechanisms, a four-layer architecture, an iceberg beneath it, and five design principles for closing it.

Read
Fine grey-blue technical drawing on warm paper of a single rack-mounted frontier-AI model unit gone dark, beside a large knife-blade breaker switch thrown to off and actuated by a sealed letter, with dashed data lines to small client terminals and a faint world map all cut by break-marks, one teal switch lever, a soft coral wash and a small human figure for scale.ResearchOpen to read
14 June 2026

When a Model Becomes a Munition

A frontier AI model was switched off worldwide by a single government letter. Reading the June 2026 shutdown through the three-body picture of digital trust: how the kill switch stopped being a metaphor, why the most governable lab was governed least carefully of all, and what single-vendor, single-jurisdiction dependence now costs a board.

Read