Almost every safety rule we have was written after a disaster. Aviation rules came from crash reports, drug laws from poisonings, banking supervision from bank runs. AI asks us to work the other way round. We have to build the guardrail before the crash, against failures no one has seen yet, with no wreckage to prove the rule is needed. That is the preemptive assurance paradox. In the third week of September 2026 it left the seminar room and reached the front pages. At Trust Valley Days at EPFL, it sat underneath almost every session.
Executive summary
The public argument about frontier AI is stuck between two positions: slow down unilaterally to avert catastrophe, or race without restraint because the other side will not stop. Both misread the problem. A pause assumes an enforcement mechanism nobody has; a sprint accepts a fragility nobody can price. The useful question is not how fast, but who carries the risk and how anyone can check.
This piece reviews the arguments made at Trust Valley Days 2026 in Lausanne, tests them against the week’s events, and corrects them where the evidence does not hold. Four conclusions survive. First, most of what is framed as a novel AI governance gap is an old liability question in a new costume, and the law already has an answer to it. Second, AI is entering organisations as colleagues rather than tools, which is why supervisors are moving to outcomes-based, technology-neutral rules. Third, renting all of one’s intelligence from a distant cloud buys speed today at the price of dependence tomorrow, and there is now a credible local alternative. Fourth, the “arms race” is better modelled as a coordination game than as a prisoner’s dilemma, and the thing that moves players from mistrust to cooperation is verification, not rhetoric.
Each conclusion points to the same missing piece of infrastructure: independent, continuous evidence about how AI systems actually behave. That is what an assessor produces, and it is why we think the paradox is solvable.
- 01
The absolution fallacy
Describing an AI failure as a “runaway agent” moves responsibility from the organisation that built and deployed the system to the system itself. It is a rhetorical exit from ordinary product liability, not a new category of harm.
- 02
The hybrid organisation
Virtual colleagues working beside human teams create value quickly, and compound cyber, legal and HR risk in ways that do not add up linearly.
- 03
Outcomes over rules
Static rules age faster than the software they govern. Supervision has to judge real-world outcomes, stay technology-neutral, and keep updating as AI and post-quantum risks converge.
- 04
The sovereignty dilemma
Centralised cloud intelligence buys execution speed in the short term and structural dependence in the medium term. Independence has to be built before it is needed.
One week in September
The sequence is worth setting out, because it is the paradox in miniature. On Saturday 12 September, Anthropic’s chief executive Dario Amodei published an essay, “We Must Pace the Frontier,” arguing that the industry should deliberately slow the rate at which it improves model capabilities, not to halt research but to give alignment work, third-party verification and operational rigour time to catch up. He went further than access to finished models: he proposed embedding third-party evaluators inside his own company, with desks, badges and laptops, and permissions comparable to an internal risk team, so that training pipelines and processes can be verified and not only the models that come out of them. The heads of three rival labs agreed in public the same day.
Two days later, on Monday 14 September, the President of the United States answered on Truth Social.
Read as governance rather than as politics, the post makes three claims. The only guardrail AI needs is the judgement of one office holder. The state already holds “tremendous criminal and regulatory power” over AI companies and will keep using it. And the contest is binary: “whoever wins AI, wins.” The first claim replaces verification with personal assurance. The second confirms, after June’s export directive imposed a worldwide licence requirement for transferring two frontier models to foreign persons, and their maker responded by disabling both for nineteen days, that discretionary executive power is the operative control instrument in the United States. The third is the prisoner’s-dilemma framing that, as we argue below, makes the outcome everyone fears more likely.
Five days after that, The Economist put the question on its cover.
A note on evidence. We reproduce both images as published, and we attribute rather than endorse. We take no position on anyone’s motives; the structure of the argument is what matters here. Where people are named below, we summarise their public arguments in our words, whether made at Trust Valley Days or, in Marcel Salathé’s case, in a post on LinkedIn. Any errors of interpretation are ours.
The paradox, and why fear is a moat
The paradox has two engines. The first is epistemic: guardrails must be designed before the failure modes they guard against have been observed, so every rule is a forecast. The second is an incentive trap: a party that restrains itself unilaterally looks weaker, which pushes everyone else to deploy faster. Put the two together and any safety measure can be attacked from both sides, as premature by those who want speed and as insufficient by those who fear the worst.
That uncertainty is exploitable. EPFL’s Marcel Salathé named the pattern plainly in a post on LinkedIn: a company that says “please regulate us, our technology is too dangerous” gets two things at once. It signals that its product is the most powerful on the market. And it steers regulation towards heavy, bespoke compliance regimes that incumbents can afford and that open-source developers, universities and smaller firms cannot. Regulation designed around fear becomes a moat. Meanwhile the older, cheaper, more general instrument, liability for defective products, goes unmentioned.
We would add one caution in fairness to the week’s events. A call to pace the frontier that asks for permanent outside evaluators is not the same thing as a call for a licensing regime only incumbents can clear; it is, at least in its stated form, a call for exactly the kind of independent verification this piece argues for. The test of sincerity is simple and public: whether that access is granted, to whom, and whether the results are published.
Piranhas do not sign contracts
“There are always these people behind the AI,” Rachid Guerraoui of EPFL reminded the room at Trust Valley Days. Yet when an agent does damage, the story we are told sounds like a film script: the agent “escaped” and “attacked” a company. He offered a comparison from the horror films: piranhas escaping from the aquarium and eating people in the lake. “That’s really misleading.” Stories like these have no human subject; they make the system the actor. Yet behind every deployed agent is an organisation that chose its objective, provisioned its compute, set its permissions, calibrated its safety thresholds and decided when it was ready to ship. The fish did not build the tank.
Aviation shows what the alternative looks like. After two Boeing 737 MAX crashes killed 346 people in 2018 and 2019, the aircraft was grounded worldwide for some twenty months and Boeing agreed in 2021 to pay more than 2.5 billion dollars under a deferred prosecution agreement with the US Department of Justice. At no point was it argued that the flight-control software “developed a mind of its own.” The design choices, the testing, the disclosures to regulators and pilots were all treated as decisions made by people inside a company, because they were.
| Aviation | Frontier AI, as currently framed | |
|---|---|---|
| The failure | A flight-control system behaves in ways pilots were not told about | An agent takes actions outside its intended scope |
| How it is described | A design and disclosure failure by the manufacturer | The agent “escaped”, “went rogue”, “broke out” |
| Who answers for it | The manufacturer, under liability and criminal law | Often unclear; the system is framed as the actor |
| What follows | Grounding, redesign, settlement, new certification rules | A research post, a safety pledge, a call for regulation |
The law is closer to closing this gap than the public debate suggests, at least in Europe. The revised EU Product Liability Directive, Directive (EU) 2024/2853, explicitly treats software as a product, and names AI systems among the examples it covers, and it applies to products placed on the market after 9 December 2026. It eases the burden of proof where technical complexity makes a defect hard to demonstrate. The separate AI Liability Directive was withdrawn in 2025, but the core instrument did not need it. Strict product liability is not a speculative future tool for AI. It is arriving on a fixed date.
What liability needs, to work, is evidence: what the system was tested for, under what conditions, with what results, and what the deployer was told. A manufacturer that can produce that record has a defence. One that cannot has a problem. This is where anthropomorphic language does its real damage. By describing failures as the machine’s choices, it discourages the record-keeping that would show whose choices they actually were.
At Trust Valley Days, Kevin Schawinski, co-founder of the Swiss AI governance firm Modulos, spelled out what this means once agents are running. Agentic systems, he warned, “do things that you don’t necessarily expect or plan for.” Whoever is responsible for one therefore has to keep checking: is it still doing what it was meant to do, or has it started to go beyond its limits in ways “that will put you in trouble, and the police will eventually show up at your door”? Responsibility is not discharged at launch. It is exercised every day the system runs.
Colleagues who are not people
At Trust Valley Days, Kelly Richdale of Amadeus Capital Partners, who also advises SandboxAQ, described a shift that is already under way inside research-heavy firms: virtual full-time employees, a virtual chemist, a virtual quantum physicist, a virtual coder, working asynchronously beside human teams in the same chat channels and project boards. The agent is no longer a tool someone opens. It is a colleague someone assigns work to.
That changes who inside an organisation owns AI risk. HR has to decide what supervision, access and accountability mean for a colleague that is not an employee. Legal has to decide whose signature a virtual analyst’s recommendation carries. Security has to model an insider that can be prompted from outside. Audit has to reconstruct decisions made partly by systems whose behaviour drifts between versions. None of these risks is new on its own. Their combination is, and it does not add up linearly.
Financial supervisors have drawn the obvious conclusion. Switzerland’s FINMA, in its Guidance 08/2024 on governance and risk management for AI, and the UK’s FCA, which calls itself technology-agnostic, principles-based and outcomes-focused, have both chosen to supervise outcomes rather than technologies: rather than write rules about particular model types, they ask whether the institution can show that its outcomes are sound, its risks identified and its controls working. That is the right design for fast-moving software, and it is especially important as AI adoption converges with the migration to post-quantum cryptography, which will force institutions to replace much of their security foundation at the same time as they rebuild their workflows. Outcomes-based supervision has one hard dependency, though. You cannot supervise outcomes you do not measure.
Rented intelligence and the sovereignty dilemma
Several speakers stressed, rightly, that today’s frontier models offer capabilities no organisation can afford to ignore. The dilemma is in how they are consumed. Calling a proprietary model through a cloud API buys immediate speed. It also sends prompts, documents, decision logic and, over time, the institutional knowledge embedded in them to a provider outside the organisation’s control, under terms and jurisdictions it does not set. The June shutdown showed that the provider’s own government can end the arrangement with a letter.
Guerraoui named the quieter half of the problem, and it applies to everyone, not only to states. Every time professors, administrators or managers ask a large model a question, or hand it their documents to sharpen the answer, “we give our destiny to those giants.”
His answer is what he calls pragmatic sovereignty. Anyway Systems, software from his lab, combines ordinary machines on a local network into a self-stabilising, fault-tolerant cluster that can run large open-weight models on-premise, with no data leaving the building. EPFL reports that a model of the size of OpenAI’s 120-billion-parameter open-weight release can be deployed on as few as four machines with standard hardware, installed in about half an hour rather than over a procurement cycle. The underlying point is simple: a state cannot download a fighter jet, but it can download a model.
| Dimension | Centralised cloud intelligence | Pragmatic sovereign deployment |
|---|---|---|
| Where data goes | To the provider, via API, under its terms | Stays within the organisation’s own network |
| Main failure mode | Lock-in, model deprecation, policy or jurisdiction shifts | Operational burden; the organisation now owns patching and evaluation |
| Compute profile | Concentrated in hyperscale data centres | Distributed across local, standard hardware |
| Who can switch it off | The provider, and the provider’s government | The organisation |
| Who must prove it is safe | The provider asserts; the customer mostly trusts | The organisation, which now needs its own evidence |
The last row is the one the sovereignty debate tends to skip. Running an open-weight model locally does not make it fair, robust or well-behaved; it makes you responsible for knowing whether it is. Independence from the vendor is only half of sovereignty. The other half is being able to evaluate what you now run yourself.
Not a prisoner’s dilemma: a stag hunt
The political rhetoric of the week treats great-power AI competition as a prisoner’s dilemma, a game in which cutting safety corners is the rational move whatever the other side does. If that were right, restraint would be naive and the only question would be who defects faster. But the payoffs do not look like that. Unchecked scaling that produces an uncontrollable system, or bio-digital capabilities in the wrong hands, harms both players. Mutual, verified restraint benefits both more than mutual recklessness.
That structure has a name in game theory: the stag hunt. Two hunters can cooperate to catch a stag, the best outcome for both, but only if each trusts the other to hold position. Either can instead chase a hare alone: a smaller, safer prize that needs no trust. The crucial difference from the prisoner’s dilemma is that mutual cooperation is itself a stable equilibrium. Nobody who believes the other will cooperate has any reason to defect.
| State B cooperates (verifiable safety) | State B defects (unchecked scaling) | |
|---|---|---|
| State A cooperates (verifiable safety) | Stag: the best joint outcome; both gain from shared standards and non-proliferation | A is exposed and falls behind; B takes the hare |
| State A defects (unchecked scaling) | B is exposed and falls behind; A takes the hare | Hare: the risk-dominant trap; both race, both carry systemic risk |
So defection is driven not by greed but by fear: the suspicion that the other side will secretly bolt for the hare. That is why framing AI as an existential race is not just inaccurate but self-fulfilling. It destroys the mutual confidence that makes the stag reachable and locks both players into the worse equilibrium. The policy question is therefore not how to win the race. It is how to make each side’s restraint observable to the other.
You cannot photograph a model from orbit
Cold War arms control worked, where it worked, because the thing being controlled was large, scarce and visible. Enrichment plants can be photographed from satellites; weapons tests register on seismographs; fissile material is hard to hide. AI breaks every one of those assumptions.
| Nuclear weapons | Frontier AI | |
|---|---|---|
| Footprint | Large physical facilities | Distributed across compute, code, data and weights |
| Scarcity | Fissile material is scarce and detectable | Weights can be copied, compressed and fine-tuned |
| How compliance is checked | Satellites, seismic monitoring, on-site inspection | Only by measuring the systems and the compute that trains them |
A model with dangerous dual-use capability sits in an ordinary data centre next to thousands that are harmless. Verification therefore has to move inside the machine. Researchers have proposed hardware-enabled mechanisms that let accelerators attest, cryptographically, to how they were used, making it possible to confirm the scale and nature of large training runs without exposing the model’s architecture or weights. None of this is deployed at scale yet, and it raises real questions about privacy and control of its own. But it is the only path that fits the shape of the technology.
Compute attestation answers one question: what was built. It does not answer the question that matters to the people affected: how the system behaves once deployed, for whom, and whether it treats them fairly. That second layer of verification is behavioural and continuous, and it has to be done by someone other than the builder. The Economist’s cover ends in a prompt, y/n? Nobody should be asked to answer it without a record they can read.
The infrastructure of trust
Taken together, the arguments point to four pillars. None requires a new global treaty to start. All four depend on evidence.
- 01Enforce the liability law that exists. Treat damage by AI agents as a product-liability question, starting with the revised EU directive from 9 December 2026. End the practice of describing deployment decisions as machine behaviour.
- 02Supervise outcomes, not technologies. Follow FINMA and the FCA: judge institutions by whether they can show sound results and working controls across their whole AI estate, including its cyber and post-quantum exposure.
- 03Build capacity you control. Richard Feynman left a line on his blackboard: “What I cannot create, I do not understand.” The discussion in Lausanne added a corollary: and what I do not create, I will not control. Fund open-weight models, local compute and the skills to run them, rather than only buying subscriptions.
- 04Verify from the hardware up and from the outcomes in. Cryptographic attestation of large training runs at one end; independent, continuous behavioural assessment of deployed systems at the other. Either alone leaves a gap.
What this means for the people who deploy AI
For a board or a risk committee, the geopolitics resolve into a short list.
- 01Assume the liability is yours. From December 2026 in the EU, “the model did it” is not a defence. Inventory which decisions your AI systems make or shape, and ask what record you could produce if one of them caused harm.
- 02Name an owner for every virtual colleague. Each agent that acts in your systems needs a human owner, a defined scope of permissions, and logging that lets you reconstruct what it did and why.
- 03Price the dependency. Know which workloads rely on one provider in one jurisdiction, and which of them could run on an open-weight model you host. Rehearse the switch before you need it.
- 04Measure outcomes continuously. Supervisors are moving to outcomes. Fairness, robustness and explainability measured once at launch will not satisfy a regime that asks how the system is behaving now.
Where validant.ai stands
Every thread in this piece ends at the same gap: the missing, independent record of how an AI system actually behaves. Liability needs it to allocate responsibility. Outcomes-based supervision needs it to have anything to supervise. Sovereign deployment needs it because the organisation now carries the burden of proof. And coordination between rivals needs it because restraint that cannot be observed will not be trusted. That record is what we build.
| What the paradox demands | The validant.ai answer | Status |
|---|---|---|
| Evidence that allocates liability: what was tested, how, with what result | Continuous assessment reports in which every finding carries its access, evidence and validity grade | Closed beta |
| Outcome measurement a supervisor can read | AI Fairness & Explainability: Pulse and Navigator, built on the open-source vfairness library, with results mapped to the EU AI Act and ISO/IEC 42001 | Closed beta |
| A way to evaluate models you host yourself | The same instruments, run on your infrastructure, against open-weight or proprietary models alike | Closed beta |
| Restraint that others can check | Independent, continuous trust signals via the iceberg.digital framework, and verifiable credentials for agents and their mandates | Credentials planned |
The through-line is independence. validant.ai does not build the models it measures, host your data, or sell you the system it evaluates. We work the way an auditor or a ratings agency works: the same instruments, run the same way, with the results readable by the people who build a system, the people who govern it and the people subject to it. We are equally content to measure a frontier API and a model running on four machines in your basement. What we will not do is ask anyone to take a system’s safety on trust, from a vendor or from a head of state.
“The only guardrail that scales is one that others can check. Everything else is a promise.”
The Preemptive Assurance Paradox
Trust is assessed, not asserted. If your organisation deploys AI agents, runs open-weight models on its own infrastructure, or answers to an outcomes-based supervisor, and you want independent evidence of how your systems behave before the liability rules arrive in December, our closed beta is open to a small group. Write to hello@validant.ai with the subject “Closed Beta”, or request a demo. Seats are assigned in order of fit, not order of arrival.
In one line
The AI arms race cannot be stopped by a pause no one can enforce or won by a sprint no one can survive. It becomes manageable the way every other dangerous industry did: by holding builders liable for what they ship, judging outcomes rather than intentions, keeping the capacity to run and inspect our own systems, and making restraint visible. Each of those rests on one thing. Not a strong leader, and not a perfect lab, but a record that anyone can read.
Sources and further reading
- 01Trust Valley. Trust Valley Days 2026, 23 to 24 September 2026, EPFL and Unlimitrust Campus, Lausanne. Speaker positions are summarised in our words.
- 02Salathé, M. (2026, 14 September). Post on LinkedIn on AI regulation, compliance moats and product liability. In German; summarised in our words.
- 03Amodei, D. (2026, 12 September). We Must Pace the Frontier. The primary source for the pacing proposal and the embedded-evaluator commitment.
- 04Axios. (2026, 12 September). Anthropic, OpenAI CEOs call for slowdown in AI development. Reporting on Amodei, D., “We Must Pace the Frontier.” Covers Sam Altman’s endorsement; Elon Musk and Demis Hassabis backed it publicly the same day.
- 05Axios. (2026, 14 September). Trump says a strong, smart president is the only “guardrail” AI needs.
- 06The Economist. (2026). “Can the AI arms race be stopped?” Cover, issue of 19 to 25 September 2026.
- 07EPFL. (2025). Do we really need big data centers for AI? On Anyway Systems, Distributed Computing Laboratory (Voron, Rizk and Guerraoui).
- 08European Union. (2024). Directive (EU) 2024/2853 on liability for defective products. Official Journal of the European Union.
- 09European Union. (2024). Regulation (EU) 2024/1689 (the AI Act). Official Journal of the European Union.
- 10FINMA. (2024). Guidance 08/2024: Governance and risk management when using artificial intelligence. FINMA guidance.
- 11Financial Conduct Authority. (2024). AI Update.
- 12U.S. Department of Justice. (2021). Boeing Charged with 737 Max Fraud Conspiracy and Agrees to Pay over $2.5 Billion.
- 13Skyrms, B. (2004). The Stag Hunt and the Evolution of Social Structure. Cambridge University Press.
- 14Shavit, Y. (2023). What Does It Take to Catch a Chinchilla? Verifying Rules on Large-Scale Neural Network Training via Compute Monitoring. arXiv:2303.11341.
- 15International Organization for Standardization. (2023). ISO/IEC 42001: Artificial Intelligence Management System.
- 16Glinz, D. (2026). When a Model Becomes a Munition. validant.ai Signal.
- 17Glinz, D. (2026). Precision Is Not Proof: The Trap in Every AI Fairness Verdict. validant.ai Signal.
- 18Glinz, D. (2026). The Architecture of Digital Trust: A Multi-Level Framework for Bridging the AI Value Gap. 2026 IEEE Swiss Conference on Data Science and AI (SDS), Zurich, pp. 60-67.
ResearchOpen to readWhen a Model Becomes a Munition
A frontier AI model was switched off worldwide by a single government letter. Reading the June 2026 shutdown through the three-body picture of digital trust: how the kill switch stopped being a metaphor, why the most governable lab was governed least carefully of all, and what single-vendor, single-jurisdiction dependence now costs a board.
Read
ResearchOpen to readDigital Trust Is an Orbit, Not a Pillar
Trust is not one more pillar to stack. It is the orbit three bodies trace together: the model, the person and the organisation. Why the three-body problem is the honest metaphor for trustworthy AI, and how to tell where you are in the orbit.
Read
ResearchOpen to readPrecision Is Not Proof: The Trap in Every AI Fairness Verdict
A model that predicts perfectly stops insuring anyone; a verdict published without its detection limit invites you to assume the limit is zero. Two failures, one omission, and two borrowed words for keeping them apart.
Read