The Loop That Never Was: What Self-Driving Car Regulation Teaches Us About AI Governance in Professional Services
The phrase “human-in-the-loop” sounds like governance. It suggests a person sitting at a checkpoint, applying judgment, catching errors before they reach the client. In practice, it has become what IBM’s Phaedra Boinodiris calls liability laundering: a human’s name gets attached to a decision they were never empowered to meaningfully review, while they are measured on speed rather than oversight.
Amazon’s Eric Brandwine, speaking to The Register in June 2026, made the same point from a different angle: humans are inconsistent, get bored, and stop paying attention when the system is right 99% of the time. The 1% error passes through unnoticed.
The critique is converging from multiple directions simultaneously. But nobody has yet offered a structural alternative that answers the question: if HitL isn’t enough, what is?
Switzerland’s Ordinance on Automated Driving (VAF), effective 1 March 2025, answers that question by implication. It regulates a domain where physical harm, clear liability chains, and a single regulatory authority forced a level of specificity that professional services AI governance has avoided. The framework it produced - defined automation levels, mandatory event logging, layered liability, continuous recertification, retroactive regulatory power - is a template for what AI governance should look like when it stops pretending that “a human reviews the output” is an adequate strategy.
The Evidence That HitL Is Broken
The clinical study. Jabbour et al., JAMA 2023. ~450 clinicians (physicians, nurses, physician assistants). Baseline diagnostic accuracy for three diagnoses: 73.0%. When shown a systematically biased AI model, diagnostic accuracy dropped to 61.7%. The human did not catch the error. The human endorsed it. Nearly 67% of participants were not aware that AI models could be systematically biased.
The BMJ analysis. Toro-Tobon et al., BMJ 2026: “Clinician in the loop: a flawed solution for AI oversight.” Placing a clinician between the AI output and the patient shifts responsibility for AI safety from developers to doctors, without giving doctors the tools or authority to exercise that responsibility effectively.
The insurance market response. In January 2026, ISO issued three new generative AI exclusions for commercial general liability policies: CG 40 47, CG 40 48, and CG 35 08. Multiple major carriers adopted them within weeks. AIG and WR Berkley have filed for exclusions; Berkshire Hathaway and Chubb have already secured approval to attach them. Industry commentators have drawn the parallel to “silent cyber” from a decade ago - unintended, unpriced coverage lurking in traditional policy lines, now being explicitly excluded.
The latency trap. Errors in work produced in 2026 may not surface until 2030 or later. Professional indemnity policies are typically claims-made: the claim attaches to the policy in force when the claim is made, not when the work was done. If by 2030 the standard market excludes AI-related losses, the 2026 work is uncovered. The liability lands on the firm. Not the vendor. Not the model. The firm.
The Swiss Framework: HitL Regulation Designed for Actual Risk
Switzerland’s VAF ordinance entered force on 1 March 2025. It permits three specific use cases: motorway pilots (requiring a driver who can intervene), driverless vehicles on approved routes (supervised by a control centre), and automated parking without a driver present in marked car parks.
Five structural elements transfer directly to AI governance in professional services.
1. Defined Automation Levels
The ordinance anchors on SAE J3016, the international standard defining six levels of driving automation from L0 (no automation) to L5 (full automation, everywhere, always). Each permitted use case specifies its level, and the level determines the governance requirements.
This is the specificity that “human-in-the-loop” lacks. The phrase collapses at least three distinct risk profiles into one meaningless label:
| SAE Level | Driving Definition | Professional Services Analog |
|---|---|---|
| L2 | System handles steering AND speed; human monitors continuously | AI drafts documents; human reviews every output before use |
| L3 | System handles driving under defined conditions; human must respond to intervention requests | AI produces first drafts; human must be available to catch errors when the system flags uncertainty |
| L4 | System handles all driving in defined conditions; no human needed for routine operation | AI handles entire bounded process (e.g. standard contract generation); human oversees exceptions and edge cases |
A firm using AI to draft standard contracts and having a partner review each one is operating at L2. A firm using AI to handle entire routine litigation workflows end-to-end, with a senior solicitor monitoring dashboards and intervening on exceptions, is operating at L4. Each level requires different governance, different liability allocation, different insurance coverage, different training for the human reviewer. Most firms have not made this distinction.
2. The Driving-Mode Memory (Black Box)
Art. 7 VAF requires automated vehicles to record specific events: emergency manoeuvres, system failures, collisions, activation and deactivation of the automation system. Each record must include the type of event, time stamp, and position. Art. 25g para. 3 SVG restricts who can read the data and for what purpose - only competent authorities for accident investigation or traffic-rule violations.
This is a template for agentic AI logging in professional services. Today, when an AI system produces an output that later proves defective, most organizations cannot reconstruct which model version was used, what input context was provided, what intermediate steps the system took, who reviewed the output, how long they spent, or whether they modified it before approval.
What a professional services driving-mode memory should record (author’s analytical recommendation, not regulatory text):
- Model identity and version
- Full input context (prompt, retrieved documents, tool outputs)
- Full output (including intermediate steps for agentic systems)
- Human reviewer identity and time spent on review
- Reviewer action (approved, modified, rejected)
- Confidence scores or uncertainty metrics where available
- Timestamps for every step
3. Layered Liability (Author’s Analytical Framing)
The Swiss framework does not explicitly label a “three-layer” model. However, reading Art. 58 SVG (keeper liability), the Product Liability Act (manufacturer liability), and the ordinance’s treatment of driver responsibility together, a coherent structure emerges:
| Role | Automotive Liability | Professional Services Analog |
|---|---|---|
| System manufacturer / vendor | Liability for software and sensor faults under product liability law | AI vendor / model provider |
| Operator / professional | Liability if human was controlling the vehicle at time of incident | Professional who approved or used the output |
| Keeper / firm | No-fault, risk-based liability under Art. 58 SVG - regardless of who was driving | Firm that deployed and maintains the system |
The key principle: the vehicle keeper remains liable under a no-fault, risk-based standard regardless of whether the system was driving autonomously. Automation does not transfer liability away from the operator. It redefines what constitutes operator failure. Instead of “you drove badly,” it becomes “you maintained inadequately, selected the wrong system, or failed to supervise.”
Applied to professional services: the firm remains liable for AI-generated output regardless of whether a human reviewed it. The question shifts from “did someone review it?” to “did the firm maintain appropriate oversight systems, select appropriate tools, and empower the reviewer with sufficient time, context, and authority?”
4. Continuous Recertification
The ordinance requires that manufacturers hold valid certificates for cybersecurity management (UN Regulation No. 155) and software update management (UN Regulation No. 156) for the entire supported operating period of the vehicle. Certificates expire. Recertification is required. The compliance obligation does not end at type-approval.
Art. 6 VAF goes further: ASTRA may declare new provisions applicable retroactively to vehicles already approved and in circulation - for example, after a security incident affecting certain vehicle types.
Compare this to professional services AI governance. Most firms assess a tool once during procurement and never revisit it. The vendor updates the underlying model. The data inputs shift. The regulatory environment changes. The original assessment is stale within months, and nobody recertifies.
5. The Oversight Function
For driverless vehicles, the ordinance requires supervision by operators from a control centre. The source material establishes this requirement but does not specify the operational details beyond it. The core point stands: even at L4, supervision is required, just structured differently than at L2 or L3.
The contrast with typical professional services practice is stark. The typical HitL today: a busy partner with billable hours targets, a full inbox, and fifteen minutes between meetings is expected to review AI-generated output thoroughly enough to catch subtle errors. They have no dashboard. No training in AI error patterns. No escalation protocol. No dedicated time. They are not a control centre operator. They are a bottleneck with a professional indemnity policy.
Agentic AI: The Governance Gap
Beale Law’s analysis of the EU court ruling on marketplace platforms flags the concern directly: “It should be noted, however, that not all AI applications operate on this basis; some are designed to function autonomously, without human review of individual outputs, which raises distinct questions about accountability and risk allocation.”
The Swiss framework provides the language for this distinction. An L2 system allows the human to drive and the system to assist. An L4 system allows the system to drive and the human to oversee exceptions. These require fundamentally different governance structures.
IDC argued in March 2026 that agentic AI constitutes critical infrastructure: once a system can see proprietary knowledge, shape work products, and connect to tools, it stops being a productivity layer and becomes part of the operating core. BCG published an analysis in June 2026 arguing that agentic AI requires a new approach to data risk management because agents take actions, not just produce outputs.
The Swiss framework’s answer: defined automation levels, mandatory event logging, continuous recertification, clear liability allocation, and oversight structures proportional to the level of autonomy. Every element applies.
The Governance Framework, Adjusted
Phase One: Identify and Assess - With Automation Levels
The Swiss ordinance forces every vehicle to declare its automation level. Every professional services AI system should do the same.
Practical step: Classify every material AI system by SAE-equivalent level (L1-L5). Document the classification and the rationale. The classification determines: the required human oversight model, the logging requirements, the recertification cadence, the liability allocation, and the insurance implications.
Phase Two: Design and Implement - With a Driving-Mode Memory
Before deployment, define what events are logged, what data elements are recorded, who can access the logs, and how long they are retained.
Practical step: Create a logging specification based on the functional requirements of Art. 7 VAF (design pattern, not regulatory mandate). Required fields: event type, timestamp, model version, input context, output, human reviewer identity, reviewer action, time on task.
Phase Three: Manage and Monitor - With Continuous Recertification
Define the recertification cadence for each system before deployment. Quarterly for L4 and high-risk L3 systems. Annually for L2 and low-risk L3 systems.
Practical step: Run a blind error-detection test at least annually. Introduce a known error into AI-generated output. Measure whether the reviewer catches it. Track the detection rate over time. If it drops below a defined threshold, escalate to a higher oversight level or suspend deployment. This is the test that most HitL claims cannot pass, because they have never been measured.
Where the Comparison Breaks
Physical harm gets regulated faster. A self-driving car that crashes kills someone. The causal chain is visible. The regulatory response is politically necessary. AI that produces a flawed legal argument, an incorrect structural calculation, or a misdiagnosis that goes unnoticed for years causes harm that is diffuse, delayed, and harder to attribute. Regulators move slower when the harm is invisible.
No single ASTRA exists for AI. The Swiss framework benefits from a single regulator with clear authority over a bounded domain. AI governance spans multiple regulators, jurisdictions, and use cases. No single body can declare retroactive requirements across all AI deployments. This fragmentation makes internal governance more important, not less - the firm must be its own ASTRA, conducting conformity checks and maintaining the authority to suspend systems.
The core rhetorical move - using physical-harm regulation as a template for diffuse-harm professional services regulation - is argued intentionally and honestly. The analogy is instructive precisely where it holds and where it breaks.
TL;DR
- HitL is failing on documented evidence: Jabbour et al., JAMA 2023 showed diagnostic accuracy dropped from 73.0% to 61.7% with biased AI - and 67% of clinicians didn’t know AI could be biased. Toro-Tobon et al., BMJ 2026 argues “clinician in the loop” is a flawed model that shifts responsibility without empowering reviewers.
- IBM calls the current approach “liability laundering”; Amazon argues humans are inconsistent and HitL fails through normalization of deviance.
- Insurance market is already responding: ISO AI exclusions CG 40 47, CG 40 48, CG 35 08 effective January 2026. AIG and WR Berkley have filed for exclusions; Berkshire Hathaway and Chubb have secured approval.
- Swiss VAF ordinance (March 2025) provides a structural alternative: defined SAE automation levels (L2/L3/L4), mandatory driving-mode memory (Art. 7), layered liability with no-fault keeper responsibility (Art. 58 SVG), continuous recertification (UN R155/R156), and retroactive regulatory authority (Art. 6).
- Agentic AI requires distinct governance: IDC calls it critical infrastructure; BCG argues it rewrites data risk management rules because agents act rather than just produce output.
- The comparison breaks in two places: physical harm gets regulated faster than professional harm, and no single AI regulator exists - making internal governance more critical, not less.
- Practical action: classify every AI system by automation level (L1-L5), implement driving-mode memory logging before deployment, run blind error-detection tests annually, define recertification cadences, and build governance as if a single regulator will audit it - because eventually, someone will.
The self-driving cars are driving under rules that professional services AI has not yet imagined for itself. The ordinance passed in March 2025. The vehicles are on the road. The question is whether professional services will build its equivalent before the first claim demonstrates that “a human reviewed it” was never a sufficient answer, or after.