06 · From model evaluation to system assurance
Assurance is a structured, challengeable argument that evidence supports a specified claim about a particular system under stated assumptions. Its central discipline is matching the strength and scope of the claim to what was actually established.
A model evaluation can provide evidence about task competence, instruction handling, or other measured behavior. Better judgment could prevent a wrong proposal, recognize conflicting evidence, or reduce the work of review. I then ask which system outcome the evaluation predicts or constrains, and under which configuration.
In the account example, a model may interpret test tickets correctly while the deployment uses a different identity mapping, retrieves stale notes, or allows a separate manual route. Whether the score predicts service outcomes is an empirical question. Anything the evaluation leaves out needs additional support if the deployment claim relies on it.
Three different promises#
The word “safe” can conceal materially different assertions. Consider three more precise claims about the hypothetical workflow:
| Promise | Path and mechanism evidence | Outcome evidence and limits |
|---|---|---|
| A change without the specified authorization cannot execute under named conditions | Coverage of in-scope paths, justified exclusions, and evidence that the mechanisms enforce the property | Tests can expose bypasses; passed attempts alone cannot establish universal coverage. A valid in-scope bypass defeats the claim. |
| Effects remain within a declared aggregate limit during an incident | Evidence that constraints support the combined maximum, including interacting paths, shared limits, accepted work, and relevant rates and intervention times | Measurement must cover accumulating effects over the declared period. An in-scope exceedance defeats a hard bound; unobserved work leaves it unsupported. Finite observations alone cannot establish a universal bound. |
| A design reduces unauthorized changes compared with an alternative while meeting service and cost requirements | An account of affected paths and mechanisms, sufficient to interpret the comparison and identify displaced harm | A credible comparison measures harmful outcomes, legitimate completion, delays, and full costs, with uncertainty. Insufficient benefit, unacceptable cost or utility, or displaced harm can undermine the claim. |
A comparative claim can permit residual risk. A hard bound promises a limit throughout its named conditions. A statistical bound needs a stated population, exposure, and supporting assumptions. It must specify whether it bounds the probability of exceeding a loss threshold, confidence in an estimated risk bound, or both.
For statistical and comparative claims, aggregate measurement can cover interacting paths together; a separate outcome estimate for every path is unnecessary. A narrowly specified bound may also be easier to assess than a general assurance about the whole service.
I judge the required evidence against the severity and reversibility of the loss, the exposure the entrusted authority permits, and the uncertainty that remains. Comparative improvement can still leave an unacceptable risk. For severe or irreversible losses, I require an explicit case for accepting residual risk: the paths driving it, the prevention or limits that can be supported, and the consequences of alternatives, delay, or refusal.
For severe or irreversible losses, I treat a case resting only on comparative estimates or proxies as insufficient unless it shows that a supportable prevention or limit on those paths is unavailable or would itself cause greater harm, including harm from delay or refusal. Meeting that exception still requires the residual-risk case. Categorical or bounded promises strengthen it only where their scope addresses the loss. If the evidence cannot support the commitment, the entrusted work or authority must narrow, or reliance must wait.
Formal reasoning can establish a property under an explicit model and assumptions. We then need evidence connecting that model to the implementation in use. The policy itself also needs justification: a service can follow a rule that leaves out a protected interest. The account-recovery conflict shows why enforcement and justification need separate arguments.
What an adversarial evaluation establishes#
A useful attack result identifies what the attacker controlled and what effect occurred. Controlling a retrieved message differs from changing a tool implementation or the identity service. Those starting conditions determine which trust boundary the test examines. An evaluation that grants a stronger foothold can usefully stress a dependency, provided the conclusion preserves that foothold.
A deliberately adversarial model policy lets researchers test controls against subversion without assuming that naturally occurring behavior has the same intent. AI Control: Improving Safety Despite Intentional Subversion compares protocols in a constructed programming-task setting where the red team seeks exploitable backdoors. It measures useful task performance alongside its scenario-level safety outcome. AI Control, §2, pp. 3–4
The protocol uses a protected GPT-4 channel to simulate limited trusted human labor, with a budget covering up to 2% of problems, primarily for auditing. It can also run test cases for one submitted solution per problem, while the attack pipeline has no tools and retains no state between calls. Usefulness and safety are measured in separate runs, using the original GPT-4 and the attack policy respectively.
The audit depends on a further assumption: because auditor and attacker both use GPT-4, the experiment assumes the auditor can construct an exploit exactly when the attacker can. An audit therefore detects a successful backdoor under that assumption; failed backdoor attempts may go undetected. AI Control, §2 and §2.1, p. 4, footnote 4; §2.2, p. 5 Actual human auditing and broader deployment remain outside this experiment. AI Control, §5.1.1, p. 16
The strength of such a result depends partly on the challenge presented. An attacker adapting to a known defense tests a different condition from a fixed suite chosen before the defense was built. Knowledge, access, retries, and adaptation budget therefore belong to the claim. An unspecified adversary makes it difficult to know whether a test was demanding, representative, or simply different from the situation being discussed.
We also need to establish what happened: an attempted action, its execution, and the specified loss are different outcomes. A reassuring transcript cannot establish downstream state it never observes. Independent state checks can improve the observation, but their coverage still limits the conclusion.
Zero observed incidents means that no incident was observed within the measurement and exposure represented. Estimating a risk from that result requires assumptions about selection, dependence, coverage, and detection. Benign work may provide substantial evidence of useful performance while offering little exposure to a particular attack. Concentrated successful attacks on selected tasks establish possible failures without supplying a population prevalence estimate.
Compare against a serious alternative#
The strongest alternative to a more elaborate control design is a competent model operating through a simpler interface with fewer opportunities for consequential error. That design might meet the declared requirements at lower total cost. The argument in this paper permits it to win.
In a prospective comparison, I would hold the entrusted work, protected interests, and test population fixed while varying a named control or design choice. It would include a credible behavioral-safeguard baseline and, where applicable, a simpler interface. Each condition would receive a stated and comparable opportunity for adversarial adaptation. The evaluation would preserve failures and abandoned attempts, rather than selecting only workflows in which the proposed control succeeded.
Legitimate completion, unauthorized outcomes and their severity, delay, refusals, review labor, and displaced work would be measured separately. The useful-performance floor, cost limits, meaningful improvement, and precision needed for a conclusion would be declared in advance. Rare severe outcomes may require more evidence than common task successes; a blended score should not hide them.
This is a proposed study, not an experiment performed for this paper. The evidence reviewed here does not establish the general superiority of external controls over behavioral safeguards at matched utility and cost. A repeated finding of little marginal benefit would narrow the proposed harm-reduction hypothesis. A control that shifts necessary work into a less observable route could lose its apparent advantage once that work is counted.
When a comparison is infeasible for a rare, severe loss, the comparative claim remains unresolved. Narrow prevention or limit claims and scoped adversarial tests can support parts of a reliance decision. Precursors need a justified connection to the loss; transferred or pooled observations need comparable conditions and measurement, with dependence accounted for. These inputs must meet the requirements of the claim they support. Remaining gaps require narrowing or deferral.
A result showing that a specified constraint prevents an otherwise reachable loss at acceptable cost is useful, even when other losses remain open. It can justify reliance within that scope, with the remaining dependencies explicit.
Carry evidence into operation#
I preserve four scopes of inference from AI Vision & Future: demonstrated capability, repeatable evaluation, bounded deployment, and scaled operation. Each asks for evidence at the level claimed; each also depends on the quality of the study within that level. A deployment report can be weakly measured, and a controlled experiment can be precise about a narrow mechanism. Framing
The claim concerns a particular configuration, so changes matter. A new model version, tool, retained-state policy, identity process, workload, or delegation pattern can alter a path we rely on. I would examine what the change could invalidate, which evidence needs renewal, and which results still apply.
A reader should be able to follow the promised outcome to the mechanism supporting it, its assumptions, and the evidence that could change the conclusion. Keeping that argument current as the service changes is a governance responsibility.