01 · Safety, security, and the losses that matter
I start with what we are trying to protect. In the account example, a customer can lose access, have private information disclosed, or have an account changed without authorization. The service also has to bear the cost of operating its protections. Those interests and costs belong in the same analysis.
I use safety to mean keeping risk of harm to people and other protected interests within explicitly justified limits. That includes harm arising from intended operation, error, misuse, and attack. I use security to mean protecting assets, information, and essential functions against access, manipulation, disclosure, disruption, or appropriation that the applicable authorization does not permit. Whether that authorization is itself justified is a question for safety analysis and governance.
Security controls can contribute directly to safety. Preventing an unauthorized recovery change can protect a customer's access and information. Safety analysis also examines the consequences of faithfully following the service's own rules. That broader starting point follows STPA's orientation toward stakeholder losses and the hazards, constraints, and control relationships that bear on them. STPA Handbook, pp. 14–16, Figure 2.2 I use STPA's orientation here, leaving its full vocabulary and procedure to the handbook. That lineage does not establish the effectiveness of the AI controls examined below; their effects need evidence from evaluated systems.
Correct execution can preserve a policy conflict#
In one variant of the hypothetical, an internal rule closes every case when the customer cannot receive a code at the old address. Assume the assistant follows the rule exactly. A legitimate customer who lost that channel then receives no staffed recovery route, contrary to the service's declared commitment.
Following the closing rule more faithfully would leave this customer locked out. The conflict is between the policy and the service's commitments. A solution must satisfy both: establish the customer's authorization and provide the promised recovery route.
Reliability concerns dependable performance against stated requirements under named conditions. Alignment, as used here, concerns the relationship between intended objectives and constraints and the behavior produced. Both require attention to what has been specified. A system can reliably implement an objective whose consequences remain unacceptable under the interests the service has undertaken to protect.
Now suppose the assistant routes these cases to a staffed vendor that requires government identification and keeps it indefinitely. Assume this meets both service commitments. The owner's criteria cover account access and unauthorized disclosure, but leave out the vendor's retention and permitted reuse of identification. The applicant may care deeply about that use even when every rule is consistent and faithfully followed. I would require that interest to enter the justification and challenge process. This example identifies an omitted interest; it supplies neither a legal verdict nor a general objection to identity verification.
That is why the owner's approval cannot settle which losses matter. I treat the justification of objectives and accepted risk as a governance responsibility: affected people, relevant obligations, and opportunities to challenge decisions belong in it. A service may have limited visibility into downstream harms. That limits what it can assure; it does not justify disregarding the interests it cannot see.
Count omissions and the cost of control#
Harm can follow from an action, an omission, or a delay. A difficult case that never reaches a worker can leave a customer without the promised recovery route, even if no account field changes. Preventing an immediate unauthorized action may still leave legitimate work that needs a way forward.
I count useful performance alongside harmful outcomes, including the people and tasks a control makes harder to serve. Review labor, false alarms, interruption, and work moved to another channel are possible costs to measure in the workflow.
Threat terminology helps organize part of that analysis. NIST's adversarial-machine-learning taxonomy distinguishes attacker objectives, access, and lifecycle surfaces, while explicitly excluding non-adversarial design and implementation failures from its scope. NIST AI 100-2e2025, Executive Summary, PDF pp. 12–13 Ordinary error therefore needs its own place in an account of this breadth.
With those interests and commitments clear, I can ask what we have delegated that could affect them. That sets the scope of the control claim.