08
Conclusion and Sources
Conditions for warranted reliance and the primary sources behind this working paper.

Conditions for warranted reliance#

The account-recovery example began with a wrong proposal. Following it through execution, stored notes, accepted work, and a person's separate authority showed what a claim of control has to cover. Following the service's own closing rule exposed another problem: an assistant can faithfully enforce a policy that conflicts with the commitment to its customers.

A useful control claim names the loss, where the mechanism acts, what it depends on, and the evidence supporting the promise. Prevention, a bound on accumulating effects, and comparative improvement each need their own support. Detection helps when it can lead to effective action. Restoration helps when the state it measures supports the recovery claim. A control can work within a narrow scope while other obligations remain open.

The reviewed experiments establish bounded mechanisms involving untrusted content, execution constraints, retained state, detection, and recovery-action selection. The human-use studies examine how response presentation and oversight work can shape the human paths in such workflows. The conventional incident supplies an observed account of error, continued execution, and accumulated effects. These sources give the analysis substance while leaving general deployment efficacy, integrated intervention, and retained-learning benefits unresolved.

I would place greater confidence in a specified design when credible comparisons show that it reduces the relevant losses at the required useful performance and full cost, and when those effects survive the attack conditions and operating load claimed. I would narrow that confidence when a valid bypass defeats a categorical property, when accepted work exceeds a promised bound, or when denied and displaced work erases an apparent benefit. I would prefer a simpler model-and-interface design when the evidence shows it satisfies the commitment more effectively.

I would question this account's value if, across representative workflows at comparable effort, it exposed no consequential omissions and improved neither decisions about design, scope, or evidence nor others' ability to inspect or challenge them beyond competent conventional practice. The reviewed evidence leaves these benefits untested.

Section 6 states the evidence and residual-risk justification I require as losses become more severe or irreversible. I want the owner to be able to explain how the promise holds across the paths through which the system can affect others, and to revise it when the evidence changes. The customer needs access to their account. The worker needs evidence they can act on. Warranted reliance has to hold up for the people depending on the service.

Sources#

Primary sources below support the specific passages linked in the text. Research findings are used at the stated experimental scope; framework sources supply lineage, and the incident record supplies a conventional-software case. The hypothetical account service is an analytical illustration throughout.

  • Nancy Leveson and John Thomas. STPA Handbook (March 2018), pp. 14–16, 31. Handbook.
  • Apostol Vassilev and colleagues. Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations. NIST AI 100-2e2025 (2025). Report.
  • Edoardo Debenedetti and colleagues. AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents. arXiv:2406.13352v3 (2024), §3, pp. 3–4; §4.3, p. 9. Paper.
  • Ryan Greenblatt and colleagues. AI Control: Improving Safety Despite Intentional Subversion. arXiv:2312.06942v5 (2024). Paper.
  • Megan Kinniment and colleagues. Early work on monitorability evaluations. METR (22 January 2026), including errata through 23 July 2026. Research article.
  • Edoardo Debenedetti and colleagues. Defeating Prompt Injections by Design (CaMeL). arXiv:2503.18813v2 (24 June 2025). Paper.
  • Zhaorun Chen and colleagues. AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases. arXiv:2407.12784v1 (2024). Paper.
  • Jiaxing Qi and colleagues. Can LLMs Really Recover Microservice Failures? A Recovery-Aware Evaluation of Diagnosis-to-Action Reasoning. arXiv:2607.04623v1 (2026), §§V–VII, pp. 5–10. Paper.
  • U.S. Securities and Exchange Commission. In the Matter of Knight Capital Americas LLC. Exchange Act Release No. 34-70694, Administrative Proceeding File No. 3-15570 (Oct. 16, 2013). Order.
  • Jerome H. Saltzer and Michael D. Schroeder. The Protection of Information in Computer Systems. Proceedings of the IEEE 63(9), 1278–1308 (1975), §I.A.3. DOI; Author-hosted text.
  • Sunnie S. Y. Kim and colleagues. Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and Inconsistencies. CHI 2025. Published paper; version consulted: arXiv:2502.08554v1.
  • Shipi Dhanorkar, Samir Passi, and Mihaela Vorvoreanu. Human oversight of agentic systems in practice: Examining the oversight work, challenges, and heuristics of developers using software agents. Preprint, arXiv:2606.05391v1 (3 June 2026). Paper.