01 · Framing
Overview#
I use a small set of operational definitions to limit category drift and keep this paper's claims tied to behavior under deployment constraints: latency, cost, data access, permissions, incentives, and accountability.
Three separations govern this paper:
- Model vs system: the model is the learned component; the system is the deployed loop around it.
- Capability vs reliability: capability describes what can be done under favorable conditions; reliability describes how dependably it can be done under operating conditions.
- Evidence vs inference: evidence supports the level observed; each further inference requires additional observations.
Working definitions#
These definitions are intentionally narrow. They serve analysis and system design; broader claims about intelligence require a different frame.
- Model: a learned component that maps inputs to outputs using parameters acquired through training, with behavior shaped by learned generalization.
- AI system: a deployed loop containing one or more models and the non-model components that shape outcomes: deterministic software, tools, data, state, infrastructure, interfaces, people, evaluation, permissions, and policy.
- Generative AI (GenAI): models or systems that produce new artifacts—such as text, code, images, audio, or structured data—conditioned on context. Quality is judged through correctness, utility, and constraints appropriate to the artifact.
- Agentic system: an AI system that selects and executes actions over time toward a specified objective, with some choice of action or control flow delegated at runtime. Runtime delegation defines agency here; recovery, memory, controls, and evaluation determine deployment viability.
- AGI-adjacent claim: a claim of broad task competence or transfer across domains with limited task-specific adaptation. Here, the relevant questions are breadth, task horizon, sensitivity to scaffolding, and the new supervision, data, or engineering required to sustain reliability.
“Agentic” is a procedural description of delegated action selection under constraints. Questions of intent, consciousness, and general intelligence require separate definitions. Section 04 develops the full agency envelope.
The evidence ladder#
I use four levels of operational evidence, each supporting a broader claim:
- Demonstrated capability: a model or system completes a task under constructed or favorable conditions.
- Repeatable evaluation: performance recurs for a defined configuration, test set, and set of conditions.
- Bounded deployment: performance persists in a live workflow with named tools, permissions, review, and error limits.
- Scaled operation: performance survives greater volume, duration, variation, and organizational use with costs and incidents visible.
I use this ladder to keep claims tied to observations. A successful trial gives us a reason to test more broadly; each transition needs additional evidence. The levels describe how far evidence reaches; study design determines confidence. Section 03 defines measurement validity and monitorability: what observations let us decide, and whether we can intervene in time.
Value at the observed scope#
Realized value means benefits remain after the full costs of integration, supervision, infrastructure, governance, and failure are counted. Assess it at the scope observed: a bounded deployment can yield value; a claim of value at scale requires evidence at scale.
Name the level at which value is assessed: task, person, organization, market, or society. Speed, outcome quality, distribution of benefits, and labor-market effects are distinct empirical questions within and across those levels.
Why the system boundary matters#
The same model can produce different outcomes inside different systems. An incorrect answer may be caught in review or acted on immediately. Reliability, security, cost, and accountability depend on the whole loop.
Controls can detect, contain, or route around model error. They can also amplify it, conceal it, or introduce failure through tools, state, permissions, interfaces, and incentives.
That is why the deployed system is the unit of analysis and operational accountability.
What would change this framing#
This framing should weaken or change if repeated evidence shows that:
- capability routinely becomes reliable in changing workflows with little additional evaluation, integration, or governance;
- controlled evaluations predict scaled outcomes across context shifts, including cost and severity-weighted failure;
- system controls and organizational process have little effect relative to raw model capability;
- or the distinctions among bounded deployment, scaled operation, and realized value cease to explain observed differences.
These distinctions guide my reading of benchmarks, reported use, and measured workflow outcomes in the reviewed evidence. I will revise them if they stop explaining what we observe.