04 · Agentic
Overview#
When I say “agentic,” I mean a system property: the system selects and executes actions over time in pursuit of an objective, using tools under constraints.
Conventional software can adapt to changing inputs and select actions dynamically. The practical question here is which choices a workflow delegates to a model at runtime: actions, tool selection, sequencing, and responses to intermediate results.
- In a single-turn setting, a mistake often appears as an output that can be rejected.
- In an agentic setting, a mistake can become the next step's context, alter an external system, trigger other actions, or cross a trust boundary before review.
Agentic systems connect the Flywheel and Helix through delegated action. They can compress the interval between observation, decision, and action; as actions become faster and cheaper to attempt, the ability to observe, evaluate, and intervene can become the binding constraint.
Required distinctions#
These distinctions make it possible to test what we mean:
- Automation vs autonomy: automation describes work carried out by software; autonomy describes the scope of action selection delegated to it.
- Tool invocation vs goal-directed behavior: invocation calls a tool as a substep; goal-directed behavior chooses tools, order, and state changes to satisfy an objective under constraints.
- Stateless vs stateful execution: stateless runs are independent; stateful systems retain context or change an environment in ways that affect future choices.
- Bounded vs open-ended agency: bounded agency has a constrained action space, explicit limits, and observable outcomes; open-ended agency combines broad permissions or horizons with weakly specified success and limited control.
I focus on bounded agency because reliability and governance can plausibly scale in that setting.
Agency and general intelligence are separate claims. A narrow system can select actions over time. It can also appear capable while remaining unsuitable for deployment at the required reliability.
The agency envelope#
I treat agency as a multidimensional envelope: what we let the system do, how far it can proceed, and how we keep that work under control.
- Action space and permissions: which tools, data, systems, and side effects are available, and under whose identity.
- Horizon: how long the system can operate, how many decisions it can make, and when it must stop or request review.
- State: what the system can read, retain, retrieve, or change across steps and runs.
- Delegation and concurrency: whether it can delegate subtasks or pursue multiple action paths at once.
- Reversibility and containment: which actions can be rolled back, isolated, retried safely, or prevented from propagating.
- Oversight capacity: whether relevant behavior is legible and can be reviewed or interrupted on the timescale the workflow requires.
The operating boundary states the commitment: entrusted work and authority under acceptance criteria and error limits. The agency envelope describes how we implement that commitment through delegation and controls. Controls can change while the commitment remains fixed.
These dimensions can vary separately, and their effects interact. A longer horizon may be acceptable with narrower permissions. Broader tool access may require approvals, stronger containment, or a smaller action budget. Each change should identify its effect on entrusted work and control.
Enabling conditions#
Technical#
- Stable tool boundaries: predictable schemas, explicit failure modes, and idempotent operations where possible.
- Constrained identity and permissions: action scope deliberately narrower than the environment's full capability.
- Bounded state: defined purpose, provenance, access, retention, and expiry for memory and retained context.
- Containment and recovery: gates for high-consequence action and a way to isolate, reverse, or escalate partial failure.
- Monitorability: records meet the Section 03 operational definition for reconstructing relevant behavior and supporting timely intervention.
- Multi-step evaluation: tests of trajectories, tool use, constraint adherence, intermediate state, and final outcomes.
Organizational#
- Clear accountability: someone owns workflow correctness, security posture, intervention, and recovery.
- Defined error budgets: acceptable failure frequency and severity are stated for the delegated scope.
- Intervention rules: approval, escalation, shutdown, and rollback conditions exist before an incident.
- A change process: observed failures can change evaluations, tools, permissions, prompts, models, or human process.
Constraints and limits#
Agentic systems break down or become uneconomical under common conditions:
- Weak success criteria: reliable progress depends on measurable outcomes; otherwise activity can masquerade as progress.
- High-cost irreversible actions: if rollback is unavailable and mistakes are expensive, the envelope should contract or action should return to human approval.
- Tool boundary brittleness: changing schemas, permissions, or data shapes can invalidate apparently reliable behavior.
- Sparse or delayed feedback: when outcomes arrive late or remain ambiguous, evaluation and recovery lag behind action.
- Monitoring saturation: action volume, horizon, concurrency, or opacity can exceed the capacity to review and intervene.
- Hidden supervision: untracked human correction can make a system appear more autonomous and reliable than it is.
These limits often appear as an integration ceiling: the model can attempt more work than the surrounding system can govern at the required cost and reliability.
Evaluate reliability against that whole envelope. Completing the task supplies only part of the evidence; the path taken and the consequences of its actions matter too.
Failure modes (common)#
- Compounding error: later steps remain locally consistent with an incorrect earlier state.
- Hidden state: memory or environment state affects action in ways that cannot be reconstructed or reviewed.
- Brittle tool boundaries: ambiguous schemas, partial failures, or non-idempotent operations make safe retry difficult.
- Unobserved delegation: subtasks move to tools or other systems without preserving what was done, why, and with what evidence.
- Control lag: the system acts faster than monitors, reviewers, or incident processes can respond.
- Human-in-the-loop illusion: reliability depends on untracked human correction whose cost and judgment are absent from the system account.