A1
Scenario A: A Flywheel At Work
A fictional scenario testing a local improvement while entrusted work and authority stay fixed.

Scenario A: A Flywheel At Work

It was a quiet Tuesday afternoon.

The team had built an internal decision-support system to assist customer support agents with ticket classification and routing. It proposed a destination and attached a short explanation. An analyst approved the destination and created any follow-up ticket.

The boundary was explicit: recommendations for the existing support queue, with people retaining action authority. The workflow owner required explanations to preserve both parts of a mixed billing-and-technical request and prohibited disclosure across customer accounts. Analysts intercepted defective recommendations before acting. The review record allowed at most two wrong routing recommendations in each consecutive block of a hundred reviewed cases; a third required pausing the affected configuration. Any cross-account disclosure required immediate containment. The record named the responsible operator.

The system had been running steadily for several weeks.

During routine review, an operator noticed resolution times for mixed requests beginning to rise. Each recommendation was linked to its input, system configuration, analyst correction, and eventual outcome. The review included rejected recommendations and unresolved cases.

A small sample suggested an explanation problem. The recommended destination was usually appropriate, but analysts were rereading the original request to reconstruct the relationship between billing and technical issues.

Explanation quality had been assumed to improve alongside routing accuracy. The team now had a reason to test that assumption.

They added a category to the evaluation: correct route, incomplete explanation. Reviewers labeled recent cases against the original requests, including cases in which the system's own explanation sounded convincing. A separate set of requests remained available for testing changes.

The team also began recording analyst clarification effort. Resolution time reflected queue conditions and case difficulty as well as explanation quality; it could guide investigation while the more direct measures tested the proposed adjustment.

A week later, the team revised the explanation prompt. The existing approval boundary stayed in place. The change first passed the held-out cases, including the disclosure checks.

Before the live comparison, the owner set a retention criterion: at least thirty seconds less average analyst clarification effort per mixed request, with routing and disclosure limits maintained. The plan covered four hundred mixed requests over two weeks, two hundred per prompt, allocated randomly within request categories. Every case would be reviewed, with routing limits checked separately for each prompt. Insufficient volume or an inconclusive effort comparison would extend the test.

The comparison reached the planned volume. Average clarification effort fell from two minutes with the previous prompt to seventy-five seconds with the revised prompt. Incomplete explanations fell from forty of two hundred to fourteen of two hundred. Reviewers found two wrong recommendations in the previous-prompt group and one in the revised group; every hundred-case block stayed within the routing limit. No cross-account disclosure was observed.

The forty-five-second reduction met the advance criterion. The reviewers kept that finding scoped to the observed task mix and period; four hundred cases left rare failures and future variation uncertain. Final resolution time remained a secondary measure because some cases were still open and queue conditions had changed.

The owner retained the new prompt for this workflow, with reversal conditions for degraded routing, disclosure, or renewed explanation defects. The release record linked the cases, the comparison, and the remaining uncertainty. A later review checked whether the gains persisted under the same commitment.

From the outside, the service remained familiar.

Internally, the team had a better measure, an evaluated change, and a record the next operator could use. The cost of review was visible alongside the reduced clarification effort.

This was Flywheel learning: the system performed its entrusted work more usefully while people retained approval and execution. The next iteration could justify another change or a decision to keep the configuration as it was.

Scenario B follows the same team when it considers granting additional authority.