Every AI marketing tool sells the same promise: it does the work for you. The promise is not wrong, but it is measured on the wrong axis. For anything that spends a budget or emails a stranger, the interesting question is not how much the agent can do unsupervised. It is which actions it can never take unsupervised, and whether that boundary is enforced in code or merely described in a policy.
Sorting actions by consequence, not by confidence
Model confidence is a poor gate. A model can be confident and wrong, and the cost of being wrong is wildly uneven: pulling a keyword report and launching a campaign to two thousand strangers are not the same kind of action, however sure the model is about either.
A more useful sort is by consequence. Reading and reporting is reversible and cheap. Tuning a bid inside a cap you already set is low-consequence and reversible. Launching a campaign, publishing an article or sending a first email to a new list changes something in the world that cannot be pulled back. Raising a spend cap, connecting an account, or buying something is in a category of its own.
Those tiers earn different treatment: automatic and logged, batched into one daily approval, approved individually with the reasoning shown, and never automatic under any circumstances.
Why the top tier does not graduate
It is tempting to let a well-performing agent earn its way up. A hundred good decisions, then it gets the keys. This works for reversible actions and fails badly for irreversible ones, because the record of good behaviour tells you nothing about the cost of the first bad one.
Raising a spend cap is the clearest case. An agent that can raise its own cap does not have a cap. Whatever number you set is a suggestion, and the guarantee you thought you had is gone. So that action stays with the human permanently, regardless of track record.
The one exception worth making
There is a case for acting without asking: when the protective action is strictly reductive. If spend is running well past plan at two in the morning, pausing is the safe move and waiting for the morning approval is the expensive one.
The constraint that makes this safe is directional. The agent may pause or reduce. It may never start, send, publish or increase. A guard that can only ever make things smaller cannot be turned into an attack on your budget by a bad inference, and every such action still surfaces in the next brief with its reasoning attached.
What this costs
Approval is friction, and friction has to be paid for in attention. That is a real cost, and the answer is not to remove the approval but to make it fast: batch it into one moment a day, group the low-consequence items so they clear in one tap, and reserve individual decisions for things that genuinely deserve one.