The more capable the agent, the more important its boundary.
Most agents are not following a complete procedure you defined. Under human supervision, they inspect the environment and assemble that procedure while they work: selecting tools, inferring commands, retrying failures, choosing deployment steps, and deciding what counts as success.
That improvisation is easy to miss when the result looks correct. But as teams delegate larger changes, it makes execution harder to predict, review, reproduce, and trust. It also creates the exploration, failed attempts, elapsed time, and token cost that accumulate behind every successful run.