Here is a test worth running on your current agent initiatives.
Take the most advanced one. Ask the team to write down the steps the agent performs in a typical successful run. Not the architecture — the steps.
If they can produce that list, and it does not change much between runs, you have not built an agent. You have built a workflow with a language model deciding what happens next, and you are paying for that decision on every execution.
What autonomy actually costs
The enthusiasm for agents treats autonomy as a capability being added. It is more accurate to treat it as determinism being removed, and determinism was doing several jobs.
Testability, first. A workflow with fixed steps can be tested. Given this input, that output, every time. Remove the fixed steps and you can no longer test the system — you can only sample it. That is a different discipline, most organizations do not have it, and it does not arrive with the platform.
Then cost predictability. A deterministic process has a known unit cost. An agent that decides its own path has a cost distribution, and the tail of that distribution is where an agent loops, re-reads context, and calls a tool eleven times because the tenth response was ambiguous.
Then debuggability. When a workflow fails you know which step. When an agent fails you have a transcript and a hypothesis, and reproducing the failure may not be possible.
And finally auditability. Regulated processes need an answer to why this decision was made. "The model chose this path" is a description, not an explanation, and it will not survive an examiner.
None of these makes autonomy wrong. They make it expensive, and the question is whether you are buying something with the money.
When you are
Autonomy earns its cost under one condition: the path genuinely cannot be enumerated in advance.
That happens when the input space is open-ended and the right next step depends on what was just discovered. Investigating an incident where each finding determines what to look at next. Research across sources you cannot list beforehand. Handling an inbound request that could be any of two hundred things. Diagnosis, broadly — where the process is genuinely a search rather than a sequence.
The common feature is that a decision tree would have to be impractically large, or would have to be rewritten constantly as the domain changes.
Most enterprise processes are not like this. They are sequences with branches, and the branches are known, because somebody wrote a procedure document for the humans who used to do it.
The uncomfortable middle
The strongest counter-argument deserves stating properly: many processes look enumerable but are not, because the documented procedure covers eighty percent of cases and the remaining twenty are handled by experienced people using judgment nobody wrote down. Automating only the documented part delivers much less value than expected.
That is true and it is the real case for agents in the enterprise. But notice what it implies about sequencing. The productive move is a deterministic workflow for the enumerable majority with an escalation path for the rest — not an autonomous agent for the whole thing on the grounds that part of it is hard.
Hybrid is unfashionable because it does not demo well. It is also what most of these systems converge on after their first year in production.
A rough decision
| If this is true | Build |
|---|---|
| You can write down the steps and they rarely change | A workflow. Use a model for the hard sub-tasks — extraction, classification, drafting — not for control flow. |
| Steps are known but the branching is wide | A workflow with model-based routing at the branch points. Deterministic skeleton, judgment at the joints. |
| The documented path covers most cases, judgment covers the rest | A workflow for the documented path, an escalation for the rest. Revisit once you have data on what the exceptions actually are. |
| The next step genuinely depends on what was just found, and the space is open | An agent. This is what autonomy is for. |
The cheaper route
The prediction is not that agents fail. It is that a large share of current agent projects will quietly become workflows, and the teams will describe this as a maturity milestone rather than as a correction.
That is a fine outcome and an expensive route to it. The cheaper route is to ask the question at the start: can we write down the steps? If yes, the autonomy is buying you flexibility you do not need, at a cost in testability, predictability and auditability that you will pay on every run for the life of the system.
Reach for autonomy where the problem is genuinely a search. Everywhere else, a workflow with good judgment at the branch points will be cheaper, more reliable, and considerably easier to explain to whoever eventually asks.
This is an argued position rather than a research summary. It cites no studies and names no platforms, and reasonable practitioners disagree with it.
Related Reading
Independent. No sponsorships. Unsubscribe anytime.