The Control Layer
How many have crashed their boats on the rocks having heard the siren’s song of agentic software factories:
Come, leave your careful plans behind.
No need to chart what I can find.
Give me the goal, release the helm.
I’ll rule the paths within your realm.credit: GPT5.5 Instant
The overwhelming majority of software factories being built today follow the pattern of putting the agent in control of the delivery loop. Implementations may differ, but the common practice is to give the agent a goal, a working environment, tools, and instructions and let the agent’s model decide how to proceed.
The architecture makes sense when the path is uncertain. The agent can inspect the situation, choose tools, delegate work, and adapt as it goes. Similar territory where curiosity is useful.
But, a software factory introduces a different question. Once the delivery process is understood, how much of that control should the agent own?
Context gathering, planning, implementation, validation, review, repair, acceptance, and hand-off are not new decisions for each feature. In most engineering teams that process is already known. There is still uncertainty and judgement required at each stage but not around the order of them.
That is the difference between agent-controlled orchestration and workflow-controlled orchestration.
In the first, the model interprets the process at runtime and decides how the work moves. In the second, the process lives in the control layer. The agent is invoked where judgment is required, while sequencing, gates, fan-out, retries, and completion are enforced outside the model.
Where control lives
Both orchestration approaches require the delivery process to be defined. Context must be gathered from relevant sources, a plan created and reviewed, code written, tests run, and so on until acceptance.
In an agent-controlled system, much of that process is encoded as instructions. Skills are used to communicate the process which the model has to interpret every time the agent runs so it knows the order of operations and what’s expected. For example, a skill might tell the agent to run tests after implementation, launch several reviewers, wait for them to finish, consolidate the findings and add them to the backlog.
A workflow-controlled system can encode the same process directly. Build is followed by validation because the workflow requires it. Reviews run in parallel because the workflow defines that fan-out. The next stage waits until the required reviews complete because the dependency is part of the graph. The process design is largely the same in both cases. What changes is where the control logic executes.
The autonomy tax
Moving known process control into the agent introduces an autonomy tax because it adds extra cognition, latency, variability and failure around the model when it reasons about how to execute a process that is already understood.
The model has to answer questions about what should happen next and if it should proceed based on its understanding of the process. This reasoning happens each turn during the agentic loop which can be helpful during peak uncertainty, but less useful when the answer is already known.
Unsurprisingly, research on agentic workflows repeatedly shows that additional orchestration, self-reflection, delegation, and control can increase token use, model calls, latency, and failure opportunities without producing a step-size improvement in solution quality.
Context gathering is a good example of where the autonomy tax appears.
Below you see both controllers receiving the same input and tasked with producing the same build-ready context output. A workflow-controlled system can declare the relevant sources in advance and how to retrieve them. A model is still used to interpret or compile the context, but it does not need to scan the entire code base or decide how the retrieval process itself should work.
In an agent-controlled system, those same requirements are usually expressed as instructions. The model interprets the skill, decides where to look, decides whether it is useful, decides whether additional context is needed, and eventually decides when enough has been gathered.
That’s a lot of decision points.
Wang et al. describes about 6× more tokens at formulation under agent control, plus an order of magnitude more model calls. The 12k/72k totals and stage splits apply that ratio to this software-factory example.
The added cost sits in Plan and Gather. The agent-controller interprets the skill, selects a source, makes sense of the result, then decides whether to keep searching. The workflow-controller spends no tokens choosing those steps because that logic lives in code.
Explicit stages make failures easier to see
Workflow control does not guarantee the output will be correct. A retrieval step can still return the wrong file or follow an out-dated process, and the model it uses can still misunderstand what it reads. But, attributing a failure back to a workflow step is easier because inputs and outputs can be inspected independently.
When the entire process is buried inside one long agent session failures can collapse into the same outcome, which is…

the agent got it wrong
Tests are not suggestions
Ever broken the build on an engineering team? Chances are you likely paid for it by buying everyone a round that evening. CI/CD makes this much harder to do now, but the point is validation is not optional. From linting to unit/integration tests, there is little value in shifting that procedural control to an agent when the gates are already objective. Even with explicit instructions the agent can decide to skip tests it doesn’t think are required.
Good luck getting the agent to buy you a beer.
Review is orchestration too
Assume every completed build must be evaluated for security, code quality, specification compliance, performance, and test coverage. An agent-controller can be instructed to launch five reviewers, but that puts the onus of interpreting the instruction, spawning the reviewer sub-agents, tracking their progress, collecting the responses, etc. back on the agent.
That’s a large orchestration overhead for little value. The time, tokens, and cognition should be spent on the reviews themselves vs. coordinating a known process.
Reviewers still share blind spots, including across providers. Different models help less than this example suggests. Assigning distinct jobs, such as security, performance or API design, also spreads coverage.
Known work belongs in the workflow
The agent does the work that needs judgment. The workflow owns the known process around it. As more of that process becomes well understood, there is less value in having the agent control it.
References
- Cemri, M., Pan, M. Z., Yang, S., et al. (2025). Why do multi-agent LLM systems fail? arXiv:2503.13657
- Gorinova, M. I., et al. (2026). Position: Coding benchmarks are misaligned with agentic software engineering. SE 3.0 Workshop, KDD 2026. arXiv:2606.17799
- Huang, J., et al. (2024). Large language models cannot self-correct reasoning yet. ICLR 2024. arXiv:2310.01798
- Kapoor, S., et al. (2024). AI agents that matter. arXiv:2407.01502
- Kim, E., et al. (2025). Correlated errors in large language models. ICML 2025. arXiv:2506.07962
- Kim, Y., et al. (2025). Towards a science of scaling agent systems. arXiv:2512.08296
- Li, W., et al. (2026). Rethinking mixture-of-agents: Is mixing different large language models beneficial? TMLR. arXiv:2502.00674
- Panickssery, A., et al. (2024). LLM evaluators recognize and favor their own generations. NeurIPS 2024. arXiv:2404.13076
- Wang, J., Liu, S., Fu, G. & Savic, D. (2027). Where should control sit? Reliability-cost trade-offs in delegating water distribution network optimisation to LLM agents. Water Research 308, 126823. doi:10.1016/j.watres.2026.126823
- Xia, C. S., et al. (2024). Agentless: Demystifying LLM-based software engineering agents. arXiv:2407.01489
- Yang, J., et al. (2024). SWE-agent: Agent-computer interfaces enable automated software engineering. NeurIPS 2024. arXiv:2405.15793
- Yang, Y.-T. & Zhu, Q. (2026). Toward reliable design of LLM-enabled agentic workflows: Optimizing latency-reliability-cost tradeoffs. arXiv:2605.23929
- Zhang, J., et al. (2025). AFlow: Automating agentic workflow generation. ICLR 2025. arXiv:2410.10762