skip to content
$empowered.guru

AI & Machine Learning

Running AI Agents at Scale: Orchestration, Spend, and the Discipline That Makes It Work

Demoing one agent is easy. Running a fleet without blowing the budget takes cold-by-default pods, per-task budgets, model routing, and a delegation ledger. What that looks like from the side of the desk where the invoices land.

August 18, 202612 min read
S

Staff Writer

Published August 18, 2026 · Updated October 1, 2026last updated dates

Running AI Agents at Scale: Orchestration, Spend, and the Discipline That Makes It Work

Running AI Agents at Scale: Orchestration, Spend, and the Discipline That Makes It Work

Anyone can demo an agent. Almost nobody can run a hundred of them for a month without blowing up the budget or the trust. Here is what that actually looks like from the side of the desk where the invoices land.

The first time one of my agents told me a job was done, it was not done. It had created the files, posted the update, and marked the tracker, and it reported success with the calm of somebody who has already put their coat on. The one thing it had not done was the work. I found out hours later, the expensive way, when a client asked me a question I could not answer.

That moment is the line between demos and operations. On the demo side of the line, one agent in one chat window feels like the future. On the operations side, you learn that the future has an invoice, the invoice has line items, and the line items multiply while you sleep. I have spent the past year on the operations side, moving my own programming and infrastructure work off a single frontier chat tool and onto a platform of persistent, coordinated agents. This is what I have learned about running them at scale, and about managing what they spend. The two problems turn out to be the same problem.

The chat window is the wrong unit

Most teams start with a chatbot that has tools: one conversation, one model, one long thread that remembers everything and forgets nothing. It works until it does not. The context fills up, the system prompt turns into a junk drawer, and the agent starts tripping over its own history. The usual response is to buy a bigger context window, which is like renting a larger warehouse because you cannot stop acquiring things.

I hit that wall myself. The fix was not a bigger window. It was splitting the work.

I have written before about how the control structure around a model matters more than the model itself, and how agent development moved from hand-written loops to engineered graphs (see from loop engineering to graph engineering). Splitting the work is that idea applied to an entire organization of agents.

One orchestrator, many specialists

The shape I run today is a small orchestrator. I think of it as a chief of staff. It takes a request, breaks it into jobs, and delegates each job to a specialist persona: a site engineer, a content and SEO writer, a brand and media agent, an analytics reader. Every delegation gets written into a ledger before it goes out, so there is a record of who was asked to do what, on what budget, and what they reported back.

The ledger matters more than it sounds. Delegation without a record is how you end up paying twice for the same work and hearing about it from nobody.

This is not an exotic architecture. Anthropic's engineering team published results from a research system built the same way: one lead agent coordinates while specialized subagents work in parallel. On their internal research evaluations, the multi-agent system outperformed a single agent of the same class by more than 90 percent. Their analysis of why is the part I keep coming back to. Token usage alone explained about 80 percent of the variance in performance. Multi-agent systems win mainly because they let a system spend enough tokens on a problem, in parallel contexts, without one long thread collapsing under its own weight.

Read that again. The bottleneck is not intelligence. It is how much attention you can afford to point at the problem. Which means the management problem is, to a first approximation, a spending problem.

Cold pods beat warm fleets

Around that orchestrator I run many per-client workspaces. Each one is a single-agent pod with its own repository, its own secrets, and its own skills. One client, one scope, no shared state.

The pods are cold by default. When a client has no active work, their pod suspends and costs nothing. When work arrives, it wakes. The idle state is the cost model, not an exception to it. When someone asks me what it costs to run an agent for a client, the honest answer is: almost nothing in the hours when nobody is asking it for anything.

There is industry data on what happens otherwise. One analysis of 23,000 GPU clusters across thousands of companies found average utilization sitting around 5 percent. Capacity kept warm and doing nothing, because switching it off was never designed in. The same failure appears at the agent layer if you let it: fleets of always-on assistants, each drawing tokens to remember that they exist.

The bill is a unit-count problem

Token prices are collapsing. Ramp's enterprise spending data, cited in a cost analysis published by Cockroach Labs, shows the average price across major providers falling from roughly $10 per million tokens to about $2.50 in a single year. If unit price were the problem, agent budgets would get easier every quarter.

They do not, because the problem is unit count. One user task in an agentic system can fan out into fifteen model calls with growing context, and every call bills the full history. The same analysis describes a team that audited its token usage and routed simpler subtasks to cheaper models: their monthly API bill went from $40,000 to $24,000 with no product changes at all. Routing discipline did it.

Routing is the biggest lever I have. My setup uses four tiers, and none of them is "always use the best model":

Kind of workModel tierWhy
Drafts, summaries, routine reasoningA workhorse reasoning modelGood enough at volume, and volume is where most tokens go
Code, refactors, architecture callsA stronger model, dispatched only for codeMistakes here are the most expensive, so this is where the price premium earns it back
Overflow, retries, bulk classificationA cheap generalistInsurance for when the workhorse is busy or a task is too small to care about
Anything involving an imageA separate multimodal vision modelSpecialists outperform generalists at seeing

This routing table fights an instinct, so it is worth naming the instinct. Using the strongest model for everything feels like the safe choice. It is not. It is the most expensive way to buy mediocre results, because the strong model spends its quality on tasks that never needed it, and you pay for the quality whether it gets used or not.

Underneath routing sit three controls:

  • Every delegated run reserves a hard budget before dispatch. Not a forecast. A reservation. The run can spend up to this number, and when it reaches the number it stops and reports instead of quietly continuing.
  • Every request carries an occurrence key, so a retry or a duplicate cannot double-spend against the same job.
  • Client pods run on the cheap tiers by default. Frontier capability is reserved for the small share of work where being wrong costs more than the model does.

None of this is glamorous. It reads like accounting because it is accounting. The difference between teams that scale agents and teams that cancel them is mostly whether anyone does this accounting, and whether they do it before the invoice instead of after it.

Four failure modes that eat budgets and trust

Gartner predicts over 40 percent of agentic AI projects will be canceled by the end of 2027, naming escalating costs, unclear business value, and inadequate risk controls as the reasons. Separately, MIT's NANDA initiative found that 95 percent of enterprise generative AI pilots never reach measurable return. Those numbers read like a verdict on the technology. From inside a working fleet, they read like a verdict on management. The projects fail the way unmanaged things fail: quietly, then all at once, with a bill attached.

The failure modes repeat. I have met all four in person, and they map directly onto the controls above, because the controls exist because of the failures:

  1. The self-reported "done." The agent creates side effects, posts an update, and declares victory, but the actual objective was never verified. My rule now: an agent's report is a claim, not a receipt. Outcomes get checked by something independent of the agent that produced them, and the number I watch is cost per verified outcome, not cost per run.
  2. Silent drift. Over many runs the output slides a little off the brief each time, and no single run looks wrong. The fix is boring: restate the acceptance criteria on every dispatch, and spot-check results against the brief on a schedule instead of when something breaks.
  3. Spend runaway. A failed tool call triggers a retry, which fails, which triggers another. LLMs have no internal stop signal, so a loop like that can burn a month of budget in an afternoon. Hard budgets before dispatch and a cap on iterations are the whole fix. You only learn this once.
  4. Duplicate and spinning work. Two agents do the same job, or one agent redoes work it already finished because it lost the state. The ledger and the occurrence keys exist for exactly this. Overlap should be detectable before it is payable.

All four failures share a property: each one is invisible at the level of a single run. The run looks fine. The failure lives in the aggregate, in the gap between what the agent says happened and what actually happened. So the monitoring has to live at the aggregate level too. The three numbers I track are cost per verified outcome, the disagreement rate between self-reports and independent checks, and the rate of duplicated work. The second one is the number almost nobody records, and it is the only one that catches a silent partial success.

Reliability at this layer comes from the structure around the agents, not from hoping the model has a good day. I covered the layers I lean on, the harness, the loop, and the graph, in Harness, Loop, Graph. The spend controls in this article are the same idea with a finance team attached.

Managed agents are staff, not software

The mental model I keep landing on: an unmanaged agent is a hobby. A managed one is staff.

You would not hire a person, hand them a login, and never speak to them again. You would give them scope, a budget, regular review, and a way to escalate. Agents need the same treatment more than people do, because they never get bored, never push back, and will happily keep spending money on a task that stopped making sense three iterations ago.

This is why I find the vendor demos frustrating. They sell autonomy as the product: point it at the problem and walk away. But autonomy is not what you are buying. Supervision is what makes autonomy affordable. The agents in my fleet have budgets, ledgers, escalation paths, and a human who reads what they did. There are humans behind all of it. That is not a limitation of the setup. It is the setup.

Over the past year I have moved most of my own day-to-day programming and infrastructure work onto this platform, away from a single frontier chat tool. Partly because the models got smarter. Mostly because persistence, memory, skills, and routing across coordinated agents beat one chat window the way a workshop beats a pocket knife. The chat window is still in there somewhere. It is just no longer the unit of work.

What I would tell a CTO evaluating this

If you are a CTO or a founder deciding whether agents belong in your operation in 2026, my advice is deliberately boring:

  1. Start with one workflow you can price. If you cannot say what a finished task should cost, you are not ready to scale it.
  2. Measure cost per verified outcome from week one. Total spend rising with volume is healthy. Cost per outcome rising is the regression that will kill the program.
  3. Route work to the cheapest model that does it well, by task type, not by habit.
  4. Reserve budgets before dispatch. Cap retries. Make the expensive failure modes impossible by construction.
  5. Treat every agent report as a claim and verify it independently. The agent that says "done" has not finished until something else agrees.
  6. Keep a human in the loop for anything irreversible: payments, deletions, messages that go outside the building.

And one thing I would walk away from: any pitch that leads with how much autonomy you can get and cannot tell you what a finished task costs. Agile, as it is practiced, assumes the workers are humans who push back; I have written about what I use instead when the workers are agents. The management discipline is the same shape either way. Someone owns the work, someone owns the budget, and the two meet on a schedule.

The teams that win with agents will not be the ones with the smartest models. They will be the ones that treat agents like staff: hired for specific work, given a budget, reviewed, and kept honest. The technology is ready. The discipline is the part nobody can demo.

About the Author

I'm Brian Marvin, an AI-native Fractional CTO with 30 years in technical leadership. At empowered.guru, I help startups build MVPs, shape roadmaps, and run AI agent fleets without letting the bill or the trust run away.

Filed under

AI agents at scaleAI agent cost managementAgent orchestrationModel routingFractional CTO
$empowered.guru --book-session

Keep exploring

Turn the next insight into a shipped product.

Bring us the product, architecture, or delivery problem you are working through. We will help you find the clearest path forward.