skip to content
$empowered.guru

AI & Machine Learning

Deterministic vs Stochastic Agents: How to Run a Fleet That Behaves

Same skill, same model, three runs, three different outcomes. That's fine for drafting and fatal for reconciliation. How to split deterministic and stochastic work across an agent fleet.

August 20, 20267 min read
S

Staff Writer

Published August 20, 2026 · Updated October 1, 2026last updated dates

Deterministic vs Stochastic Agents: How to Run a Fleet That Behaves

Deterministic vs Stochastic Agents: How to Run a Fleet That Behaves

Give the same task to the same agent three times and you can get three different outcomes. That variance is a feature when you are drafting copy. It is a liability when you are closing the books. Running a fleet of agents means deciding, workflow by workflow, which mode you are in.

Picture an accounting office that runs on AI. The close process lives in a skill with four steps: gather the invoices, reconcile against the bank feed, flag exceptions, produce the close report. Clean and simple.

Now hand that skill to three different LLMs. You get three different closes. One model reconciles first and reads the rest of the steps through that lens. Another flags an exception the first one silently resolved. The third formats the report differently and drops a table the first one included.

Here is the part that surprises people: hand the same skill to the same model three times and you still get three different closes. Most of the time they are all fine. Sometimes one of them does something you did not expect, and because the process is stochastic, you cannot reproduce the failure on demand. That is the whole problem in one paragraph.

Why the variance does not go away

The first instinct is to turn the randomness down. Set the temperature to zero, get greedy decoding, same token every time. Except that is not actually true.

Even at temperature zero, modern inference is not bit-stable. Floating point math on GPUs is not associative: (a+b)+c and a+(b+c) can round differently. Inference engines optimize for throughput, so they switch parallelization strategies based on current batch load. The same request can take a different compute path depending on what else the system is doing at that moment. The numerical difference is tiny, right up until argmax turns it into a cliff: a difference of 0.000001 in two logits flips which token gets selected, and from there the whole generation branches. Anyone who has compared two "identical" runs at temperature zero has seen exactly this.

Stack two more layers on top. Different models make different decisions from the same written procedure, because each one was trained differently and carries different priors. And the same model changes behavior when the vendor updates it. Your perfectly tuned workflow from March can drift in April without you touching a thing.

None of this is a defect to be fixed in a later release. Sampling is how these models work. The engineering lesson is not "find the setting that makes it deterministic." It is "design the system so the parts that must be deterministic do not depend on a model being deterministic."

The cleanest split I know of

Anthropic's engineering team drew the line precisely in their building effective agents write-up. Workflows are systems where LLMs and tools are orchestrated through predefined code paths. Agents are systems where LLMs dynamically direct their own processes and tool usage, maintaining control over how they accomplish tasks.

Both are legitimate. The mistake is treating every job as agent work. A workflow runs the same way every run because the path is code; the model fills in the blanks but never chooses the route. An agent chooses its own route, which is exactly what you want for open-ended problems and exactly what you do not want for a bank reconciliation.

Gartner's projection that over 40% of agentic AI projects will be canceled by the end of 2027, with escalating costs and inadequate risk controls named as causes, is what happens at fleet scale when organizations skip this decision. The projects that get canceled are usually the ones where everything was stochastic, so nothing could be trusted, audited, or costed.

At fleet scale, variance compounds

One agent improvising is a review problem. Ten agents improvising is an operations problem.

I run a fleet of client-facing pods, and the rule for spinning one up is simple: it happens the same way every time. A pod, or a copy of a pod, comes out of a fixed, ordered procedure with gates that verify each step before the next one starts. Not because the agents are bad at provisioning. Because provisioning is the kind of work where "slightly different" means "broken in a way nobody notices for two weeks."

The math is not kind to stochastic fleets. If a single step has even a 5% chance of being skipped or misread, a forty-step provisioning run has roughly an 87% chance of at least one bad step. Every time. Per pod. Now multiply by every pod in the fleet and every run of every recurring workflow. Variance that looks harmless in a single demo becomes a certainty across a fleet.

This is also why reproducibility is a debugging requirement, not a nice-to-have. When a workflow fails once in a stochastic system, you cannot replay it, you cannot bisect it, you cannot even be sure what it did. When the path is deterministic code, a failure is a stack trace with a line number.

The design rule: deterministic skeleton, stochastic joints

I sort every workflow in the fleet into two buckets, and I am strict about it.

Deterministic, no debate: anything involving money, compliance, or infrastructure. Invoicing, reconciliation, payroll inputs, deploys, migrations, cron setup, pod provisioning, backups. If an auditor might ask how it happened, it happens the same way every time, in code, with the model filling slots rather than choosing routes.

Stochastic, deliberately: anything where I want options. Drafting, design, research, naming, messaging, exploratory analysis. Variance is the product here. Three different drafts of a landing page is a feature, because I can react to them.

The part most teams miss is the joint between the two. When a stochastic step hands work to a deterministic one, the handoff must be structured: a schema, typed fields, validated output. Not prose. Prose is where stochastic leaks into deterministic, and it is where most fleets quietly rot. The drafting agent can be as creative as it likes; what crosses the boundary into the invoicing step is a validated object with an amount, a date, and a counterparty.

How to enforce it in practice

  1. Encode set-in-stone procedures as runbook-style skills: fixed step order, idempotent steps, verification gates that refuse to report done until each step checks out against reality. The model operates the procedure; it does not re-derive it.
  2. Pin model versions for deterministic workflows. If a workflow must behave the same in April as it did in March, it cannot silently ride a model update.
  3. Use structured outputs at every boundary between a stochastic step and a deterministic one. Validate against the schema before anything downstream runs.
  4. Keep a ledger of what actually ran, with results, so a fleet-wide process is inspectable after the fact. Determinism you cannot verify is just optimism.

I wrote about the first of these in detail in Your AI Agent Shouldn't Be Improvising Your Deploys. The short version: when we moved pod provisioning out of a prose checklist and into a gated runbook, the pods started coming out identical every time, and the class of failure where a step quietly got skipped disappeared.

The mindset shift is the real takeaway. Stochastic is not the enemy, and determinism is not something you beg the model for at temperature zero. Determinism is an architecture decision: you build it into the system around the model, and you spend the model's judgment where judgment is actually worth something.

About the Author

I'm Brian Marvin, an AI-native Fractional CTO with 30 years in technical leadership. At empowered.guru, I help startups build MVPs, shape roadmaps, and make AI-powered technology decisions that scale.

Filed under

AI AgentsAgent OpsDeterminismLLM WorkflowsSkills
$empowered.guru --book-session

Keep exploring

Turn the next insight into a shipped product.

Bring us the product, architecture, or delivery problem you are working through. We will help you find the clearest path forward.