skip to content
$empowered.guru

AI Agents

How I Test Tool Calling Before an Agent Touches Production

Most agent failures are not model failures. The agent picked the wrong tool or the right tool with the wrong argument. The test routine I run before anything ships.

September 30, 20269 min read
B

Brian Marvin

Published September 30, 2026 · Updated October 1, 2026last updated dates

How I Test Tool Calling Before an Agent Touches Production

How I Test Tool Calling Before an Agent Touches Production

Most agent failures I get called about are not model failures. The model picked the wrong tool, or the right tool with the wrong argument. Here is the test routine I run before anything ships.

I work as an AI-native fractional CTO, and I spend a lot of time cleaning up agents that worked in a demo and fell apart on real traffic. The pattern is almost always the same. In the demo, the agent called the right functions with tidy arguments. In production, it called something close but wrong, or sent an order id that looked right and was not. Nobody had tested the calling itself. They had only read the final answer and said looks good.

So I test tool calling as its own layer now, before I argue about prompts or models. This article is the routine. It is written for teams running agents with real tools, and it borrows nothing from anyone. It is just what held up for me across client builds.

For related thinking on the layers around this, I wrote more in harness loop graph reliable AI agents and how I make AI workflows hold still.

Agents fail in two places, so I test both

Every tool call has two decisions inside it. Which tool to use, and what to send it. Those fail independently, and they need different tests.

Wrong tool looks like this. A support agent can look up orders, issue refunds, and search help docs. A customer asks what the refund policy says, and the agent calls the refund function instead of the docs search. The arguments might be perfect. The choice was wrong.

Right tool with wrong arguments looks like this. The agent picks the order lookup, then sends ord_7282 when the customer asked about ord_7281. The choice was right. The value was wrong.

I also test a third case that teams skip. The case where the agent should not call anything at all. A customer asks what a status word means, and the model already has enough context to answer. If my suite only checks whether a tool call exists, an agent that calls a random function passes. So I include no-tool cases and grade them as strictly as the rest. Did it call something. It should not have. That is a fail.

Deterministic checks first, judge second

I check everything in code before I let another model grade anything. Code is cheap, repeatable, and it does not have opinions.

The first code check is the tool name. When one tool is clearly right for the case, I compare the returned name against the expected name in plain code. No model needed. If several tools could solve the request honestly, I do not force an exact match. I mark those cases for the judge instead, which I cover below.

The second code check is argument structure. I take the same schema I sent the model and validate what came back against it. Missing required fields, wrong types, extra fields the tool does not accept. All of that is mechanical. A boolean that arrived as a string fails here. A call missing its id fails here. I catch the shape problems before I argue about meaning.

The third code check is argument values, and only where I know the answer. If the case says order ord_7281, I compare the parsed argument against ord_7281 directly. Shape and value are separate checks on purpose. A value can satisfy the schema and still be wrong. ord_7282 is a fine string. It is the wrong order. I have seen teams stop at schema validation and declare victory. That is how wrong-but-valid calls reach production.

I go deeper on this code-first mindset in prompt engineering is software engineering. The short version is that tests belong in code wherever code can answer the question.

A judge for the calls code cannot grade

Some calls have no single right answer to compare against. A search query, a summary length, a date range, a choice between two tools that both fit. For those I use a separate model as a judge, with no reference answer. It reads the request, the tools available, and what the agent did, then decides whether the choice was reasonable.

The judge only works if I treat it like a test fixture, not an oracle. I write down what counts as correct before I run it. I keep the same judge and the same instructions for every candidate model I compare. And I calibrate first. I hand-check a small set of cases myself, run the judge over them, and read every disagreement. If the judge disagrees with me on cases I am sure about, I fix the instructions before I trust it on the rest.

I also keep the judge small in scope. If an equality check, a schema check, or a business rule can answer the question, I use that and skip the judge. Judges cost money and drift. Code does not. The judge handles meaning. Code handles everything else.

Test the sequence, not just single calls

Single-call checks miss a whole class of failure in multi-step agents. Each call looks fine alone, but the order is wrong or a required step never ran.

A refund flow might need lookup, then eligibility check, then the refund itself. If the agent skips the middle step, every individual call passes its own test and the workflow still fails. So for flows where order matters, I grade the sequence directly.

I match the strictness to the requirement. When the steps must run in order, I enforce the order. When two lookups can run in either order, I check that both ran instead of demanding my exact sequence. And when several paths reach the same valid end, I grade the end state rather than the path. The question is always what the business required. If the requirement was the final state, testing the exact path rejects good work for no reason.

Run the same suite across models without changing the rules

Once the cases and graders exist, comparing models is straightforward. The discipline is in keeping everything else still.

Same test cases for every candidate. Same grading rules. Same settings where it matters, including how much reasoning effort each model is allowed, because effort changes calling behavior. Same routing setup, so one model is not quietly hitting a different provider path than the other. I include the no-tool cases in every run, since unnecessary calling is one of the first things that differs between models.

I score selection, structure, and values separately, then roll them into a pass per case. A case passes when the tool choice is right, every call is structurally valid, and the known values match. For no-tool cases, the tool choice is the whole grade. That split tells me what to fix. If selection is weak, I work on tool descriptions and routing. If structure is weak, I tighten schemas and required fields. If values are weak, I look at how context reaches the call.

For background on what that comparison costs at volume, see orchestration and spend.

Mistakes that waste weeks

These are the ones I keep seeing.

Checking that a tool call exists instead of checking which one. A model that calls the wrong function still produces a call. Existence proves nothing.

Demanding one exact sequence when several are valid. That turns the suite into a test of your reference, not the agent. Grade the required calls or the end state.

Treating schema-valid as correct. Structure and values are different checks. Run both.

Skipping the no-tool cases. Unnecessary calling is a real failure with real cost. Every wasted call burns latency and money and sometimes triggers side effects.

Changing the judge or the settings between candidates. If the grader moves, the comparison means nothing. Freeze the judge, the instructions, and the settings, then compare.

FAQ: How do I tell a wrong tool from a wrong argument

Check the tool name first. If the name does not match the expected one, stop there. That is a selection failure, and fixing argument prompts will not help. If the name matches, parse the arguments and validate structure, then compare values. I fix selection with clearer tool descriptions, fewer overlapping tools, and better routing. I fix arguments with tighter schemas, required fields, and cleaner context.

FAQ: When should I use an LLM judge instead of code

Use code wherever a fixed answer exists. Tool names, known ids, enums, required fields. Bring in a judge only where correctness depends on context or several answers are valid, like search queries or close tool choices. Keep the judge instructions written down, keep the same judge for every comparison, and hand-check a small calibration set before trusting it broadly.

FAQ: How many test cases are enough to start

Start with ten to twenty cases that cover your real traffic. Include the three happiest paths, two no-tool cases, and five cases drawn from actual failures you have seen. That small set catches most regressions. Grow it every time production surprises you. Each incident becomes one new case, so the suite compounds instead of rotting.

FAQ: How do I compare two models fairly

Freeze everything except the model. Same cases, same graders, same judge and instructions, same reasoning effort, same routing. Run the no-tool cases too. Score selection, structure, and values separately so you know what actually differs. If you change any part of the harness between runs, rerun both sides.

About the Author

I am Brian Marvin, an AI-native Fractional CTO with 30 years in technical leadership. At empowered.guru, I help startups and small businesses build MVPs, shape roadmaps, and make AI-powered technology decisions that scale.

Filed under

AI AgentsTestingEvalsTool CallingFractional CTO
$empowered.guru --book-session

Keep exploring

Turn the next insight into a shipped product.

Bring us the product, architecture, or delivery problem you are working through. We will help you find the clearest path forward.