Skip to content
Journal / education

Agent observability: why every tool call should show its attempt chain

Routing transparency for AI agents: attempt chains, per-call prices in every response, and why tool infrastructure should never be a black box.

Here is a question you should be able to answer about any tool call your agent made yesterday: which provider actually served it, who was tried first and failed, and what did it cost, in dollars, that specific call? If your infrastructure cannot answer that, you do not have observability. You have vibes and a monthly invoice.

I run a tool router, and the single most consequential design decision we made was not about routing algorithms. It was that every response includes the full story of how it was produced. This post is about why that matters and what to demand from any tool layer, ours or one you build yourself.

The black box problem

Any layer that adds failover or provider selection is, by definition, making decisions on your behalf. That is the value. It is also the risk, because a layer that silently decides is a layer you cannot debug.

The failure mode looks like this: your agent’s answers got slightly worse on Tuesday. Was it the model? The prompt? Or did your search layer quietly fail over from its primary provider to a weaker one for six hours during an outage you never heard about? Without per-request routing data, you cannot distinguish these, and teams burn days debugging prompts when the actual story was infrastructural.

The fix is structural, not heroic: the routing decision must travel with the response. Every response from our router carries a routing.attempted array, the full chain of providers tried, in order, with what happened at each hop. If Serper timed out and Brave answered, the response says so. Not in a dashboard you check later. In the response your code just received, where your agent’s own logs and traces can pick it up.

What an attempt chain gives you in practice

Silent degradation becomes visible. Failover is supposed to be invisible to your users, but it should never be invisible to you. A chain that reads “provider A failed, provider B answered” showing up in 20% of yesterday’s requests is a fact you want in your monitoring, because it explains latency shifts, quality shifts, and cost shifts before you go blame the model.

Retries stop being mysterious latency. A request that took eight seconds instead of two is alarming until the chain shows two timeouts before the success. Now it is not a mystery, it is a documented bad afternoon at two vendors, and you can decide whether your failover budget is set correctly. (Ours caps failover with per-category wall-clock budgets, 20 seconds for search up to 330 for video, so a chain can never hang your agent indefinitely.)

Circuit breaker behavior is auditable. Our router skips providers whose error rate exceeds 30% over a five minute window. When that happens, you can see it in the traffic: the skipped provider stops appearing in chains. Infrastructure that changes its behavior should leave evidence.

Price belongs in the response too

The second half of transparency is money. Every response from our router includes the exact price of that call. Not an estimate, not a monthly rollup: this request, this provider, this many dollars.

This sounds like an accounting nicety. It is actually an agent observability primitive, because it lets you attribute cost to the same traces you already collect. Tag each tool call’s price with the agent session that made it and you can answer the question that monthly invoices never can: what does one run of this agent cost, and which step is eating the budget? I walked through that exercise in auditing your agent tool spend; it is only possible because the per-call price exists at the moment of the call.

Failover makes per-call pricing non-negotiable, by the way. If cheapest-first routing sends a search to a $1.20 per 1k provider but failover lands it on a $6.00 per 1k one, your unit economics for that request just changed by 5x. You should not have to discover that at the end of the month. And the flip side of billing honesty: attempts that failed appear in the chain but not on the bill. You pay for the call that succeeded, and the response shows both facts side by side. More on that philosophy in what honest billing looks like.

Demand this from any layer, including one you build

None of this requires our router. If you build your own fallback wrapper, give it the same properties from day one:

  • Return the attempt chain in-band, with the response, not only in logs.
  • Attach the actual per-call cost, derived from the provider and units used.
  • Make automatic behaviors (retries, skips, provider swaps) leave visible evidence.
  • Never charge yourself, or bill your budgets, for attempts that failed.

The pattern costs little to implement early and is miserable to retrofit. The reason I harp on it is that routing layers earn trust by being auditable; the value proposition is “we make decisions for you,” and the only acceptable version of that sentence ends with “and we show you every one.” Our routing docs document the exact response shape, and the pricing model (list plus a public 20%) is published for the same reason: a layer you cannot verify is a layer you eventually stop trusting.

If you want to see attempt chains in practice, sign up, make a few calls with the $2 in free credits, and read the routing block in what comes back: dashboard.