ALPNAISWISS INTERNATIONAL

SPEND PROOF / AGENT-PLATFORMS

Compare variants before changing your routing.

A model’s price does not measure the cost of a completed task.

An agent platform can lower the price of a call while increasing retries, searches and failures. Spend Proof compares a baseline with a proposed variant on matched tasks. The report helps identify what deserves a controlled trial before you change a production router or workflow.

Evaluate a different model-and-tool combination

Include model, tool and retrieval charges in each attempt’s cost. Then compare cost per successful result: a budget model that needs more tool calls no longer looks artificially economical.

Make retry-policy costs visible

Group attempts under the same task. Check whether a retry strategy produces more usable results or mainly consumes budget. Repeated attempt rows do not become additional successful tasks.

Reject a comparison that is not ready

The diagnostic flags unmatched tasks, insufficient samples, declining success or a breached duration ceiling. Even when thresholds pass, it recommends a controlled trial and never authorizes a production deployment.

What you receive

  • Comparison by workflow, including matched tasks and tasks missing from either variant.
  • Cost per success, retry counts and P95 of summed attempt durations when every duration is available.
  • An explicit decision, detailed checks and the evidence still needed to continue.

The minimal export contains execution metrics and opaque identifiers. It needs no prompts, user content, provider access or production keys. Supply complete costs and consistent success criteria across variants.

Does Spend Proof automatically change our model or budgets?

No. The current tool analyzes supplied traces; it does not connect to your infrastructure or alter routing. A favourable finding remains an invitation to run your own controlled trial.

Why measure tasks instead of API calls?

A task may need multiple calls and retries. Dividing costs by calls alone can conceal expensive failures. The report totals all attempts and divides by distinct successful tasks; success labels come from you and are not independently verified.

Can we already monitor a fleet of agents continuously?

Not with the currently available product. Continuous collection and provider connections remain to be built. The USD 99 monthly monitoring hypothesis is not a commercial offer; current API access is for sandbox trials.

Compare my variants