AI agent analytics should start before a team scales an automation. This post treats an agent as a workflow with inputs, tool calls, decisions, outputs, and a human fallback. The aim is not to crown a model or claim a benchmark result. The original post did not name a model, record prompts, publish a test set, or include run-time and cost data. Those numbers should not be reconstructed after the fact. Instead, the two existing visuals show the measurement system a team needs before it trusts an agent with more volume.
That distinction matters. A run can finish, return text, and still create a bad business outcome. A lead may be routed to the wrong owner. A support summary may omit the detail that changes the reply. A CRM update may add duplicates. Agent analytics connects the visible output to the work required to correct it.
- What the analytics test sets out to check
- Seven metrics to collect
- Parameters and evidence to record
- What the two existing outputs show
- Choosing a model after measurement

What the analytics test sets out to check
An AI agent analytics test should check the complete task, not only the final sentence. Start with a small, frozen set of real examples that represent normal work and known trouble spots. For a lead-routing agent, that might include complete forms, missing company names, duplicate contacts, non-business addresses, and requests that need a human owner. For a support triage agent, use routine requests, ambiguous requests, policy-sensitive cases, and messages that must escalate.
Every case needs an expected business outcome before the run. That could be a correct queue, a required field set, a summary with named facts, or a handoff. It does not have to mean one exact string. A reviewer can score whether the result met the acceptance rule. OpenAI’s evaluation guidance makes the same practical point: define the task and criteria, run representative inputs, then inspect results before changing the system.
Keep the test set stable while comparing a prompt, tool schema, or model. Then add new failures as regression cases. This turns one-off operator feedback into a durable check. If a routing rule breaks after a model change, the team can see the difference instead of relying on anecdotes.
7 AI agent analytics metrics that matter
1. Valid completion rate
Count a completion only when the workflow reaches a valid end state. A successful HTTP response is not enough. The record must land in the expected place, contain the required fields, and avoid a disallowed action. Report valid completions divided by all attempted cases, then split the result by case type. One blended percentage can hide a weak edge case.
2. Decision accuracy
Measure whether the agent made the right classification, routing choice, extraction, or recommendation. Compare the output with the expected outcome in the test set. If reviewers disagree, log the reason and tighten the acceptance rule. This gives the team a quality number that is separate from whether the automation merely finished.
3. Human handoff rate
Handoffs are not failures by default. A payment dispute, a safety question, or a sensitive customer request may require one. What matters is whether the agent escalated the right cases and supplied enough context for the person taking over. Track unnecessary handoffs and missed handoffs separately. They point to different fixes.
4. Retry and tool-error rate
Record retries by step: model call, search, database lookup, CRM write, or notification. A high retry rate often means an integration, timeout, or data-shape problem rather than weak reasoning. This metric prevents teams from spending time changing prompts when the real fault sits in a tool call.
5. Time to valid completion
Measure elapsed time from intake to a valid result. Keep model latency separate from queue time, tool time, and reviewer time. An agent that writes a correct answer in seconds may still be slow if an upstream lookup stalls. The breakdown makes the bottleneck visible.
6. Rework rate
Track how often an operator edits a draft, corrects a field, reruns a job, or reverses an action. Rework is often the most honest measure of usefulness because it captures work that a completion counter misses. Sample the edits too. Repeated changes to one field usually signal a missing instruction, validation rule, or source of truth.
7. Cost per valid outcome
Use the total cost of model calls, tools, retries, and human correction divided by valid outcomes. Do not substitute a provider list price for observed cost. This post has no recorded Wiro model runs, so it cannot report a defensible runtime or cost per output. Once runs exist, store the model version, input and output settings, tool calls, elapsed time, and charged cost alongside each case.

Parameters and evidence to record
A useful test record names the agent version, model and version, system instructions, user input, temperature or other generation settings, tool definitions, tool responses, retries, time stamps, and final output. It also records the expected outcome, reviewer decision, failure label, and cost. If a run calls more than one model, preserve each call as its own event. Without that detail, a team cannot tell whether a change improved the model, the prompt, or an integration.
Do not invent precision. If Wiro or the run log does not show a duration or cost, report it as unavailable. Compare like with like: same test case, same tool access, same validation rule. Microsoft’s agent orchestration guidance is useful here because it treats the workflow, handoffs, and orchestration pattern as part of the system, not background detail.
What the two existing outputs actually show
The first inline image is a conceptual analytics dashboard. It emphasizes that run volume is the easy number and that reliability and usefulness need their own fields. It does not show results from a named model, a prompt, a fixed data set, or a billed Wiro run. The correct reading is directional: build the dashboard before volume makes root-cause analysis expensive.
The second image presents the three signals that should sit beside one another: completion, handoff, and retries. It does not demonstrate that one model beat another, and it should not be treated as a benchmark chart. Its practical message is sharper: investigate a workflow with high completion but high rework, or low retries but too many escalations. Those patterns can reveal a system that looks healthy from one metric and fails at the operator desk.
Choosing a model after the measurement
Pick a model only after the task and pass criteria are clear. Use a smaller or faster option when the task is structured, the expected output is narrow, and validation catches mistakes. Use a stronger reasoning model when the test set shows that ambiguity, multi-step planning, or exception handling drives errors. A workflow that handles sensitive actions should still retain explicit approval and handoff rules, whatever model it uses.
Run the same cases under the same settings, then compare valid completion, decision accuracy, latency, rework, and observed cost per valid outcome. Anthropic’s guidance on defining success criteria similarly recommends measurable criteria before testing. The right model is the one that meets the accepted outcome at the needed speed and cost, not the one with the most impressive single demo.
For workflow-specific examples, see lead generation workflow agents and app review backlog AI agents. Apply the same measurement discipline to each workflow, then scale only after the failures are understood.