6 min read

How to measure whether your AI agents are actually working

Task counts prove nothing. Four metrics that show whether an agent system is producing value, and the vanity numbers to ignore.

Last updated 11 August 2026

Most agent dashboards report activity. Activity is not value, and a system that ran 412 tasks tells you nothing about whether it was worth paying for.

Four metrics that mean something

  • Approved output per week. Not produced — approved. Work you rejected cost money and delivered nothing.
  • Approval rate over time. If it is climbing, the system is learning your standards. Flat or falling after a month is a real problem.
  • Hours returned. Estimate honestly what the approved output would have taken you. This is the number that justifies the invoice.
  • Time to first response, if you automated anything customer-facing. It is the metric most directly tied to revenue.

Vanity metrics to ignore

  • Tasks executed. Cheap to inflate, meaningless alone.
  • Tokens consumed. A cost, presented as an achievement.
  • Number of agents. More agents usually means worse reliability, not more capability.
  • Uptime, if the output is mediocre. A reliably unhelpful system is still unhelpful.
Set the bar before you start

Write down on day one what day thirty should look like in numbers. Teams that skip this argue about impressions later and usually cancel something that was quietly working.

Our dashboard answers three questions on purpose — what it did, what it made, what needs you. See it.

Keep reading

18 Aug 20268 min read

Why multi-agent AI systems fail in production

Most multi-agent demos work and most multi-agent deployments do not. The reason is error compounding — and the fixes are structural, not model upgrades.

Read
Want this done rather than explained?

We build the agent team, connect it to your accounts and supervise the output. Twenty-minute call · See pricing