Hadto note

Operating Notes · 2026-05-08

How to measure an AI-run business without fooling yourself

Once agents handle intake, quoting, and follow-up, output numbers improve almost automatically. Here are the owner-side measures that show whether the business is actually getting stronger: rescue rate, retained learning, recovery time, and shared exposure.

Who this is for

This is for small-business owners running AI-assisted operations who want to know which numbers prove the business is getting stronger and which numbers just prove the software is running.

Once AI runs part of your business, your standard numbers start flattering you, because the things dashboards usually count (drafts, replies, completions, response times) are exactly the things AI makes cheap.

ai operationsowner operatorsoperator dashboardsownership systemsautomation

The week you turn on AI for intake, quoting, or follow-up, your numbers improve. Response time drops. More estimates go out. More messages get answered. Every draft, reply, and task completion costs almost nothing to produce now, and those are precisely the units most dashboards count. So the counts rise whether or not the business underneath them got any better. Emad Mostaque makes the macro version of this point in The Last Economy: cheap cognition can multiply visible throughput without creating a stronger human position anywhere in the system, and legacy activity metrics can stay healthy-looking while the reality underneath them gets worse. I think the small-business version of that warning is more practical than the macro one, and this essay is my attempt at it: which numbers still mean something once software produces the activity, and which numbers have quietly become vanity metrics.

What a green board can hide

Picture a home-services dashboard tracking booked jobs, technician utilization, callback resolution, and response speed. Useful numbers. Now give the system one target, keep throughput rising, and watch what it learns to do. Rushed diagnostics start slipping through because faster closes help the board. Callbacks get tolerated because recovery work still fills the schedule. Customers get over-messaged because contact volume reads as engagement. And you stay trapped in exception review because your intervention keeps saving the metric. Every tile is green. The board is active, the trend lines point up, and the business is getting weaker in four specific ways the board was never asked to show.

Name the wins that must not count

Mostaque’s Chapter 4 example is GDP, which can book disaster recovery, sickness treatment, and attention capture as growth. He calls it a dashboard for insanity. A business dashboard fails the same way through the same mechanism: give an optimizer one target with no stated exclusions, and it will find the substitute wins before you name them. The fix I use is a three-part contract for any metric an automated system is allowed to optimize. State the target. State the protected values the target may not trade away, things like diagnostic honesty and customer trust. Then state the anti-goals, the wins that must never count: callback recovery counted as utilization, owner interruptions counted as responsiveness, stretched customer attention counted as engagement. Anti-goals sound like ethics language, but they belong on the instrument panel, because they name the cheats in advance. When the protected values stay implicit, the system discovers them through damage.

Watch the rescues, not the volume

Output metrics answer one question: did activity happen? Owner metrics answer a different one: could this business run well without you? Those call for different instruments. On the owner side, the numbers I trust most are first-pass decision quality, meaning how often the operator’s call stood without rework; rescue rate, meaning how often finishing a job cleanly required you or a senior expert to step in; and how many exceptions got resolved from shared records instead of somebody’s memory or a direct-message thread. Rescue rate is the sharpest of the three. An AI layer can draft the estimate and classify the intake while borrowing all of its judgment from context that still lives in your head, and the borrowing shows up in a place you can count: the same edge cases keep coming back to you. If output climbs while rescues stay flat, the system automated motion without transferring capability.

Book progress on three axes

Our rule for calling any AI gain “progress” is that it has to show up in three places at once. Speed: the loop actually moved faster. Compounding: the gain got preserved somewhere the next run can use it, in the script, the estimate standard, the queue design, rather than living in one person’s prompt. Structural trust: the system became easier for another operator to inspect and govern. Each axis fails silently on its own. Speed without compounding is a burst you will pay to rediscover. Compounding without trust builds an expert-only system nobody can inherit. Trust without speed is governance work that has not yet earned its keep. When one axis is missing, say so in the review instead of rounding up.

Ask what each run leaves behind

A loop that runs every morning is not necessarily learning anything. Mostaque’s Chapter 6 argument is that the systems that persist are the ones that compound information and keep knowledge from slipping away over time. The negative case is a loop that spends its intelligence reconstructing its own context each cycle: what matters, what changed last time, which decision became policy. It produces a plausible report every day and retains nothing. So after each run, ask a narrow question: what durable artifact now exists that did not exist yesterday? A decision attached to its evidence, a rule changed in a place the next operator can find, an exception preserved with its resolution. If the answer is repeatedly “none,” you are measuring repetition and calling it learning.

The warnings arrive before the break

Mostaque’s Chapter 2 argues that failing systems advertise themselves before the obvious break, through slower recovery from disruption, higher variance, thrashing between states, and failures spreading across boundaries. Translated to your operation: a queue that used to return to baseline in a day now takes three. Revenue looks stable while rework and cycle time swing wider. A lane runs clean under one dispatcher and falls apart under another, which usually means the real method still lives in private judgment. A single broken intake field starts showing up in quoting, billing, and closeout at once. Each event gets explained away locally; together they are the system talking. These are measurable now, before anything breaks: track time-to-recover after exceptions, track the spread between best and worst case rather than only the average, and track which lanes hold only under one specific person.

Measure the portfolio, not the tiles

Ten automations can each pass their own health check while all ten depend on the same model vendor, the same aging pricebook, and the same two people for escalation. That is not ten capabilities; it is one shared failure domain wearing ten labels. Mostaque’s warning that old dashboards miss the shape of the system they claim to govern lands hardest here, because per-loop status is the view that hides shared exposure. So review the portfolio on its own schedule. Which dependencies are common enough that one break damages several loops? Which artifacts are old enough that “working” means “working from stale memory”? Who terminates most escalations, and could the business survive their vacation? Which loops have a real fallback route, and which have a timeout followed by panic?

The owner scoreboard

Pulled together, the panel I would actually run:

  • First-pass decision quality: how often the operator’s call stood without rework.
  • Rescue rate: how often clean completion needed you or a senior expert.
  • Exceptions resolved from shared records rather than memory or direct messages.
  • Retained learning per cycle: the durable artifact each automated run left behind.
  • Time-to-recover and spread: recovery after disruption, and best-to-worst variance, not just averages.
  • Person-dependent lanes: which workflows hold only under one specific person.
  • Shared exposure: providers, stale artifacts, and escalation paths that several loops have in common.
  • Anti-goal checks: attached to every target an automated system optimizes.

Keep the output numbers too. Volume, response time, and close rate still matter for capacity planning and customer experience; the argument is that they can no longer run the review, not that they are worthless.

Two honesty notes to close. First, The Last Economy argues at macro scale, and whether my translation of it down to a twenty-person service company holds is an open question; no controlled comparison backs these metrics yet. So treat the scoreboard as a starting panel: adopt the two or three measures that name your current blind spot, and keep only what changes a real decision within a quarter. Second, the test that outranks every number on the panel: when a metric improves, ask whether the gain made the next operator more capable of running this business, or just made the failure easier to harvest. AI is working when it produces a more capable owner. More output was never the hard part.


Source evidence used in this note: Emad Mostaque, The Last Economy, for the arguments that GDP-style measurement can count disaster recovery, sickness treatment, and attention capture as growth (Chapter 4), that cheap cognition can raise visible throughput without strengthening the human position inside the system, that persistent systems compound information over time (Chapter 6), and that failing systems show slower recovery, higher variance, state thrash, and boundary-crossing failures before the obvious break (Chapter 2). Hadto interpretation: the anti-goal contract, the owner-metric set, the three-axis progress rule, and the portfolio review are operating judgments translated from the book’s macro claims, and they have not been validated by controlled comparison.

← Back to all notes