A large tan percentage-style circle on the left sits above an empty outlined box, while on the right a smaller terracotta-outlined circle sits above a box filled with horizontal lines — a claim without a denominator beside a claim with its evidence.

How to Verify an AI Agent Vendor's Claims (2026)

Every healthcare AI vendor publishes a percentage. Almost none publish the denominator. Here's how to read automation claims, what to ask before signing, and how to measure whether an agent actually worked once it's live.

Prasad ThammineniHealthcare
9 min read

Every healthcare AI vendor publishes a percentage. Almost none publish the denominator.

The short answer

An automation percentage is unverifiable without three things: what it is a percentage of, over what period, and what happened to the remainder. "90% automated" can describe the same system as "40% automated" depending on whether the denominator is all transactions or only the ones the vendor already supports. Before signing, get the denominator, the exception path, and the measurement plan in writing. After signing, the only claim that means anything is one measured against a baseline you recorded before go-live.

Key facts

  • A percentage without a denominator is a marketing number, not a measurement
  • Staff-time savings appear within days; denial-rate changes need a full billing cycle, because claims in flight were filed under the old process
  • Realistic automation rates are ranges, not points: ~70–80% for message routing, ~90% of major payers for eligibility (both Agentman measurements — see the denominators below)
  • Independent cost benchmarks exist for the workflows vendors quote against: the 2023 CAQH Index prices a manual eligibility verification at $7.97 to the provider and a manual prior authorization at $10.97
  • Automation claims about eligibility rest on a public standard — the X12 270/271 transaction set — so "we automate eligibility" is a checkable statement, not a proprietary one
  • Baselines must be recorded before deployment — retroactive baselines are not baselines

How to read a percentage

Four questions turn a headline number into something checkable.

1. Percentage of what? "90% of eligibility checks automated" and "90% of eligibility checks from supported payers automated" are different claims. If a vendor supports sixty percent of your payer mix, the second sentence describes fifty-four percent of your actual work.

2. Measured when, and for how long? A number from a two-week pilot at one practice is an anecdote. A number across a quarter and multiple practices is a measurement. Both get printed the same way.

3. What counts as automated? If a check runs electronically but a human reviews every result before it posts, is that automated? Vendors answer this differently, and almost none define it on the page.

4. What happened to the remaining share? This is the question that matters most operationally. The ten percent that did not automate does not vanish — it lands on someone. Whether it arrives as a clean exception queue or as a surprise is the difference between a good deployment and a bad one.

The five questions to ask before signing

  1. What is the denominator behind your headline percentage?
  2. What happens to the cases you do not handle — who does that work, and how does it reach them?
  3. What does failure look like, and how will I see it? An agent that fails loudly is safer than one that fails silently. Ask to see the error path, not just the happy path.
  4. What is the fully loaded price? Including minimums, setup or onboarding fees, and overage rates. In healthcare voice AI specifically, most vendors do not publish any rate at all, so this has to be asked directly.
  5. What will you measure with me after ninety days, and what result would count as a failure? A vendor who cannot name a failing outcome has not defined success either.

Get the answers in writing. A vendor describing a product will answer all five; a vendor describing a demo will deflect at least two.

Comparing vendor claims honestly

Claim shapeWhat it tells youWhat to ask
"90% automated"Nothing yetOf what denominator, measured over what window?
"Saves 15 hours a week"More useful — a unit you can checkAt what patient volume, and doing which specific task?
"65% fewer denials"Useful if scopedIn which denial categories, over how many billing cycles?
"$107K–$149K saved"A range is a good signPer what — practice, provider, year? Projected or realized?
"Trusted by N practices"Nothing about performanceWhat is your retention rate at twelve months?

A range usually signals more honesty than a single figure. Real deployments vary by payer mix, specialty, and starting state; a vendor quoting one precise number across all of them is quoting a best case.

How to measure it yourself after go-live

This is the part most practices skip, and it is the only part that produces a real answer.

Before deployment, record the metric you expect to move — for at least two weeks:

  • Staff hours on the specific task, timed rather than estimated
  • The relevant denial category as a share of total claims
  • Time from arrival to completion for whatever the agent handles

After deployment, re-measure at the right interval:

  • Time metrics: 30 days. These move fast and are easy to attribute.
  • Denial metrics: a full billing cycle, minimum. Claims filed before go-live are still returning under the old process; comparing too early measures noise.

Then state a verdict plainly — it worked, it did not, or the result is confounded by something else that changed. That last one is common and honest: if you added staff, changed payers, or shifted your specialty mix in the same quarter, the comparison is compromised and should be labeled as such rather than claimed as a win.

Without a pre-deployment baseline, every post-deployment number is unfalsifiable. That is why baselines get recorded first, and why retroactive ones do not count.

Applying this to our own numbers

Fair is fair — the same test, run on the figures we publish:

Our claimDenominatorWhere it comes from
70–80% of messages auto-routedAll inbound faxes, voicemails, and portal messagesInbox triage agent; a range because the ambiguous share varies by practice
~90% of eligibility checks automatedChecks against major commercial and government payers, not all payersEligibility agent
65% fewer denialsEligibility and prior-auth-related denials only, not all denialsValley Diabetes case study
$107K–$149K annual savingsPer physician, projected from measured time and denial changesSame case study; projected, and labeled as such
$0.50 per eligibility checkList price, no minimumPricing

Two of those are projections rather than realized cash, and the denial figure covers one category rather than all denials. Stating that costs us nothing and tells you exactly how much weight the numbers carry.

Frequently Asked Questions

What does a 90% automation claim actually mean?

On its own, nothing verifiable — the number is meaningless without its denominator. Ninety percent of what? All eligibility checks, or only the ones from payers the vendor already supports? Measured over which period, at which practice, counting a check as automated at what point? The same underlying performance can be presented as 90% or 40% depending on what sits under the line, so the first question about any percentage is what it is a percentage of.

What should I ask a healthcare AI vendor before signing?

Five questions separate a real claim from a demo. What is the denominator behind your headline percentage? What happens to the cases you do not handle — who does that work? What does a failure look like and how will I see it? What is the fully loaded price including minimums, setup fees, and overage? And what will you measure with me after ninety days? A vendor who answers all five in writing is describing a product; one who deflects is describing a demo.

How do I measure whether an AI agent actually worked?

Baseline before you deploy, not after. Record the metric you expect to move — staff hours on the task, denial rate in the relevant category, time to complete — for at least the two weeks before go-live. Then re-measure after a stated interval, typically thirty days for time metrics and a full billing cycle for denial metrics. Without a pre-deployment baseline, any post-deployment number is unfalsifiable.

Why do denial improvements take longer to show up than time savings?

Because claims in flight were filed under the old process. A change to eligibility verification affects claims submitted after go-live, and those take weeks to adjudicate and return. Staff-time savings appear within days; denial-rate changes need a full billing cycle before the comparison means anything. Vendors who report denial improvements in the first two weeks are measuring noise.

What is a reasonable automation rate to expect?

It varies by workflow and should always be a range rather than a single number. Message routing runs about 70 to 80 percent automated because the remainder is genuinely ambiguous. Eligibility verification reaches roughly 90 percent of major payers because coverage is structured data. Any vendor quoting a single precise figure across every practice and payer mix is quoting a demo, not a deployment.

Should I trust a case study from a vendor's own customer?

Trust it as far as it is specific. A useful case study names the practice, the specialty, the starting state, the measurement window, and what did not improve. A weak one gives a percentage and a logo. Ask whether you can speak to the practice directly — vendors with real results usually say yes.

What to do next

Take the five questions to every vendor on your list, including us. Record a baseline before you deploy anything — two weeks of honest measurement is enough, and it is the difference between knowing whether something worked and believing it did.

If you want to see what a published denominator looks like, our pricing and per-agent pages state rates and scope without a form.

Ready to automate your back office?

See how production-grade AI agents handle your toughest workflows.