AI Agents ยท Payments ยท Research

Why every agentic payments number you have seen is measuring something else

No one has measured the agentic payments market. The filters that separate real payments from machine noise delete agent behaviour by design.

Vin Lim Founder, Astralab

We build AI agents for a living, so when we started work on a payment rail for them, the first thing we wanted was a number. How big is agentic commerce today, in dollars that moved because software decided to spend them? There is no good answer. Nobody has measured the category, and the one rail that has been measured carefully turns out to be mostly manufactured. After a few weeks of looking we stopped treating that as a gap in our research and started treating it as a fact about how the payments industry counts.

Why do stablecoin bot filters delete agentic payments?

Because every serious attempt to separate real stablecoin payments from on-chain noise filters out exactly the behaviour an agent exhibits.

BCG, working with Allium, excludes any wallet with 1,000 or more transactions, or $10m or more in volume, in any 30-day window, on the assumption that it is a bot. Visa Onchain Analytics, also using Allium, publishes the identical threshold. McKinsey makes the same move from the other direction, distinguishing its own estimates from volumes published elsewhere on the grounds that those may include bot-based payments, intra-exchange volume, and other high-frequency activity.

This is good methodology, and the numbers it produces are the reason anyone knows how thin real stablecoin payment volume is. Raw on-chain volume is mostly machine noise. BCG takes 2025's $62tn down to $11.7tn by stripping bots, likely bots and intermediaries, then down to $4.2tn after removing a further $7.5tn of internal transactions, and narrows again from there to $350bn to $550bn of observable bilateral payments for goods and services. McKinsey lands independently at about $390bn annualised. The raw headlines therefore overstate real payment volume by roughly 90x in McKinsey's case, $35tn against its $390bn, and by 113x to 177x in BCG's.

Two caveats their authors would want kept. BCG says it deliberately biases downward and calls its figures directionally robust lower bounds. McKinsey's $390bn is annualised from a single month of December 2025 activity, per its own footnote, rather than summed across a year. The convergence is also a little softer than it looks: the two houses work with different analytics partners, Allium and Artemis respectively, but BCG's own exhibit footnotes cite an Artemis stablecoin update for at least one assumption, so the inputs are not fully independent.

Now apply that filter to the thing we were trying to size. An AI agent paying per API call is a high-frequency wallet, and that is not a side effect of the design, it is the design. An agent buying data, compute and tool calls as it works will pass 1,000 transactions in 30 days without doing anything unusual. It gets classified as a bot and removed before the payments figure is computed.

We ran a case-insensitive search across the full text of BCG's paper for agent, agentic, autonomous, machine, MEV and arbitrage. Zero occurrences. A paper whose entire purpose is separating real payments from machine traffic contains no concept of a machine making a real payment. The industry's best measurements of genuine payment activity are, by construction, blind to this category.

Stablecoins are not a special case here. Bitcoin's economic-volume metrics apply the same kind of filter with different mechanics. Glassnode's entity-adjusted analysis cuts raw on-chain volume by 75.5% across 2016 to 2019, implying under a quarter of recorded volume is transfer between genuinely distinct parties. Coin Metrics builds its adjusted transfer value from named heuristics, one of which drops any output spent within 60 minutes of its creation. An agent that receives value and spends it inside the hour, which is the ordinary shape of a pay-as-you-work loop, is deleted by that rule. Every cleaning methodology we looked at removes rapid, high-frequency, machine-driven flows, because until very recently those were always noise.

Don't the agentic rails publish their own numbers?

They publish transaction counts, and the counts are real. The problem is what a count of transactions turns out to represent once someone checks.

x402, the HTTP 402 payment protocol backed by Coinbase and Cloudflare and now a 40-member foundation, reports enormous scale. Chainalysis measured well over 100 million cumulative transactions on Base through Q1 2026, and by August 2026 the x402scan dashboard was reporting roughly 75 million transactions in a trailing 30 days.

Then you read the authenticity work. A population-scale census of Base covering 280 days to 23 June 2026 found 136,708,672 settlements carrying $44,121,383.81 gross, which is about $0.32 each. Of those settlements, 21.20% are fictitious and 63.78% are internal settlement within a linked cluster. Genuinely independent value sits somewhere between $187,861 that demonstrably reaches a nameable service and $20,258,746 that is merely not provably manufactured. The Gini coefficients for payer, recipient and value all exceed 0.98.

Two things are easy to garble here, so they are worth stating plainly. The percentages are shares of settlement count, while the dollar bounds are shares of value. And the upper bound is an absence of proof rather than a positive finding. The paper's own conclusion is the sentence worth keeping: "Settlement count measures manufacturability, not adoption."

To its credit, Chainalysis says as much in its own post. Much of the growth was meme coin farming driven by a pay-to-mint game, where Base's near-zero gas fees let users repeat the payment loop hundreds of times. Weekly wallet retention reached roughly 87% for a week before falling to 5% once the incentive faded.

So the largest agentic payment rail on earth moved somewhere between $188,000 and $20 million of genuinely independent value in nine months. The lower bound says the category barely exists, and the upper bound puts it somewhere below a mid-sized Shopify store.

Why do the 2030 forecasts disagree by 8x?

Not because anyone disagrees about growth rates, but because each house is counting a different thing and none of them says so in the headline.

SourceFigureWhat it counts
eMarketer *~$20bn 2026, $144bn by 2029Checkout completes inside the AI platform
Morgan Stanley *$190bn to $385bn by 2030Autonomously executed only
Bain$300bn to $500bn by 2030 (US)Initiated, influenced, or completed by agents
Juniper$1.5tn globally in 2030Scope undisclosed

* The eMarketer and Morgan Stanley figures reached our research as secondary reporting rather than fetched primary sources, so treat them as directionally indicative only. Both other rows were read from the publisher's own page.

Hold the year constant and the 2030 estimates run from $190bn to $1.5tn, a spread of roughly 8x on figures that are supposed to describe the same market in the same year. Even that flatters them, because Bain's band is US-only while Juniper's is global, so two of the four rows are not comparable at all. Quoting the widest possible gap, eMarketer's 2026 figure against Juniper's 2030 one, would make the spread look like 75x, and most of that would just be four years of growth.

We also dropped a fifth row while fact-checking this piece. A $3tn to $5tn figure is widely attributed to McKinsey, and it does not appear in the McKinsey article everyone cites for it. That article is about stablecoin supply and contains no reference to agents at all.

Bain's number is the one most often quoted as an agentic payments market, and the one that least supports that reading. Its stated scope is purchases initiated, influenced, or completed by third-party AI agents or retailer-hosted agents, excluding shopping journeys that only use AI-assisted search or discovery. An influenced purchase is one where a human sees an agent-generated recommendation and then checks out normally with their own card, which carries zero agent-initiated payment volume. Bain's cited empirical anchor is a vendor statistic about AI influencing $3bn of US Black Friday sales, itself an attribution metric. Its published page carries no CAGR, no baseline 2025 figure, and no methodology section, which we checked on the page itself rather than taking anyone's word for it.

Neither Bain nor Juniper discloses a bottom-up build: no agent counts, no transactions per agent, no average ticket. Juniper's published methodology detail is that the report draws on 38,000 datapoints over a five-year period, which describes a spreadsheet rather than a model.

One more trap, and we nearly fell into it ourselves. Juniper publishes its quantities separately from the $1.5tn: 1.3 billion agentic commerce users by 2031 in a second release, up from under 300 million in 2026, and 120 billion transactions by 2031 in a third from August 2026. It is tempting to divide the $1.5tn by the 120 billion to get an average ticket. That is three releases across two horizons, none of which states its scope. The arithmetic works and the result means nothing.

Is anyone measuring this properly?

Visa is, and the numbers are not public. The most interesting artifact we found was Visa Onchain Analytics' agentic payments dashboard. It tracks x402 and MPP across six metrics: total volume, organic volume, total transactions, buyer wallets, merchant wallets, and median transaction size. Its own page copy describes AI agents autonomously paying for compute, data, and API services as a new economic primitive.

Two details there are worth more than the dashboard itself. The dashboard's own format string for median transaction size, $$.4f, renders to four decimal places, so whoever built it expects sub-cent payments and designed the display around them. That tells you more about the shape of this market than any forecast we read. And the numbers never arrive: the page ships its six metric labels and then renders loading skeletons where the values should be, in the server payload and in a real browser alike. The page's structured data names Allium as the provider, so Visa is publishing what Allium measures. The measurement exists, and it sits behind a data agreement.

What should you measure instead?

If you are building here, you cannot borrow a number, and the useful move is to stop trying. We build ours bottom-up and show the arithmetic: paying agents, times purchases per agent per year, times average ticket, with every input either sourced or explicitly flagged as an assumption. The output is a band rather than a point, which is less impressive to present and much harder to argue with.

Definition does most of the work. We count payments executed by an agent, not influenced by one, because that is the only definition under which an agentic payment rail is the thing being bought. That rules out the largest published numbers, which is the point of choosing it.

Two smaller habits are worth copying. When you multiply four uncertain factors together, compounding all the lows against all the highs produces a range so wide it says nothing, so take the geometric mean for the centre: the error is multiplicative, so the middle should be too. And separate observed figures from modelled ones everywhere, including visually, with one treatment for anything carrying a source and a date and another for anything downstream of an assumption. That is mostly a device for not fooling yourself, which is the real risk when the data is this thin.

So is the market real?

There is a version of this piece that concludes agentic payments are hype, and that is not what we think. The category is unmeasured rather than absent, and those are very different situations for anyone deciding whether to build. A market with 40 serious institutions building rails into it, a top-tier security venue publishing audits of its infrastructure, and a major card network standing up a dashboard to watch it, has not failed. It has not yet produced the thing worth counting.

The uncomfortable implication is for anyone already claiming share. If the leading rail's genuinely independent volume is bounded at $20 million and might be two orders of magnitude below that, nobody has won anything yet. The transaction counts that look like a moat are, on the measurement literature's own reading, a record of how easy those transactions were to manufacture.

Sources

Every figure above traces to one of these. We have marked what each one is, because provenance matters more than usual when a category has no audited data.