The week the catalogue tripled and found three of our rules wrong
State of the directory, 2026-09-09. All figures as of 08:00 UTC today unless dated. Previous report: week-2026-09-02. Corrections live at /methodology.
On Saturday we imported 3,325 routes on 1,716 hosts — not by count, but by evidence: every one had been paid by at least two independent wallets on the Coinbase Bazaar. Inside three days the new supply found three of our probe rules wrong. All three are corrected, dated, and below. Then a wave bought from 383 of them, a second instrument checked every purchase, and we learned what a pre-spend verdict can and cannot see. That is most of this week.
Now: 4,933 listings on 1,884 hosts, 4,549 verified, 283,656 probes in the last 24 hours at a 94.0% pass rate, every listing re-probed every 25.7 minutes. 1,020 endpoints bought from and confirmed delivering, out of 2,550 attempted. Two verdicts sold, both to ourselves.
1. What we got wrong this week
We lead with this because a directory that grades others has to grade itself first.
Correction #16 — a zero-address payTo (2026-09-06). An imported listing declared payTo: 0x000…000. Our USDC index dutifully watched it and attributed $12.87M of USDC burns to one listing over ten and a half hours. The listing's page said so, in public, for those ten hours. Fix: sentinel addresses are refused at the index and the rollup, and a verdict on a listing that names one is avoid. Eight daily rows deleted; the economy line for 09-05 restored to $16,515.
Correction #17 — validation before payment graded as failure (2026-09-06, widened the same day). A route that answers 400 or 422 to a placeholder request is telling you it validates input before it asks for money. We counted that as a failed probe, five in a row made it failing, and 48 listings on 25 hosts wore the wrong status for about twelve hours. Fix: it is now the fourth neutral class — no score, no streak, no status. A 403 whose body is a JSON error is the same thing wearing the wrong code; a templated route answering 404 to its literal placeholder is the same thing again. Since Monday the prober also reads the seller's own error body for an example and probes that — which is how one host (api.ukintel.uk) went from unreadable to verified, 110 of 110 probes yesterday. It is also, so far, the only host in the whole validate-first population that prints a usable example.
Correction #15 — a pass without a challenge (2026-09-02, after the last report). 75 listings held verified on 2xx answers alone, never having produced a 402. Verified now requires a parsed challenge at least once.
Method changes, in order, all dated on /methodology: 404 or 410 on both verbs is absence, not neutrality (09-04); neutral probes leave the hot tier (09-04); a 403 that the same host does not give other clients is a refusal of our prober, not a fact about the listing (09-05); ephemeral tunnel hosts leave the public counts (09-05); the three validation cases above (09-06); testnet listings leave the public counts (09-07); a delisted route that answers 402 again comes back on its own (09-07); templated routes are probed with the seller's own example (09-07).
And one we did to ourselves on Tuesday. Wave 7 (below) attempted 1,461 routes, and 245 of its failures were our client's pre-flight verb: 127 answered 404 and 118 answered 405 to a GET pre-flight, on routes that serve their 402 only to POST — a fact the directory already knew for every one of them. That single mistake is why the census's pass rate on payment-required endpoints reads 41.4% this morning instead of the 54.8% it read last week. The verb fix shipped that afternoon; the wave's records stand as recorded, labelled as our limit. We do not re-grade a wave after the fact.
One near-miss, not a correction: /v1/listings?limit>100 threw for about 22 hours after the import, a bound-parameter cap in the database. Found by our own scout, fixed the same hour.
The sentence that recurs: a neutral class needs a check for the case where we are the variable. We wrote it three times this week.
2. The import: evidence, not count
The Bazaar carried 15,299 routes we did not, on 1,707 hosts. We took every route with two or more distinct payers, capped at 25 per host, plus one representative route for every remaining host: 3,325 routes, one refused at the door (DNS). 93% answered 402 on their first probe.
What we learned from them:
- 608 of our active Base listings are not on the Bazaar. We are not a mirror, and it is not a superset.
- Four hosts expose 500+ routes each; 240 are route templates (
/tx/:hash), and 165 of those (69%) answer 402 to the literal placeholder — they charge before they validate. The other 71 validate first, which is what correction #17 was about. - 70 listings answer their 402 with a testnet (Base Sepolia, two spellings). They verify correctly and a payment there carries no value. They are now probed, served, flagged
testnet,cautionin the verdict, and out of the public counts — which moved the public figures by rule: 4,978 → 4,910 listings, 60 hosts. - Our own verdict route arrived in the import, indexed by the Bazaar from its first sale. Delisted as first-party, with a guard so it cannot happen again.
Cost of carrying it: the cron went from every five minutes to every two, probes from ~115k to ~285k a day, and the re-probe promise (30 minutes) holds at 25.7. The on-chain index now watches ~1,200 more payTo addresses, and their history is still backfilling: 1,726 listing pages currently say "watching since 2026-09-05" instead of showing a 30-day figure, because until Tuesday they showed nothing at all — and nothing reads as "nobody pays this" on exactly the routes that were selected for being paid.
3. "126 new listings; one was real"
On 09-04 our own count showed 126 new listings. 123 were one automated onboarding loop: three listings per ephemeral tunnel host, about 41 tunnels a day, submit → probe → delete, no contact address. 77 of that day's 110 pass-streak verifications were tunnels that no longer existed. A third-party crawler indexed several minutes after deletion.
We did not block it — it is a seller testing exactly the flow we document. Ephemeral hosts are now probed and served but excluded from listings_total, verified and hosts_total, and their badge reads ephemeral. Real new supply that day: one listing. This week: 11 to 17 a day, including a seller who submitted four routes by script on Monday night, deleted them, and resubmitted seven an hour later — which is the flow working.
4. The economy line: twice the addresses, the same money — then the best days yet
The index counts USDC paid on Base to the payTo addresses our listings declare. Cleaned of correction #16:
| UTC day | payments | payTos paid | payers | USD |
|---|---|---|---|---|
| 09-04 | 9,019 | 181 | 588 | $21,147 |
| 09-05 | 6,839 | 226 | 535 | $16,515 |
| 09-06 | 14,580 | 486 | 559 | $16,357 |
| 09-07 | 13,092 | 299 | 606 | $24,933 |
| 09-08 | 13,397 | 443 | 672 | $22,253 |
Saturday's import doubled the addresses we watch and the dollars did not move: 486 payTos paid on 09-06 for the same $16k. The money is concentrated. Then 09-07 and 09-08 were the two largest days we have recorded, with the widest payer counts — 672 distinct paying wallets on Monday. We do not know why yet, and we are not going to guess in this report.
Two things the line still cannot say: it does not yet split EIP-3009 x402 settlements from plain transfers (one gift-card seller's payTo takes ~$12.7k a day from 81 wallets, mostly plain), and it is Base only. Ninety Solana listings and every non-Coinbase facilitator are outside it by construction.
5. The verdict route: four days at the toll gate
A one-word paid verdict — pay, caution, avoid, with reasons — has been for sale at $0.005 per listing since 09-04. Results:
- Sales: 2, both ours. Gate log since 09-05: ~2,900 challenges from ~90 distinct clients; zero external sales; one external payment attempt, from a settlement-census bot paying the literal
:idtemplate, correctly refused and never charged. - Who is at the gate: twelve named prober and index projects, plus three paid-verification and search peers. All read the price. None pay. The unnamed clients (
node) went from 4 to 9 in four days and stopped. - Registered with the Bazaar (one consolidated template), x402scan (after two conformance rounds), Poncho and AgentCash.
Honest read: liveness probing is a commodity — twelve instruments showed up in one afternoon — and paid-delivery verification is not. The verdict is the directory's own x402 endpoint, its funnel gauge, and its ledger; it was not designed as revenue and it is not producing any. On Tuesday we added a second paid route aimed at the other side of the market, on-demand paid verification: $5 buys a real purchase from your endpoint within 24 hours, graded exactly as a wave, published whatever it finds. No requests yet.
6. The seam: what the wire says versus what the chain says
Paddock answers "should this agent pay this address" from the chain side, before the payment. We answer from the wire side, after it. On Tuesday's wave every purchase carried her verdict, stamped with her schema version, inside our attested record.
On 376 paid calls where Paddock said route: true, we received usable data on 337 — 89.6% agreement between a pre-spend verdict and a paid outcome, same endpoint, same minute. The 39 disagreements are all one shape — things a pre-payment check structurally cannot see:
| our outcome where she said route | n |
|---|---|
paid, answered, data too old to be what was advertised (stale_data) | 18 |
| paid, answered, not JSON | 13 |
| paid, answered, an error object | 4 |
| paid, rejected 400 | 2 |
| paid, 500 | 1 |
| paid, empty body | 1 |
There was no case where she said route and the payment itself failed. Seven more came back inconclusive from her and delivered for us — she errs conservative, which is the direction to be wrong in. Her schema moved from 0.5 to 0.6 during the run; every record is stamped, and nothing needed re-running.
The seam, with a number on it: she is right about payability; we are right about delivery; the gap is about 10% of paid calls. Anyone routing on either signal alone should know which 10% they are missing. One caveat we owe: this population was cheapest-first over never-purchased routes, so it skews toward small, simple endpoints. The next wave is built to test a different population (§8).
7. Manifests, and "there is no agreed word"
We surveyed /.well-known/x402 on 2,000 listed hosts: 1,265 (63%) serve a JSON manifest. Melchiorre Oliva ran the same question over his own pinned 1,521: 868 (57%). Two populations harvested differently, four points apart: the convention is far more widespread than the working-group thread assumed, almost certainly because the CDP templates emit one.
What they say about networks is another matter. Under any key mentioning network, chain or accepts: 319 of our 1,271 manifests name networks, under 30 different spellings (network 231, chain_id 57, network_name 53, networks 52, accepts 49, a tail of one-offs). His: 146 of 868 under 18. Only accepts entries — 16 hosts, 21 entries — carry a network beside scheme, asset and payTo, which is the shape every x402 parser already reads off the wire. So the proposal on the table is: the outer key is accepts, and per-version scoping (x402Versions) goes on each entry. Two hosts carry that shape today, deskcrew.io and us. We withdrew our earlier "51 declare" figure; it was a loose match.
8. Method-aware probing, one week on
Since 09-03 a GET that does not produce a 402 is retried as POST with an empty body. This week that retry ran on ~56,000 probes a day and passed ~49,000 of them — routes we would have called dead a week earlier. Twelve hosts carry ten or more POST-only listings each.
Auto-relist. A route delisted by a probe rule now stays in the queue for 30 days, at the back, and a parsed 402 brings it back. 32 are in the window; two came back in the rule's first five minutes. We had written that this rule "would have caught" twenty routes a peer relisted by hand last week. Checked against our own record: of his twenty, six were ours before Saturday, and one — vape.juxtaposition1.deno.net, recovered 09-01 — would have been caught. One of twenty, not twenty. He was right.
Same route, both right, different resolution. On the four live routes we and he both watched through August, three match to the day. On the fourth, binance-crypto-24h-stats, his once-a-day probe saw a 502 on 6 of 20 days (30%); at five-minute resolution we recorded 407 502s in 1,207 probes (34%). Same route, both right, and the rate agrees. On 08-11 that route answered 402 thirty-five times and 502 a hundred and fifty-five times inside one day. A streak rule sees that; a daily sample cannot. On the mixed days the resolution is the finding.
9. Demand, honestly
- The top search by count — 872 asks this week for "ecb reference conversion usd eur plain text" — is one seller watching its own rank (three clients, one operator, who also owns the listing it returns). It is not demand. Ranked by distinct clients, the real top query is
ukintel(8 clients), the UK company-data host that correction #17 made verifiable. - An agent running in VS Code used our MCP server to find an FX rate, and then judged the seller's "updated hourly" claim against ECB reality (reference rates publish once a day). Our record supplied the facts around it; the judgement was the agent's. That is the directory doing its job.
- MCP: 28 queries from 8 clients on 09-07, the highest day yet; sessions now run continuously, several tool calls at a time. The "queries" line counts searches only and undercounts sessions.
- First programmatic multi-listing claim (six routes in three minutes), first seller resubmitting its catalogue by script, and a peer's nightly catalogue read that announces its licence in its own user-agent.
10. Peers we now name
- Paddock — chain-side pre-spend verdicts; §6. Her
paidFulfillmentfield reads our census. - nsgoods — a weekly payability observatory that scanned our catalogue and wrote to us with numbers. Two rounds later: his "22.5% malformed 402" was a labelling flaw he corrected himself (989 of 1,003 return no 402 to a GET); we could answer the POST question directly — of our listings that answer no 402 to a GET, 360 of 391 in the 405 bucket, 332 of 428 in the 404 bucket, 95 of 184 answering 200 produce a clean 402 on POST; genuinely no challenge on either verb: 167. He also found a hole in our changes feed we could not have found ourselves: a seller adding a Solana rail beside an unchanged EVM payTo produces no event at all, because we store one network and one address per listing. 77 of his 125 "wallet changes" are invisible to us by construction. Fixing it means storing the whole
accepts[]per listing; it is on the list, and we are not going to pretend a flag covers it. He also sent 85 Solana listings whose receiving token accounts never existed or were closed, checked at slot 445145988. We cannot verify that; it is his finding, and the multichain gap quantified by someone else's instrument. - ScoutScore, x402.direct, ShortForge — paid-verification and search peers we will list and grade like anyone else where they charge.
- Glama has inspected our MCP server daily since 09-04 and grades it A. A third party keeping a longitudinal record of us is a pleasant symmetry.
11. What we will not claim
- Whether the two record days on the economy line are a trend.
- A paid-delivery rate by transition type. 23 of the 1,020 paid-verified routes have been bought twice; none across a change. The ledger is wide, not deep. Next week's wave buys a second time from routes that changed since their first purchase, and then there will be a number.
- Anything about Solana, or about any facilitator we do not index.
- A rule for hosts that refuse our prober permanently. One host, twelve neutral probes, no rule yet.
Figures, this morning
| listings (active, probed, not ephemeral, not testnet) | 4,933 |
| hosts | 1,884 |
| verified | 4,549 |
| probes, 24h | 283,656 |
| probe pass rate, 24h (four neutral classes excluded) | 94.0% |
| re-probe cadence | 25.7 min (promise: 30) |
| endpoints paid-verified (delivered on latest attempt) | 1,020 of 2,550 attempted |
| pass rate on payment-required endpoints | 41.4% (last week 54.8%; §1) |
| paid-verification waves this week | 1 (wave 7: 1,461 attempts, 383 paid, 344 delivered, $0.99) |
| Paddock agreement on paid calls | 89.6% (337 / 376) |
| verdict sales | 2 ($0.01), both ours |
| corrections since launch | 17; three this week |
| listings with on-chain section | 4,638 (1,726 partial: watched since 09-05) |
Twin: week-2026-09-09.json. Method: /methodology. Every figure above is recomputed from the live database at publication.