Methodology 1.0 · schema 2026-08-11 · archive through 2026-08-12

Evidence first. A grade only when supported.

MCP Scores ranks the evidence observed for an MCP server or x402 endpoint. It is not a probability, endorsement, security audit, or guarantee of future behavior.

The publication floor, in full
Archive days required7
Successful observations required3
Archive coverage held2 day(s)
Scale0–1000
State below the floorNR · no number
Both conditions must be met. Neither can be waived, bought, or estimated.

The method in four decisions.

Start here. The full model, thresholds, weights, constraints, caveats, and correction process follow below.

Can it be reached?

Calls test whether the listed infrastructure answers and whether its protocol response is valid.

What evidence exists?

Identity, tenure, stability, tools, payment terms, and observable financial activity are recorded where applicable.

Does the record clear the floor?

Without 7 archive days and 3 successful observations, the public state remains NR with no number.

Do hard constraints apply?

Recent silence, disappearance, or material changes can cap the result after the weighted blend.

§ 01

The scale, drawn to size.

Band widths are the model's real score ranges, and the tally under each band is the live count of listings published in it right now.

MCP Scores grade scale, 0 to 1000 Grade E covers 0 to 399 with 0 published listings; Grade D covers 400 to 549 with 0 published listings; Grade C covers 550 to 699 with 0 published listings; Grade B covers 700 to 849 with 0 published listings; Grade A covers 850 to 1000 with 0 published listings Grade E: scores 0 to 399. 0 published listing(s). E 0 0 published Grade D: scores 400 to 549. 0 published listing(s). D 400 0 published Grade C: scores 550 to 699. 0 published listing(s). C 550 0 published Grade B: scores 700 to 849. 0 published listing(s). B 700 0 published Grade A: scores 850 to 1000. 0 published listing(s). A 850 0 published

A ≥ 850 · B ≥ 700 · C ≥ 550 · D ≥ 400 · E < 400 · NR = not rated

Grade E deliberately occupies the bottom 40% of the scale. Under Model 1.0 an endpoint that cannot be reached does not collect a middling grade; the availability rules below force it down.

§ 02

What is observed.

Archive coverage today: 2 day(s), 2026-08-11 → 2026-08-12. Every window claim on this site is read from that record, never assumed.

Unpaid probes

The observatory calls listed x402 endpoints on a repeating schedule and archives the HTTP status, the round-trip latency, and whether an unpaid request produced a structurally valid x402 payment challenge. Probes are unpaid by design: they measure whether the listing behaves as listed, not whether a paid call returns correct data.

What counts as a qualifying response

A qualifying response is 2xx, 3xx, or 402 Payment Required. For an x402 endpoint answering an unpaid probe, 402 is the correct protocol answer, so it counts. Every other 4xx is a refusal and does not qualify — 400, 401, 403, 404, 405, 410 and 429 all count against Reliability, as does every 5xx. An endpoint that refuses every caller cannot earn a high Reliability component.

MCP servers

Registry-listed MCP servers are measured by completing an initialize handshake and calling tools/list. A server that answers with an auth challenge is recorded as reachable but un-inventoried; a server that does not answer at all is recorded as such. Registry metadata — publisher, creation date, update date — supplies identity and tenure evidence.

Money

Where a listing declares a payee wallet, inbound USDC transfers to that address on Base are counted inside a disclosed block range, along with the number of distinct payers and the share taken by the largest one. Observed inflow is never labelled revenue, and a declaration in listing metadata does not establish legal ownership of an address.

§ 03

Components and weights.

Five components, each measured 0–100, combined into the 0–1000 composite. Hosts and MCP servers weight differently because financial evidence only exists where a payment wallet is declared.

Componentx402 hostMCP server
R Reliability 30% 40%
T Tenure 20% 25%
F Financial evidence 20% 0%
S Stability 15% 20%
I Identity 15% 15%

Reliability

Recency-weighted share of probes that produced a qualifying response, with a 0.9 decay per observation so recent behaviour dominates.

Tenure

Log-scaled observed survival, plus registry age for MCP servers and seller-owned domain registration age for hosts. Shared platform domains earn no tenure credit — a listing on someone else's domain does not inherit their history.

Financial evidence

Parseable payment terms, then observed inflow, payer diversity and concentration. Credit is cut to 30% when the payer set is narrow, the top payer dominates, or the wallet is shared across listings.

Stability

Distinct archive days on which the listing's metadata changed. A declared-payee change triggers a time-bounded score cap.

Identity

Registry namespace control, domain evidence, published documentation, and observable identity artifacts.

Technical composite formula

Each applicable component is measured from 0–100. The pre-constraint composite is the rounded weighted sum multiplied by ten:

round(10 × (R·wR + T·wT + F·wF + S·wS + I·wI))

For MCP servers, financial evidence has zero weight and the remaining weights sum to 100%. Hard constraints are applied only after this blend.

§ 04

The publication floor.

A listing stays NR until the archive holds at least 7 days of coverage and at least 3 successful observations for it.

Below the floor there is no grade and no number. Not a provisional figure, not a greyed-out score, not a hint. The scorecard, this site's lists, the JSON API, the MCP tools and the badge all withhold the same figure at the same moment, because they all call the same function.

Confidence is reported separately and never changes the score: low below 3 successful observations, medium from 3 to 30, high above 30.

Publication floor progress Archive coverage 2 of 7 required days. Successful observations 2 of 3 required. ARCHIVE DAYS 2 / 7 OBSERVATIONS 2 / 3

Archive-wide coverage against the floor, and the best-covered candidate currently evaluated.

Published right now: 0

No evaluated listing currently clears the publication floor, so there is nothing to rank. This is an accurate empty result, not a failure: MCP Scores does not publish a grade until the floor is met, and it does not invent one to fill this list.

§ 05

Hard constraints that override the blend.

Some observations are disqualifying regardless of how the components score. These caps are applied after the weighted composite.

ObservationEffect
No HTTP response on the 7 most recent probesPublic grade forced to E (score capped at 399)
No active registry entry in the latest catalogue snapshot, with a recorded gone eventPublic grade forced to E under the listing-availability rule
No HTTP response on the 3 most recent probesScore capped at 399
No probe served a qualifying responseScore capped at 549
Declared payee wallet changed within 60 daysScore capped at 549 for 60 days
MCP server not answering and not auth-gatedScore capped at 399
§ 06

What a score cannot prove.

  • Unpaid probes measure reachability and protocol behaviour. They do not prove that a paid call returns correct data.
  • Observed on-chain inflow is not revenue, and not every transfer is an x402 sale.
  • A grade is not a security audit, a performance guarantee, or a recommendation to transact.
  • Absence of a record means the listing has not been observed — not that it is unsafe.
  • The score is an ordinal Model 1.0 output, not a calibrated probability of failure.

Pending model inputs

  • Cross-wallet funding-source analysis
  • Cohort survival percentiles, once the archive is deep enough to support them
  • Wallet-funded delivery verification (paid calls, not unpaid probes)
  • 90-day probability calibration and a quarterly published backtest

Model changes are versioned. A score computed under a later model is labelled with that model version, and prior records are not retroactively rewritten.

§ 07

Corrections and appeals.

Operators may contest any measured statement. Write to corrections@mcpscores.com with the listing identifier, the exact statement at issue, the supporting evidence, and a reply address.

What happens to an accepted correction. MCP Authority records it in the observatory action log against the listing identifier. Every scorecard reads that log and renders accepted corrections as dated annotations — in the HTML page, in the corrections_applied array of the JSON payload, and in the MCP tool response. Nothing is silently rewritten and no annotation is removed once published.

A correction changes the observed record. It does not buy a grade: the same model runs on the corrected inputs.

Observed, not asserted

Every figure traces to an archived observation with a date and a source table.

Read-only

This public site computes from the record and never mutates it.

Correctable

Disputes are answered on the public record, dated and attributed.

No pay-to-rate

No operator can purchase, accelerate, or suppress a rating.