Public MCP Score · methodology

How we rank public MCP servers from the outside.

Continuous validation from residential networks, a published scoring formula, and explicit limits on what the score claims. See the live scores for current rankings.

What the Score measures

Outside-in MCP reliability, not a catalog score.

The Public MCP Score measures public MCP servers the way a client would: connect, discover tools, invoke them, and record whether the session held up from a residential network path.

This is not a popularity list and not a lab capability benchmark. It is continuous validation of production MCP surfaces, scored with public math.

See also MCP validation for the product surface, and the live scores for current rankings.

01

Reach the MCP endpoint

Confirm the server is reachable from a residential origin.

02

Confirm transport / session

Exercise the MCP transport path the client would use.

03

Discover tools

List tools and verify discovery completes cleanly.

04

Invoke tools + score results

Call tools and score invocation and functional outcomes.

05

Aggregate into 30d Score inputs

Roll per-run evidence into the public composite score.

Why outside-in

What datacenter checks miss.

MCP servers often look healthy from a datacenter laptop and fail for the clients that actually matter: agents and apps on residential or carrier networks, behind real DNS and TLS paths, with real latency tails.

Geo

Geo / IP filtering

Server allows cloud ranges but blocks consumer ISPs — so a datacenter check stays green while real clients never connect.
SSE

Transport fragility

SSE / streamable HTTP works on a fat pipe, then stalls on residential paths with different buffering and idle timeouts.
P95

Tool timeouts

P50 looks fine in the lab; P95 from real regions blows past client budgets and agent tool-call deadlines.
½

Partial discovery

Handshake succeeds while tool lists or invocations fail intermittently — half a protocol session is still a failed client.

Score probes run on the same distributed residential network used for AgentStatus agent validation. The backend orchestrates; nodes execute the workloads.

Signals

What each column answers.

Each public column answers one plain question — and has an explicit non-claim:

Uptime (30d)

What share of MCP validation runs reached the server?

Non-claimDoes not prove tools returned useful results.

Latency P95

How slow were the slowest successful probes from residential origins?

Non-claimDoes not capture every client region equally on every run.

Tool success

Did tool calls invoke cleanly and look functionally correct?

Non-claimProbe coverage is finite; not every tool on every server is exercised every run.

Protocol conformance

Did handshake, transport, discovery, and invocation behave like MCP?

Non-claimDoes not certify full MCP spec compliance for every optional feature.

Score

Score calculation.

The Score is a composite metric from 0 to 100:

Score = weighted(uptime, latency, tool_success, conformance) + verdict_bonus
Uptime35%

reachable_runs / total_mcp_runs over the last 30 days. Run-based reachability — not a continuous synthetic heartbeat.

Latency25%

P95 response time mapped to 0–100. Perfect score at ≤200 ms; linear decay to 0 at ≥2000 ms.

Tool success25%

invocation_rate × 0.6 + functional_rate × 0.4. Falls back to whichever rate is available if only one exists.

Protocol conformance15%

Average per-run score: reachable 30% + transport 25% + discovery 25% + invocation 20%. Missing invocation evidence renormalizes the rest.

Verdict bonus

+5

Operational

Current status healthy — small bonus for an actively good surface.
+2

Degraded

Partial impairment — still reachable enough to score, not fully healthy.
+0

Down

Not earning a status bonus. Reachability and other signals still contribute to the base score.

If a component is missing, remaining weights renormalize. The verdict bonus is still applied. Final score is capped at 100.

Ranking rules

Eligibility, ranking & sample size.

Indexed

Servers explicitly opted into the public MCP index and actively monitored.

Hidden

Paused monitors are excluded from the public leaderboard.

Ranked

Any server with at least 1 scored validation run. Sort uses a confidence-adjusted Score (shrinkage toward a neutral prior).

Displayed Score

Raw composite score — what you see in the drawer and score column.

Low sample

Fewer than 7 runs — badge only. Still ranked, but dampened so thin evidence cannot dominate.

Unrated

Zero scored runs in the window (rare for listed servers).

Identity

Public rows use slug + display name only. No agent UUIDs.

Freshness

Cadence & data freshness.

SurfaceFreshness
Validation runsPer monitor schedule (public MCP fleet commonly ~6h; configurable down to 5 min)
Live leaderboardComputed from the latest 30-day window of MCP validation results
Trend deltaVs yesterday’s daily Score snapshot (00:15 UTC). Null until day 2 of snapshots
Validation runs counterReal 30-day run count for indexed servers — not a marketing session total

The split

What we are, and are not.

We are not

  • A paid placement directory
  • A full MCP spec certification body
  • An inside-out APM / tracing product

We are

  • An independent outside-in validator
  • A public reliability ranking with published math
  • The client’s perspective at residential network scale

FAQ

Common questions.

QuestionAnswer
Is this a lab benchmark of MCP servers?No. We probe production MCP endpoints from residential nodes on a schedule, the same network class a real client would use. scores reliability in the wild, not capability in a controlled room.
What goes into the score?Four weighted signals plus a small verdict bonus: uptime (35%), latency (25%), tool success (25%), and protocol conformance (15%). Missing components renormalize the remaining weights. The formula is public.
How is uptime measured?Run-based reachability over the last 30 days: reachable MCP validation runs divided by total MCP validation runs for that server. This matches the monitor’s schedule (public MCP monitors are commonly ~every 6 hours; paid agents can schedule down to 5 minutes), not a fake continuous heartbeat.
What does tool success mean?A blend of invocation success (did the tool call complete?) at 60% and functional success (did the result look correct for the probe?) at 40%. If only one signal exists, we use that one.
What is protocol conformance?Per-run adherence to the MCP surface we exercise: reachability, transport support, tool discovery, and invocation. Those pieces are weighted 30% / 25% / 25% / 20% and averaged across runs.
Why are some servers marked Low sample?Fewer than 7 validation runs in the scoring window. They still get a raw score and a leaderboard rank, but ranking uses a confidence-adjusted score so a lucky 2-run streak cannot outrank a proven fleet. Unrated means zero scored runs.
Can a server pay for a better rank?No. Rankings come from measured results only. Paused monitors are hidden from the public index entirely.
Do you publish agent UUIDs?No. The public index uses display names and stable slugs only.

Independence

Trust is measured, not purchased.

The Public MCP Score is operated by Carmel Labs / AgentStatus as an independent validator. Rankings are determined by measured validation results. Server operators cannot pay for a better score.

The only way to move up is to improve reachability, latency, tool behavior, and protocol conformance under outside-in probes.

Public MCP Score

Get listed.

Public MCP servers are indexed and monitored for free. Submit yours — we create a pilot dashboard and email you access.