How we rank public MCP servers from the outside.
Continuous validation from residential networks, a published scoring formula, and explicit limits on what the score claims. See the live scores for current rankings.
What the Score measures
Outside-in MCP reliability, not a catalog score.
The Public MCP Score measures public MCP servers the way a client would: connect, discover tools, invoke them, and record whether the session held up from a residential network path.
This is not a popularity list and not a lab capability benchmark. It is continuous validation of production MCP surfaces, scored with public math.
See also MCP validation for the product surface, and the live scores for current rankings.
Reach the MCP endpoint
Confirm the server is reachable from a residential origin.
Confirm transport / session
Exercise the MCP transport path the client would use.
Discover tools
List tools and verify discovery completes cleanly.
Invoke tools + score results
Call tools and score invocation and functional outcomes.
Aggregate into 30d Score inputs
Roll per-run evidence into the public composite score.
Why outside-in
What datacenter checks miss.
MCP servers often look healthy from a datacenter laptop and fail for the clients that actually matter: agents and apps on residential or carrier networks, behind real DNS and TLS paths, with real latency tails.
Geo / IP filtering
Transport fragility
Tool timeouts
Partial discovery
Score probes run on the same distributed residential network used for AgentStatus agent validation. The backend orchestrates; nodes execute the workloads.
Signals
What each column answers.
Each public column answers one plain question — and has an explicit non-claim:
Uptime (30d)
What share of MCP validation runs reached the server?
Latency P95
How slow were the slowest successful probes from residential origins?
Tool success
Did tool calls invoke cleanly and look functionally correct?
Protocol conformance
Did handshake, transport, discovery, and invocation behave like MCP?
Score
Score calculation.
The Score is a composite metric from 0 to 100:
reachable_runs / total_mcp_runs over the last 30 days. Run-based reachability — not a continuous synthetic heartbeat.
P95 response time mapped to 0–100. Perfect score at ≤200 ms; linear decay to 0 at ≥2000 ms.
invocation_rate × 0.6 + functional_rate × 0.4. Falls back to whichever rate is available if only one exists.
Average per-run score: reachable 30% + transport 25% + discovery 25% + invocation 20%. Missing invocation evidence renormalizes the rest.
Verdict bonus
Operational
Degraded
Down
If a component is missing, remaining weights renormalize. The verdict bonus is still applied. Final score is capped at 100.
Ranking rules
Eligibility, ranking & sample size.
Indexed
Hidden
Ranked
Displayed Score
Low sample
Unrated
Identity
Freshness
Cadence & data freshness.
| Surface | Freshness |
|---|---|
| Validation runs | Per monitor schedule (public MCP fleet commonly ~6h; configurable down to 5 min) |
| Live leaderboard | Computed from the latest 30-day window of MCP validation results |
| Trend delta | Vs yesterday’s daily Score snapshot (00:15 UTC). Null until day 2 of snapshots |
| Validation runs counter | Real 30-day run count for indexed servers — not a marketing session total |
The split
What we are, and are not.
We are not
- • A paid placement directory
- • A full MCP spec certification body
- • An inside-out APM / tracing product
We are
- • An independent outside-in validator
- • A public reliability ranking with published math
- • The client’s perspective at residential network scale
FAQ
Common questions.
| Question | Answer |
|---|---|
| Is this a lab benchmark of MCP servers? | No. We probe production MCP endpoints from residential nodes on a schedule, the same network class a real client would use. scores reliability in the wild, not capability in a controlled room. |
| What goes into the score? | Four weighted signals plus a small verdict bonus: uptime (35%), latency (25%), tool success (25%), and protocol conformance (15%). Missing components renormalize the remaining weights. The formula is public. |
| How is uptime measured? | Run-based reachability over the last 30 days: reachable MCP validation runs divided by total MCP validation runs for that server. This matches the monitor’s schedule (public MCP monitors are commonly ~every 6 hours; paid agents can schedule down to 5 minutes), not a fake continuous heartbeat. |
| What does tool success mean? | A blend of invocation success (did the tool call complete?) at 60% and functional success (did the result look correct for the probe?) at 40%. If only one signal exists, we use that one. |
| What is protocol conformance? | Per-run adherence to the MCP surface we exercise: reachability, transport support, tool discovery, and invocation. Those pieces are weighted 30% / 25% / 25% / 20% and averaged across runs. |
| Why are some servers marked Low sample? | Fewer than 7 validation runs in the scoring window. They still get a raw score and a leaderboard rank, but ranking uses a confidence-adjusted score so a lucky 2-run streak cannot outrank a proven fleet. Unrated means zero scored runs. |
| Can a server pay for a better rank? | No. Rankings come from measured results only. Paused monitors are hidden from the public index entirely. |
| Do you publish agent UUIDs? | No. The public index uses display names and stable slugs only. |
Independence
Trust is measured, not purchased.
The Public MCP Score is operated by Carmel Labs / AgentStatus as an independent validator. Rankings are determined by measured validation results. Server operators cannot pay for a better score.
The only way to move up is to improve reachability, latency, tool behavior, and protocol conformance under outside-in probes.

