6-min read

Carmel Labs · AgentStatus · June 2026

The Two Failures Hiding in LLM-as-a-Judge

Calibration problems shrink with better technique. Competence problems do not. Why the dominant evaluation paradigm has a structural ceiling, and the two methods older than language models that get past it.

agentstatusagentstatus.dev | June 2026

More research

Continue reading

August 2026

The MCP Protocol Era Report

Week-one refresh across 2,481 Index servers. Legacy still ~50%. Modern ~42%. Holdouts are filterable on the live Index.

Read report

July 2026

You Can't Measure What You Can't Reach

Reachability and reliability are different problems. Almost every tool solves at most one — valid validation needs both.

Read report

July 2026

"But My Agent Can Be Wrong and Still Pass"

Two kinds of wrong. We catch the one that gives itself away through instability and rule-violation, and we're honest about the residue no user-side method catches without ground truth.

Read report

July 2026

The Residentiality Report

The datacenter blind spot: reachability from where users are — and the hard limit on what user-side can claim.

Read report

March 2026

The State of AI Agent Reliability

We monitored 3,260 production AI agents across 48 countries. 89% with perfect uptime scored 0% on quality. The full data is inside.

Read report

April 2026

The State of AI Agent Drift

88% of agents started giving worse answers at least once in 30 days. A look at how production AI agents drift - and the systemic March 29 event.

Read report

April 2026

The Anti-Synthetic Monitoring Thesis

Datacenter vs residential reachability across 6,228 matched agents — access asymmetry, not correctness-by-origin.

Read report

Research

9 Businesses You Can Build on Agent Behavioral Data

Insurance underwriting, credit scores, compliance certification, procurement intelligence - the commercial layer that sits on top of continuous agent monitoring.

Read report