6-min read

Carmel Labs · AgentStatus · June 2026

The Two Failures Hiding in LLM-as-a-Judge

Calibration problems shrink with better technique. Competence problems do not. Why the dominant evaluation paradigm has a structural ceiling, and the two methods older than language models that get past it.

agentstatusagentstatus.dev | June 2026

More research

Continue reading