Back to website
2-min read

AgentStatus × Scale AI, a quick map of how we fit

Independent verification for Scale's deployed agents.

We do two jobs: reachability from residential networks (past CDN/WAF), then reliability once reached — gold/contract and consistency checks, across the channels each platform supports, from 2,500+ nodes across 70 countries. We sit alongside Scale's Generative AI Platform, Data Engine, and SEAL benchmarks. We don't replace them.

22M
tests
8,000+
agents
2,500+
residential devices
70
countries
agentstatusagentstatus.dev | partner brief

What we understand about Scale AI

A full-stack platform for building, evaluating, and aligning AI systems.

Scale's Generative AI Platform lets enterprises build, evaluate, and control AI agents end-to-end. The Scale Data Engine powers the data work, collection, curation, annotation, RLHF, that goes into the world's leading models. Scale Labs and the SEAL (Safety, Evaluation and Alignment Lab) ship benchmarks like Humanity's Last Exam, SWE Atlas, and Audio MultiChallenge to test models against rigorous, multi-turn standards.

Scale's customers include OpenAI, Meta, Microsoft, Cisco, the U.S. Department of Defense, and the Government of Qatar. The work is concentrated where reliability matters most, defense, healthcare, financial services, and frontier model development, with a stated mission of building reliable AI systems for the world's most important decisions.

What AgentStatus is

We measure whether users can reach the agent, then whether it still passes its checks.

Reachability. Controlled validations from 2,500+ residential devices across 70 countries measure whether users can open the agent the way they do — past CDN, WAF, and bot walls. Multi-geo is observer vantage for access and last-mile latency — not answer localization by probe IP, and not agent tool egress.

Reliability. Once reachable, we run gold/contract checks when truth exists, plus rephrase, drift, and policy consistency probes when it doesn't. Dual LLM-as-judge scores open answers with a known ceiling — stably wrong but consistent still needs a domain expert.

That includes multi-turn conversations and multi-agent journeys when customer paths span tools, escalations, and handoffs. It supports governance and risk conversations when stakeholders ask what was tested, from where, and what changed.

Outside-in validation is two separate jobs

Reachability

Claims agent · residential
Status dashboard with reachability verdict and regional coverage

Outcome

Residential path

Your monitor hits the VIP lane. Users hit the WAF. Datacenter checks get blocked, throttled, or allowlisted. Residential observers take the inbound path customers take — so “up” means reachable from home networks, not from AWS.

Reliability

Claims agent · eval
Answer quality dashboard with evaluation prompts and pass fail results

Outcome

Answer quality

Reachable and self-contradicting is still broken. Rephrase flips, drift, and policy breaks need no ground truth. Gold and dual judges cover the rest when truth exists. Uptime grades none of that.

Where we fit

We sit beside the platform. We do not replace it.

01

Benchmarks vs deployed behaviour

SEAL benchmarks measure how a model performs against rigorous test sets at evaluation time. AgentStatus answers a different question: what did the deployed agent actually do for a user-like validate in production, from a residential observer vantage — and, once reachable, did the reply diverge from gold/contract or consistency checks? Observer geo measures inbound access past CDN/WAF; it does not localize tool egress or answers by probe IP.

02

Eval-time truth vs production drift

A frontier model that scores 55% on Audio MultiChallenge is a known quantity at eval time. The same model wrapped in an agent, three weeks later in production, often is not — either because residential observers can't reach it the way users do, or because answers drift against expected checks once they can. Distributed validate traffic catches both classes before they become incidents.

03

Global execution footprint

2,500+ nodes across 70 countries is the proof we are not 'synthetic from a single cloud region.' It matters for Scale's enterprise and government customers operating across regulated geographies, and for access failures that only reproduce from specific residential locations or ISPs — distinct from answer-quality checks once the agent is reached.

04

Partner-friendly integration posture

We do not assume we can 'discover' a customer agent the way some web-widget vendors can be scraped. Credential-based surfaces (chat endpoints, voice numbers, agent APIs, sandbox releases) and customer-approved monitoring are the right model, aligned with the trust posture Scale's defense and enterprise customers require.

The split

How the work divides

How the work divides

Their platform

Scale AI, Evaluation
  • Generative AI Platform
  • Data Engine & RLHF
  • SEAL benchmarks
  • Red Teaming & adversarial testing
  • SWE Atlas & coding evals

Outcome

System of record

Dashboards, exports, lifecycle tools, and orchestration remain theirs. We do not replace that surface.

AgentStatus

AgentStatus, Post-deploy
  • Continuous validate traffic
  • Expected-answer checks & drift detection
  • Multi-turn / multi-agent journeys
  • Real-network execution evidence
  • 2,500+ nodes across 70 countries

Outcome

Outside-in layer

Residential inbound path past CDN/WAF, then gold, consistency, and scoped judges once the agent is reachable.

Proof of scale

Auditable scale metrics

In about two months, we have executed on the order of 18 million validate runs across the network. We also maintain on the order of 6,000 agent records in our system, meaning rows/configurations we track, including evaluation and pipeline agents, not "6,000 paying customers."

If helpful, we can share stricter production-only definitions under NDA.

What we are not claiming

We are an independent layer that runs alongside your stack.

We are not a replacement for Scale's Generative AI Platform, Data Engine, or SEAL benchmarks. We are an independent layer that can coexist with them, and, where useful, help teams correlate outside-in validate outcomes with eval-time benchmark performance, so the gap between "passes the benchmark" and "behaves correctly in production" can be measured and closed.

What we'd like from this conversation

These three asks would move a pilot forward.

01

A 2-week sandbox pilot

A joint customer scenario, particularly in a regulated vertical, with a set of agreed prompts, expected answers, and a 2-week evaluation window. SEAL benchmark results and AgentStatus validate outcomes tell one story together across the eval-to-production boundary.

02

Security and procurement posture

How AgentStatus should connect in a way that satisfies enterprise and government security reviews. Data handling, least privilege, audit evidence, and clear test-traffic boundaries.

03

Where independent proof is most useful

Whether the right starting point is plugging into the Generative AI Platform itself, a joint customer engagement, or both.

Closing

Scale helps frontier teams build, evaluate, and align AI systems for the world's most important decisions.

AgentStatus helps those same teams prove, continuously, that the deployed agent behaves the way policy and customers require, globally, with evidence that holds up under scrutiny.

Contact·dulra@carmel.so·roman@carmel.so

Metrics are stated with explicit definitions: validate runs are scheduled executions over ~two months; agent records are database rows, not revenue customers. Public Scale AI references above reflect Scale's public product pages, research output, and customer disclosures as of the date of this note.