Back to website
2-min read

AgentStatus × AIUC

AIUC-1 certifies. AgentStatus verifies, every day.

AIUC-1 sets the bar at certification. We make sure your agents are still clearing it on day fifty, with scheduled outside-in tests from 2,500+ real consumer devices across 70 countries, checking that answers stay correct after the audit date.

22M
tests
8,000+
agents
2,500+
residential devices
70
countries
agentstatusagentstatus.dev | partner brief

Why we're reaching out

Certification is point-in-time. Reliability is continuous.

AIUC-1 is the first AI agent standard with real teeth, six pillars, quarterly updates, MITRE as a technical contributor, Schellman as the independent auditor. UiPath's certification required 2,000+ technical evaluations at audit time.

The Reliability pillar, agents behaving predictably and consistently in production, is exactly what we measure continuously. Not at certification. Every day. From outside the customer's stack, on the same networks their users are on.

That's the gap we'd like to fit into: between the audit and the renewal, generating the evidence that the bar is still being cleared.

What AgentStatus is

We continuously test AI agents and check answers with gold/contract and consistency probes (scoped — stably wrong but consistent still needs a domain expert).

Scheduled tests run from an independent network of 2,500+ consumer devices in 70 countries, with expected answers and drift detection when behaviour slips. Repeatable proof from residential observer networks — inbound reach past CDN/WAF — not from two AWS regions.

That includes multi-turn conversations and tool handoffs, so you can show what was exercised, from where, and what changed week over week.

Outside-in validation is two separate jobs

Reachability

Claims agent · residential
Status dashboard with reachability verdict and regional coverage

Outcome

Residential path

Your monitor hits the VIP lane. Users hit the WAF. Datacenter checks get blocked, throttled, or allowlisted. Residential observers take the inbound path customers take — so “up” means reachable from home networks, not from AWS.

Reliability

Claims agent · eval
Answer quality dashboard with evaluation prompts and pass fail results

Outcome

Answer quality

Reachable and self-contradicting is still broken. Rephrase flips, drift, and policy breaks need no ground truth. Gold and dual judges cover the rest when truth exists. Uptime grades none of that.

Where we fit

Certification and continuous outside-in evidence are different jobs.

How the work divides

Their platform

AIUC, Their lane
  • <strong>Defines the standard</strong> (AIUC-1, six pillars)
  • <strong>Audit + certificate</strong> via Schellman
  • <strong>Insurance backstop</strong> when things go wrong

Outcome

System of record

Dashboards, exports, lifecycle tools, and orchestration remain theirs. We do not replace that surface.

AgentStatus

AgentStatus, Our lane
  • <strong>Generates the evidence</strong>, continuously
  • <strong>Outside-in tests</strong> between audit dates
  • <strong>Prevention signal</strong>: catches drift before it becomes a claim

Outcome

Outside-in layer

Residential inbound path past CDN/WAF, then gold, consistency, and scoped judges once the agent is reachable.

One sentence. Certification answers "did we meet the bar then?" AgentStatus answers "is it still true this week, from real places on the internet?"

Proof of scale

Here is what we have already run, with plain definitions.

~10M test runs in ~2 months across the network. 8,000+ agents being tracked, including ones from companies you'd recognise (specifics under NDA).

We've also caught node operators trying to game the network with datacenter VMs instead of real consumer devices, the same kind of adversarial behaviour AIUC-1 is designed to make harder. Detection is built into the product.

What we are not claiming

We are an independent layer that runs alongside your stack.

We are not AIUC-1 auditors, not AIUC-1, and not an insurance company. We do not replace AIUC's standard or policies. We're the continuous evidence layer that sits between them.

What we'd like from this conversation

We would like to start with a two-week sandbox pilot.

01

One certified or candidate agent

A surface AIUC has worked with, Intercom, Ada, ElevenLabs, or another, with one agreed set of expected answers.

02

A fixed 2-week window

We run scheduled outside-in checks and share pass/fail rates, drift events, and geography/network-split results.

03

A 30-minute readout

Does this belong next to AIUC-1 as ongoing evidence between audits?

Closing

AIUC gives enterprises reason to sign.

AgentStatus helps them keep the story true in production, every day, from residential observer networks users actually use.

Contact·dulra@carmel.so·roman@carmel.so

"Test runs" and "agent rows" mean what we said above. AIUC descriptions are from public pages and announcements, not an endorsement by AIUC.