AgentStatus × AIUC
AIUC-1 certifies. AgentStatus verifies, every day.
AIUC-1 sets the bar at certification. We make sure your agents are still clearing it on day fifty, with scheduled user-side tests from 2,500+ real consumer devices across 70 countries, checking that answers stay correct after the audit date.
Why we're reaching out
Certification is point-in-time. Reliability is continuous.
AIUC-1 is the first AI agent standard with real teeth, six pillars, quarterly updates, MITRE as a technical contributor, Schellman as the independent auditor. UiPath's certification required 2,000+ technical evaluations at audit time.
The Reliability pillar, agents behaving predictably and consistently in production, is exactly what we measure continuously. Not at certification. Every day. From outside the customer's stack, on the same networks their users are on.
That's the gap we'd like to fit into: between the audit and the renewal, generating the evidence that the bar is still being cleared.
What AgentStatus is
We continuously test AI agents and check answers with gold/contract and consistency probes (scoped — stably wrong but consistent still needs a domain expert).
Scheduled tests run from an independent network of 2,500+ consumer devices in 70 countries, with expected answers and drift detection when behaviour slips. Repeatable proof from residential observer networks — inbound reach past CDN/WAF — not from two AWS regions.
That includes multi-turn conversations and tool handoffs, so you can show what was exercised, from where, and what changed week over week.
User-side validation is two separate jobs
Reachability

Outcome
Can we talk to it?
Residential observers take the inbound path customers take — past CDN, WAF, and bot walls that treat datacenter synthetics differently. Monitoring asks: is it healthy right now? Reliability asks: does it keep working over time? “Up” means reachable from home networks, not from AWS.
Outcome verification

Outcome
Did it do the right thing?
Reachable and wrong is still broken. Scenario — did it finish the job? Compositional — do the pieces hold together? Safety — must-not-say, policy, attack probes. Stability — same ask, same story? Scores: Consistency and Drift. Not prose matching. Optional sample review is corroboration only.
Where we fit
Certification and continuous user-side evidence are different jobs.
Your platform
- • <strong>Defines the standard</strong> (AIUC-1, six pillars)
- • <strong>Audit + certificate</strong> via Schellman
- • <strong>Insurance backstop</strong> when things go wrong
Outcome
System of record
Dashboards, exports, lifecycle tools, and orchestration remain yours. We do not replace that surface.
AgentStatus
- • <strong>Generates the evidence</strong>, continuously
- • <strong>User-side tests</strong> between audit dates
- • <strong>Prevention signal</strong>: catches drift before it becomes a claim
Outcome
User-side layer
Reachability (Monitoring / Reliability) from residential networks, then outcome verification once reached — Scenario, Compositional, Safety, Stability; Consistency and Drift scores. Not prose matching.
One sentence. Certification answers "did we meet the bar then?" AgentStatus answers "is it still true this week, from real places on the internet?"
Proof of scale
Here is what we have already run, with plain definitions.
~10M test runs in ~2 months across the network. 8,000+ agents being tracked, including ones from companies you'd recognise (specifics under NDA).
We've also caught node operators trying to game the network with datacenter VMs instead of real consumer devices, the same kind of adversarial behaviour AIUC-1 is designed to make harder. Detection is built into the product.
What we are not claiming
We are an independent layer that runs alongside your stack.
We are not AIUC-1 auditors, not AIUC-1, and not an insurance company. We do not replace AIUC's standard or policies. We're the continuous evidence layer that sits between them.
What we'd like from this conversation
We would like to start with a two-week sandbox pilot.
One certified or candidate agent
A surface AIUC has worked with, Intercom, Ada, ElevenLabs, or another, with one agreed set of expected answers.
A fixed 2-week window
We run scheduled user-side checks and share pass/fail rates, drift events, and geography/network-split results.
A 30-minute readout
Does this belong next to AIUC-1 as ongoing evidence between audits?
AIUC gives enterprises reason to sign.
AgentStatus helps them keep the story true in production, every day, from residential observer networks users actually use.
Contact·dulra@carmel.so·roman@carmel.so
"Test runs" and "agent rows" mean what we said above. AIUC descriptions are from public pages and announcements, not an endorsement by AIUC.
