AgentStatus × Decagon
Independent assurance for production customer AI, with ~10 days of evidence on a 19-brand cohort.
Two jobs: reachability from residential networks (past CDN/WAF), then reliability once reached — gold/contract and consistency checks, with a scoped ceiling on judges. We sit alongside Decagon's runtime (HTTP and streaming chat paths your customers already hit in the open). We do not replace them.
What we understand about Decagon
Embedded AI that has to work when traffic, content, and policy pressure are real.
Decagon's public product story centers on Agent Operating Procedures (AOPs), natural-language workflow specs teams can iterate quickly, plus omnichannel execution (chat, email, voice) and always-on scale. That combination is exactly where outside-in drift shows up first: policy tweaks, retrieval changes, routing edges, and model swaps that never ship as a press release but still move customer-visible behaviour.
Decagon powers concierge-grade support automation for a long tail of recognizable enterprises and consumer brands. In practice, buyers are judging not only tone and resolution rate, but whether the assistant stays reachable, coherent, and policy-aligned as traffic mixes, locales, and abuse patterns change.
For a platform like Decagon, the adjacent question an independent layer can answer is: from consumer-like networks, on the same paths users use, does the assistant still do the right thing this week, with evidence that survives a security review, not a screenshot deck.
What AgentStatus is
We measure whether users can reach the agent, then whether it still passes its checks.
Reachability. Controlled validate traffic from 2,500+ residential devices across 70 countries measures whether users can open the agent the way they do — past CDN, WAF, and bot walls that treat datacenter synthetics differently. Multi-geo is observer vantage for access and last-mile latency. It does not change agent tool egress or localize answers by probe IP.
Reliability. Once reachable, we run gold/contract checks when truth exists, plus rephrase, drift, and policy consistency probes when it doesn't. Dual LLM-as-judge scores open answers with a known ceiling — stably wrong but consistent still needs a domain expert.
Outside-in validation is two separate jobs
Reachability

Outcome
Residential path
Your monitor hits the VIP lane. Users hit the WAF. Datacenter checks get blocked, throttled, or allowlisted. Residential observers take the inbound path customers take — so “up” means reachable from home networks, not from AWS.
Reliability

Outcome
Answer quality
Reachable and self-contradicting is still broken. Rephrase flips, drift, and policy breaks need no ground truth. Gold and dual judges cover the rest when truth exists. Uptime grades none of that.
Where we fit
We sit beside the platform. We do not replace it.
Inside-out quality vs outside-in behaviour
Eval-time demos vs sustained production telemetry
Global execution footprint
Partner-friendly integration posture
The split
How the work divides
How the work divides
Their platform
- • Embedded customer AI & routing
- • Integrations & workflow logic
- • Tenant-specific configuration
- • Support outcomes & analytics
- • Platform scale & roadmap
Outcome
System of record
Dashboards, exports, lifecycle tools, and orchestration remain theirs. We do not replace that surface.
AgentStatus
- • Scheduled validate traffic
- • Gold libraries & drift signals
- • Multi-turn / streaming paths where enabled
- • Verdict-tier evidence + optional conformance detail
- • 2,500+ nodes across 70 countries
Outcome
Outside-in layer
Residential inbound path past CDN/WAF, then gold, consistency, and scoped judges once the agent is reachable.
Proof of scale
Auditable scale metrics
A. Posture first. How we operated: all activity targeted publicly reachable, customer-visible chat surfaces that use Decagon-shaped HTTP and WebSocket flows, with conservative rate limits and no attempt to bypass authentication. We did not harvest tenant back-office data. What we retained is verdict metadata, latency and pass-rate aggregates, short response previews, and conformance outcomes, the minimum needed to prove behaviour, not to reconstruct customer records.
B. What we measured (Decagon-only). Over a ~10-day window ending late April 2026, we ran 19 independent monitors of customer-visible Decagon HTTP production-shaped endpoints on a 12-hour cadence. That produced 417 aggregate rora_results snapshots and ~3,980 underlying validate executions in our telemetry, plus ~900 structured conformance rows where guardrail-style validations were enabled. In our taxonomy, UP means the run met configured health expectations; DEGRADED means transport succeeded but a configured semantic or policy check (gold, expectation, or conformance) failed, reachable but wrong or unsafe under test, not "server down." AUTH_ERROR means we hit an auth, entitlement, or quota gate on the path we used (including cases where the surface expects credentials we did not possess).
C. Honest mix and confidentiality. The snapshot distribution is mostly UP, with a non-trivial DEGRADED tail and a small AUTH_ERROR set. We show the honest mix, not a cherry-picked win rate. Per-brand verdict mix, degradation categories, and example validate rows are available on request under mutual confidentiality (we do not put customer brand names on a first-touch web page).
D. Definitions. A rora_results row is one scheduled aggregate snapshot for a monitored configuration (not "ARR customers"). Validation executions are the underlying checks that roll up into pass rate, latency, and verdict. Conformance rows are outcomes from adversarial-style validations where that program was enabled. These metrics are not Decagon SLAs, revenue, or customer counts unless separately agreed in writing.
What we are not claiming
We are an independent layer that runs alongside your stack.
We are not a replacement for Decagon's platform, routing, retrieval, or customer-specific policy. We are an independent layer that produces repeatable, externally executed evidence about how customer-visible agents behave over time, and where they start to drift.
What we'd like from this conversation
These three asks would move a pilot forward.
Validate the fit
Where would Decagon want independent assurance artifacts surfaced: partner GTM, enterprise security reviews, or joint customer success? Where should everything remain native to Decagon's own analytics?
A practical next step
A small, named cohort (sandbox or live-with-consent) where we align on gold prompts, streaming expectations, and rate limits, then agree on a shared definition of healthy that both sides can defend in front of a CIO.
Partner path
If there's a path to work together, we'd want a practical conversation about credentialed monitoring, tenant scoping, and customer approval. Those three decisions determine whether outside-in evidence is useful to your enterprise customers or just noise on the roadmap.
Decagon × AgentStatus.
Decagon helps enterprises ship and scale customer AI that actually runs in production. AgentStatus helps those same enterprises prove, continuously, that it still behaves the way security, procurement, and brand teams require, with evidence that survives scrutiny outside the demo room.
Contact·dulra@carmel.so·roman@carmel.so
Figures above reflect a time-bounded monitoring window in production (April 2026) on 19 Decagon HTTP monitors with 12-hour scheduling. Metrics are stated with explicit definitions: a rora_results row is one scheduled aggregate snapshot for a monitored configuration; validate executions are underlying checks contributing to aggregates; conformance rows reflect adversarial-style validations where enabled. These metrics are not revenue, customer counts, or Decagon-specific SLAs unless separately agreed in writing.
