AgentStatus × Parloa, a quick map of how we fit
An independent outside-in layer beside the Test stack you already built.
You already ship a serious first-party Test product. We are not asking you to replace it. We sit beside it as a neutral third party that reaches your deployed agents from real residential networks, vantage and independence that inside-out simulation and transcript analysis are not designed to provide on their own.
What we understand about your platform
Your Test product
Your AI Agent Management Platform spans design, test, scale, optimize, and secure. The Test pillar is substantive: dual-model simulation (one model as caller, one as agent), blends of historical transcripts and synthetic scenarios, scoring with both LLM-as-judge and rule-based criteria, and post-deployment evaluation of real conversations. You also partner with OpenAI on reliability. First-party testing is clearly strategic for you, not an afterthought.
That stack is strong and inside-out: it runs in your environment, grades agents you operate, and, for post-deployment work, typically analyzes transcripts and events you already hold (Insights, Data Hub, Transcripts API). That is first-party eval and conversation truth from your own surface. It is a different job from probing the deployed agent from where an external user reaches it.
For text integrations, TextchatV2 is a clear HTTP model: a dialoghook endpoint with a release id and Bearer token, per your public API specification.
What AgentStatus is
Outside-in beside your stack
Reachability. Controlled calls from 2,500+ residential devices across 70 countries measure whether end users can open the agent the way they do: past CDN, WAF, and bot walls that treat datacenter synthetics differently. Multi-geo is observer vantage for access and last-mile latency. It does not change agent tool egress or localize answers by probe IP. A sim engine does not become this by adding another scoring rubric; it is a different kind of infrastructure.
Independence. Many of your enterprise buyers (procurement, CISOs, audit) want a neutral third party attesting to production behavior. Your own Test product, however strong, is still first-party. You cannot be your own independent auditor. We can sit outside that boundary and provide the attestation they ask for.
Reliability checks once reached (gold/contract, rephrase, drift, policy consistency, scoped dual judges) are how we deepen the signal after we have a real external vantage. They are the method, not a claim that you lack testing capability. Much of that method is buildable in-house. What we bring that is hard to self-supply is the residential network and the third-party role.
Where we fit
What we add beside it
Your post-deployment eval and our outside-in probes are different jobs.
If post-deployment evaluation analyzes transcripts and events you already have, that is first-party conversation truth. Hitting the live agent from residential user conditions is a separate layer. We would like to understand how you run post-deployment eval today so we map cleanly against it rather than overlap it.
Residential reachability is a different kind of infrastructure.
Cloud synthetics and allowlisted monitors often take a VIP lane. Residential observers catch CDN/WAF/bot-wall access failures on the path your customers’ end users take. Operating a global opted-in residential fleet is an operations business, not a scoring checkbox next to dual-model simulation.
Independence is something you cannot self-supply.
For enterprise sales, “we tested ourselves” carries less weight than a neutral party’s attestation. You can keep improving first-party Test and still not be a third party about yourselves. Where your buyers ask for independent verification, we are that outside attestation next to your own evidence.
Judges as a second opinion, not a replacement for yours.
You use LLM-as-judge in scoring. So do we. We are not claiming a better judge. We offer an independent check from outside, with a scoped ceiling, as a second opinion on first-party scores, useful for the stably-wrong residue any single judge can miss.
Partner-friendly integration
We connect only with credentials and surfaces you approve, for example TextchatV2, not by scraping for your customers. That matches the security posture your enterprise buyers expect.
The split
How the work divides
How the work divides
Your Test stack
- • Dual-model simulation and synthetic + historical tests
- • LLM-as-judge and rule-based scoring
- • Post-deployment conversation evaluation (first-party)
- • Insights, Data Hub, Transcripts API
- • AMP lifecycle: design, test, scale, optimize, secure
Outcome
System of record
Your Test product stays the system of record for first-party QA. We do not compete with simulation, Insights, or transcript eval. We assume that stack is strong and strategic for you.
AgentStatus
- • Residential inbound reachability past CDN/WAF
- • Probes from real external user conditions
- • Third-party evidence for procurement and audit
- • Independent check on first-party scores (scoped judges)
- • Shared pilot metric: failures your internal testing did not catch
Outcome
The layer beside it
Not a second Test product. Residential vantage plus independence, so your buyers get outside-in evidence next to the inside-out truth you already produce.
Where we sit in the landscape
Two jobs, not one tool category
Assurance for production agents usually splits into two jobs.
The builder stack is what the team that owns the agent does (or buys): access, integration, first-party traffic. Simulation, judges, transcript analysis, inside-out observability. Same category as a serious Test or Insights product.
Outside-in is different. You exercise the deployed agent from real residential networks. No access required. It still works when the person who needs the answer does not own the agent.
We sit in that second job, beside the first. Not as a substitute for it.
Two well-funded products live in the builder stack.
Coval (coval.ai) does sim and eval, voice-first. Synthetic conversations before launch, evals on integrated live traffic, human-in-the-loop. Closest cousin to a first-party Test product. ($28M Series A led by Norwest, June 2026; ~$31M total, public press.)
Raindrop (raindrop.ai) does inside-out observability (“Sentry for AI”). Install in the app, watch real conversations, flag hallucinations, frustration, tool failures, jailbreaks. Closest cousin to Insights. ($15M seed, Lightspeed, Dec 2025, public press.)
A platform like yours already covers those jobs, or can deepen them. AgentStatus is the independent outside-in layer next to that stack.
| Builder stack (owner has access) | Outside-in (no access) | ||
|---|---|---|---|
| Raindrop | Coval | AgentStatus | |
| Category | Inside-out observability | Lab sim + integrated eval | Residential outside-in probes |
| Same job as… | Insights / production issue detection | A first-party Test product | Independent attestation (not Test) |
| When it catches a failure | After a real user hit it | In sim, or after live integrated traffic | On our schedule, before a user has to hit it |
| Needs integration / access? | Yes: code in your app | Yes: integrate | No: tests the deployed agent like a stranger |
| Whose traffic | Your real users’ conversations | Synthetic callers + your integrated traffic | Independent external probes |
| Reachability | Assumed (lives inside) | Assumed (simulated / integrated) | Measured: catches path and bot-wall failures |
| Who it can serve | The team that owns the agent | The team that owns the agent | Anyone, including a third party who doesn’t own it |
What they do well (we do not hide this)
- Coval is voice-first and mature at it. Our voice surface is roadmap. If the conversation is voice-primary, we say so. We do not imply voice parity today.
- Live monitoring sees real-user intents our probes did not invent. Raindrop and integrated Coval traffic catch behavior we did not script. That is a genuine strength of after-the-fact monitoring, and of a strong Insights layer. We complement it; we do not replace it.
- Both are further along on funding and product maturity in their category. We do not win a feature bake-off against builder-stack tools. We win when the question is category and vantage.
What the builder stack cannot self-supply
No access, no integration required
Builder-stack tools need you to own the agent and install or integrate. They cannot validate an agent they do not control, which closes the third-party lane (insurers, registries, certifiers, partners). We attach as independent evidence next to your first-party Test without becoming another SDK in your customers’ critical path.
External residential vantage
Inside-out and lab stacks largely assume reachability. We measure it from real consumer networks across geographies. Datacenter-blocked or bot-walled paths stay invisible when you only ever sit inside the path.
How this maps for Parloa
You already sit in the builder stack. Dual-model sim, judges, Insights, transcript eval: that is Coval/Raindrop territory as a category. We are not asking you to swap that for us. We sit beside it as the outside-in, independent layer that stack is not designed to self-supply.
When buyers name Coval or Raindrop: name the category split. Their question is usually “who grades the agent from the owner’s seat?” Ours is “what happens when a stranger reaches the deployed agent from a home network, and who can attest without owning the stack?”
When a buyer does not own the agent (procurement, audit, insurer, partner): access-based tools cannot serve them. We can. That is often the cleanest way to attach AgentStatus next to Parloa in an enterprise deal without displacing Test.
We do not claim voice parity or eval-depth parity with Coval. We do not claim to replace Raindrop’s production issue detection. Those are builder-stack wins. We claim residential outside-in vantage and third-party independence.
Proof of scale
Auditable scale metrics
In about two months, we have executed on the order of 18 million validate runs across the network. We also maintain on the order of 6,000 agent records in our system, meaning rows/configurations we track, including evaluation and pipeline agents, not "6,000 paying customers."
If helpful, we can share stricter production-only definitions under NDA.
What we are not claiming
Coexistence, not displacement
We are not a replacement for your Test product, Data Hub, or Transcripts API. We do not sell “we simulate conversations and score with judges” as if that were missing from your stack. You already do that well.
We are also not claiming that reliability methods (gold, consistency, conversation-level compounding) are a secret only we can implement. A team that already ships dual-model sim and LLM-as-judge can extend first-party scoring. What we bring that is hard to self-supply is residential vantage and third-party independence: one is an operations business outside your core platform; the other is something no first party can truthfully claim about itself.
Where it helps your customers, we correlate outside-in outcomes with the inside-out conversation truth you already produce. That is coexistence, not displacement.
What we'd like from this conversation
Pilot checklist
How you run post-deployment evaluation today
When Test evaluates real conversations after deploy, does that mean analyzing transcripts and events you already have, or probing the live agent from external user conditions? That answer tells us how to sit cleanly beside your product rather than overlap it.
Whether your enterprise buyers ask for independent verification
In Fortune-scale deals, do procurement or security teams ask who independently verifies agent behavior in production? If they do, we can attach as third-party evidence next to your own Test results. If they rarely do, we focus the pilot on production failures your inside-out stack is not positioned to see.
A two-week sandbox pilot with a shared success metric
Sandbox TextchatV2 (release id and API token), agreed scenarios, two weeks, no production customer data. Shared success metric: how many real failures we surface that your internal testing did not, especially reachability and whole-conversation breaks. You keep the report either way; we both see whether the layer earns its place.
You help enterprises build, test, and operate serious AI agents.
We sit beside your Test stack as the independent outside-in layer: residential vantage plus third-party attestation for the buyers who need evidence that is not only yours.
Contact·dulra@carmel.so·roman@carmel.so
Metrics are stated with explicit definitions: validate runs are scheduled executions over ~two months; agent records are database rows, not revenue customers. Public Parloa references above reflect Parloa's public product pages and documentation as of the date of this note.
