AgentStatus × Hippocratic AI
RWE-LLM, from outside the stack.
Hippocratic's RWE-LLM framework set a new bar: output testing over input validation, 6,234 clinicians, 307,038 evaluated calls. We agree with the thesis, and we're the external complement. Continuous voice-aware validations of production agents from 2,500+ network egress points in 70 countries, testing what gets through after the call leaves Hippocratic infrastructure.
Why we're reaching out
Internal output testing is the gold standard. External output testing is the next layer.
Dr. Bhimani's RWE-LLM paper makes an argument we deeply agree with: traditional input-side benchmarking is insufficient for healthcare, and comprehensive output testing across diverse real scenarios is the only path to safety assurance at scale.
The framework's four stages, pre-implementation, tiered review, resolution, and continuous monitoring, are well-developed inside Hippocratic. The continuous monitoring stage, by definition, has the largest external surface: ASR variance across network paths, IVR handoffs to third-party systems, multilingual auto-switch on degraded audio, response consistency across geographies.
That's the layer we test. Not what your 22-model constellation says. What gets through after the call leaves your infrastructure.
What AgentStatus is
We run voice-aware external tests against production AI agents.
Scheduled outside-in validations of voice agents, full audio in, full audio out, from 2,500+ network egress points across 70 countries. We measure ASR quality, response latency, audio dropouts, IVR navigation success, multilingual auto-switch behavior, and response consistency against expected answers. ~10M test runs in the last 2 months across 8,000+ agents.
Voice validating runs over programmatic telephony from independent network egress, not datacenter, not Hippocratic infra, not the customer's environment. We do not evaluate clinical accuracy and we do not replace your clinicians; we test the deployment layer between your model and the patient experience.
Outside-in validation is two separate jobs
Reachability

Outcome
Residential path
Your monitor hits the VIP lane. Users hit the WAF. Datacenter checks get blocked, throttled, or allowlisted. Residential observers take the inbound path customers take — so “up” means reachable from home networks, not from AWS.
Reliability

Outcome
Answer quality
Reachable and self-contradicting is still broken. Rephrase flips, drift, and policy breaks need no ground truth. Gold and dual judges cover the rest when truth exists. Uptime grades none of that.
Where we fit
Internal safety testing and outside-in reliability checks are different layers.
How the work divides
Their platform
- • <strong>6,234 clinicians</strong> evaluating call outputs
- • <strong>22-model constellation</strong> on Hippocratic infra
- • <strong>Clinical safety</strong>: did the agent give the right medical answer
- • <strong>Internal continuous monitoring</strong>: spot-checks, A/B, RWE feedback
Outcome
System of record
Dashboards, exports, lifecycle tools, and orchestration remain theirs. We do not replace that surface.
AgentStatus
- • <strong>Independent voice validations</strong> from outside the stack
- • <strong>70 countries</strong>, 2,500+ network egress points, scheduled and continuous
- • <strong>Deployment integrity</strong>: did the right answer make it through ASR, IVR, network, channel
- • <strong>External continuous monitoring</strong>: voice-aware, geography-split, drift-aware
Outcome
Outside-in layer
Residential inbound path past CDN/WAF, then gold, consistency, and scoped judges once the agent is reachable.
One sentence. RWE-LLM answers "did our agent say the right thing?" AgentStatus answers "did the right thing actually reach the patient from real networks — access and path integrity, not absolute omniscience?"
The deployment surface
Outside-in voice validations catch failures that stay invisible inside the stack.
Your internal eval runs on Hippocratic's infrastructure with Hippocratic's audio path. The deployment surface, the layer between Polaris and the patient, has variance your nurse panel can't see:
- • ASR degradation across network paths: transcription accuracy varies with audio compression, jitter, packet loss
- • Response latency by geography: P50/P95 latency differs across regions and network conditions in ways that affect conversational feel
- • IVR navigation drift: third-party systems (other providers, labs, pharmacies) change behavior; outside-in tests catch breaks early
- • Multilingual auto-switch behavior: Spanish at 99.83% on internal eval, but how does the auto-switch hold up on lower-quality audio paths?
- • Configuration / model-swap drift: silent regressions after infra updates, third-party API changes, or version rollouts
- • Regional outage detection: a customer in Tampa might experience failures invisible from Palo Alto
For your customers
This can become an asset Hippocratic hands to health-system buyers.
Health system procurement teams increasingly ask: "How do we continuously verify the AI vendor is performing as promised post-deployment?" RWE-LLM is the answer for Hippocratic's internal validation. AgentStatus can be the answer Hippocratic hands its enterprise buyers, independent, third-party, ongoing deployment monitoring, not vendor-self-reporting.
This makes Hippocratic's procurement easier, not harder, and it fits the published RWE-LLM thesis on output testing.
Proof of scale
Here is what we have already run, with plain definitions.
~10M test runs in ~2 months across the network. 8,000+ agents being tracked, including ones from companies you'd recognise (specifics under NDA).
We've also caught node operators trying to game the network with datacenter VMs instead of real consumer egress. Detection of adversarial behavior is built into the product, the same kind of rigor RWE-LLM applies to clinical outputs, applied to network integrity.
What we are not claiming
We are an independent layer that runs alongside your stack.
We are not clinicians. We are not adjudicating clinical accuracy. We are not replacing RWE-LLM, your clinician panel, or any part of Polaris. We do not handle PHI, tests run against demo or synthetic agent surfaces. We are the deployment-layer external complement to your internal output testing.
What we'd like from this conversation
We would like a thirty-minute conversation on methodology.
A 30-minute call with Dr. Bhimani's team
On whether external voice-aware output testing belongs as a published extension of the RWE-LLM continuous monitoring stage.
A joint readout or co-authored write-up
Hippocratic publishes papers as standard practice. We'd love to contribute the external-monitoring chapter, your voice, our data, real findings from validating one of your demo agents across geographies and network conditions.
One demo agent for a 2-week sandbox
A non-PHI surface we can validate externally for two weeks. Output: a methodology artifact showing ASR variance, latency distribution, IVR navigation correctness, and multilingual behavior across geographies.
Hippocratic argued that output testing is the gold standard.
AgentStatus is that same testing, from outside the stack, voice-aware, continuous, geography-distributed, and ready to be the published external chapter of RWE-LLM.
Contact·dulra@carmel.so·roman@carmel.so
"Test runs" and "agent rows" mean what we said above. Hippocratic AI and RWE-LLM descriptions are from public pages and announcements, not an endorsement by Hippocratic AI.
