Back to website
2-min read

AgentStatus × Liberate, a quick map of how we fit

Independent verification for Liberate's AI agents.

We do two jobs: reachability from residential networks (past CDN/WAF), then reliability once reached — gold/contract and consistency checks, across voice, SMS, and email, from 2,500+ nodes across 70 countries. We sit alongside Liberate's reporting, transcripts, recordings, and quality signals, we don't replace them.

22M
tests
8,000+
agents
2,500+
residential devices
70
countries
agentstatusagentstatus.dev | partner brief

What we understand about Liberate (public)

Insurance-native AI across sales, servicing, and claims.

Liberate positions as insurance-native AI for sales, servicing, and claims, with Voice AI plus email and SMS, aimed at end-to-end resolution and tight integration into core insurance systems.

Public messaging emphasizes always-on coverage, multilingual voice, warm transfer with context, and an operations layer that includes call transcripts and recordings, sentiment, and a proprietary "smoothness" quality measure, alongside enterprise security claims such as HIPAA, SOC 2, PCI, and GDPR.

Their integrations story is ecosystem-heavy (Guidewire, Duck Creek, Salesforce, Snapsheet, and similar), which implies deep workflow and data-plane work, not a generic "drop-in widget" deployment model.

What AgentStatus is

We measure whether users can reach the agent, then whether it still passes its checks.

Reachability. Controlled validations from 2,500+ residential devices across 70 countries measure whether users can open the agent the way they do — past CDN, WAF, and bot walls. Multi-geo is observer vantage for access and last-mile latency — not answer localization by probe IP, and not agent tool egress.

Reliability. Once reachable, we run gold/contract checks when truth exists, plus rephrase, drift, and policy consistency probes when it doesn't. Dual LLM-as-judge scores open answers with a known ceiling — stably wrong but consistent still needs a domain expert.

That includes multi-turn conversations and multi-agent journeys when customer paths span tools, escalations, and handoffs, and it supports governance and risk conversations when stakeholders ask what was tested, from where, and what changed.

Outside-in validation is two separate jobs

Reachability

Claims agent · residential
Status dashboard with reachability verdict and regional coverage

Outcome

Residential path

Your monitor hits the VIP lane. Users hit the WAF. Datacenter checks get blocked, throttled, or allowlisted. Residential observers take the inbound path customers take — so “up” means reachable from home networks, not from AWS.

Reliability

Claims agent · eval
Answer quality dashboard with evaluation prompts and pass fail results

Outcome

Answer quality

Reachable and self-contradicting is still broken. Rephrase flips, drift, and policy breaks need no ground truth. Gold and dual judges cover the rest when truth exists. Uptime grades none of that.

Where we fit

We sit beside the platform. We do not replace it.

01

Outside-in vs inside-out

Liberate gives carriers strong inside-the-platform visibility: transcripts, recordings, sentiment, smoothness, and operational reporting. AgentStatus answers a complementary question: what did the real customer channel actually do for a controlled validate from a residential observer vantage (inbound reach / last-mile path — not answer localization by probe IP), including failures that only show up outside your own telemetry.

02

Voice-first reality

Insurance voice is path-dependent (carrier routing, telephony, ASR/TTS, regional behaviour). A distributed execution mesh is built to catch regressions that single-region checks miss.

03

Global execution footprint

2,500+ residential nodes across 70 countries prove we are not synthetic from a single cloud region. That matters for national and regional carriers when access and path failures only show from real networks — distinct from answer-quality checks once reached.

04

Consent-first posture

We do not pitch covert scanning of production policyholder lines. The right model is explicit pilot fixtures: sandbox numbers, approved test accounts, agreed call/SMS/email volumes, and clear success criteria.

The split

How the work divides

How the work divides

Their platform

Liberate, Inside-out
  • Voice AI + email + SMS
  • Transcripts & recordings
  • Sentiment & "smoothness"
  • Carrier system integrations
  • HIPAA / SOC 2 / PCI / GDPR

Outcome

System of record

Dashboards, exports, lifecycle tools, and orchestration remain theirs. We do not replace that surface.

AgentStatus

AgentStatus, Outside-in
  • Continuous validate traffic
  • Expected-answer checks & drift detection
  • Multi-turn / multi-agent journeys
  • Real-network execution evidence
  • 2,500+ nodes across 70 countries

Outcome

Outside-in layer

Residential inbound path past CDN/WAF, then gold, consistency, and scoped judges once the agent is reachable.

Proof of scale

Auditable scale metrics

In about two months, we have executed on the order of 18 million validate runs across the network. We also maintain on the order of 6,000 agent records in our system, meaning rows/configurations we track, including evaluation and pipeline agents, not "6,000 paying customers."

If helpful, we can share stricter production-only definitions under NDA.

What we are not claiming

We are an independent layer that runs alongside your stack.

We are not a replacement for Liberate's dashboards, QA workflows, or customer experience management features. We are an independent layer that can coexist, and, where useful, help teams correlate outside-in validate outcomes with inside-out transcripts, recordings, and quality scores.

What we'd like from this conversation

These three asks would move a pilot forward.

01

A 2-week sandbox pilot

A test phone number, a test inbox, and a set of agreed scenarios with expected answers. Two-week evaluation window. No production traffic, no policyholder data. At the end you get a written report of what we tested, what passed, and what drifted.

02

Security and procurement posture

How AgentStatus should connect in a way that satisfies carrier security reviews, data handling, least privilege, audit evidence, and clear test-traffic boundaries.

03

Where independent proof is most useful

Whether the right starting point is Liberate-internal QA, a joint carrier design where the carrier wants third-party evidence for go-live, or both.

Closing

Liberate helps insurers automate and operate insurance-native AI across the channels that matter.

AgentStatus helps those same organizations prove, continuously, that customer-facing behaviour still matches policy and expectations in the real world, globally, with evidence that holds up under scrutiny.

Contact·dulra@carmel.so·roman@carmel.so

Metrics are stated with explicit definitions: validate runs are scheduled executions over ~two months; agent records are database rows, not revenue customers. Liberate references above reflect public marketing on liberateinc.com as of the date of this note, not an endorsement by Liberate.