Back to website
2-min read

AgentStatus × Roots Automation, a quick map of how we fit

Independent verification for Roots' insurance-trained agents.

We do two jobs: reachability from residential networks (past CDN/WAF), then reliability once reached — gold/contract and consistency checks, across the channels each platform supports, from 2,500+ nodes across 70 countries. We sit alongside Roots' AI agent library and multi-system integrations. We don't replace them.

22M
tests
8,000+
agents
2,500+
residential devices
70
countries
agentstatusagentstatus.dev | partner brief

What we understand about Roots

Insurance-trained AI agents for claims and underwriting.

Roots is the AI agent platform built for insurance, purpose-built agents for claims and underwriting, with insurance brains embedded into the platform itself. The agent library covers the lifecycle: submissions intake and triage, loss history access for underwriters, premium audit at 98%+ accuracy, FNOL/FROI automation, and policy management.

Roots positions itself as a transparent, ethical alternative to BPOs and to general-purpose LLMs, with multi-system process automation that integrates into the tools insurance teams already use daily, and a Trust Center built around the data security expectations of carrier customers.

What AgentStatus is

We measure whether users can reach the agent, then whether it still passes its checks.

Reachability. Controlled validations from 2,500+ residential devices across 70 countries measure whether users can open the agent the way they do — past CDN, WAF, and bot walls. Multi-geo is observer vantage for access and last-mile latency — not answer localization by probe IP, and not agent tool egress.

Reliability. Once reachable, we run gold/contract checks when truth exists, plus rephrase, drift, and policy consistency probes when it doesn't. Dual LLM-as-judge scores open answers with a known ceiling — stably wrong but consistent still needs a domain expert.

That includes multi-turn conversations and multi-agent journeys when customer paths span tools, escalations, and handoffs. It supports governance and risk conversations when stakeholders ask what was tested, from where, and what changed.

Outside-in validation is two separate jobs

Reachability

Claims agent · residential
Status dashboard with reachability verdict and regional coverage

Outcome

Residential path

Your monitor hits the VIP lane. Users hit the WAF. Datacenter checks get blocked, throttled, or allowlisted. Residential observers take the inbound path customers take — so “up” means reachable from home networks, not from AWS.

Reliability

Claims agent · eval
Answer quality dashboard with evaluation prompts and pass fail results

Outcome

Answer quality

Reachable and self-contradicting is still broken. Rephrase flips, drift, and policy breaks need no ground truth. Gold and dual judges cover the rest when truth exists. Uptime grades none of that.

Where we fit

We sit beside the platform. We do not replace it.

01

Insurance-trained agents vs production drift

Roots ships agents trained on insurance, that's the foundation. AgentStatus answers the next-layer question: what did the deployed agent actually do for a user-like validate today, given a specific submission, claim type, or geography, and did the answer drift from what the expected answer says it should be?

02

Premium audit accuracy vs ongoing accuracy

A 98%+ accuracy figure is a strong inside-out signal at evaluation time. Distributed validate traffic catches the cases where that accuracy starts slipping in production, before a renewal cycle is priced wrong or a claim is mis-triaged at FNOL.

03

Global execution footprint

2,500+ nodes across 70 countries is the proof we are not 'synthetic from a single cloud region.' It matters for carriers operating across multiple geographies and for access or path failures that only reproduce from specific residential networks or partner edges — distinct from answer-quality checks once reached.

04

Partner-friendly integration posture

We do not assume we can 'discover' Roots customers the way some web-widget vendors can be scraped. Credential-based surfaces (agent endpoints, sandbox environments, customer-approved monitoring) are the right model, aligned with the security posture Roots already maintains for its carrier customers.

The split

How the work divides

How the work divides

Their platform

Roots, Inside-out
  • Insurance-trained agents
  • Claims & underwriting library
  • Multi-system automation
  • Premium audit accuracy
  • Trust Center & data security

Outcome

System of record

Dashboards, exports, lifecycle tools, and orchestration remain theirs. We do not replace that surface.

AgentStatus

AgentStatus, Outside-in
  • Continuous validate traffic
  • Expected-answer checks & drift detection
  • Multi-turn / multi-agent journeys
  • Real-network execution evidence
  • 2,500+ nodes across 70 countries

Outcome

Outside-in layer

Residential inbound path past CDN/WAF, then gold, consistency, and scoped judges once the agent is reachable.

Proof of scale

Auditable scale metrics

In about two months, we have executed on the order of 18 million validate runs across the network. We also maintain on the order of 6,000 agent records in our system, meaning rows/configurations we track, including evaluation and pipeline agents, not "6,000 paying customers."

If helpful, we can share stricter production-only definitions under NDA.

What we are not claiming

We are an independent layer that runs alongside your stack.

We are not a replacement for Roots' agent library, insurance training, or multi-system integrations. We are an independent layer that can coexist with them, and, where useful, help carriers correlate outside-in validate outcomes with inside-out agent performance, so claims and underwriting leaders have continuous evidence the deployed agent is still behaving the way regulators and customers expect.

What we'd like from this conversation

These three asks would move a pilot forward.

01

A 2-week sandbox pilot

A sandbox agent (FNOL triage, premium audit, or underwriting submission), a set of agreed scenarios with expected answers, and a 2-week evaluation window. No production traffic, no policyholder data. At the end you get a written report of what we tested, what passed, and what drifted.

02

Security and procurement posture

How AgentStatus should connect in a way that satisfies carrier security reviews. Data handling, least privilege, audit evidence, and clear test-traffic boundaries aligned with Roots' Trust Center posture.

03

Where independent proof is most useful

Whether the right starting point is Roots-internal QA, a joint carrier scenario where the buyer operates under regulatory and audit requirements (primary, P&C, life), or both.

Closing

Roots helps carriers build and operate AI agents trained on insurance from day one.

AgentStatus helps those same carriers prove, continuously, that those agents behave the way regulators, auditors, and policyholders require, globally, with evidence that holds up under scrutiny.

Contact·dulra@carmel.so·roman@carmel.so

Metrics are stated with explicit definitions: validate runs are scheduled executions over ~two months; agent records are database rows, not revenue customers. Public Roots Automation references above reflect Roots' public product pages, agent library, and Trust Center documentation as of the date of this note.