Back to website
2-min read

AgentStatus × Sixfold

Independent, distributed assurance for production AI agents, alongside Sixfold's underwriting AI platform.

Two jobs: reachability from residential networks (past CDN/WAF), then reliability once reached — gold/contract and consistency checks, with a scoped ceiling on judges. We sit alongside your platform. We do not replace them.

22M
validations
8,000+
agents
2,500+
residential devices
70
countries
agentstatusagentstatus.dev | partner brief

What we understand about Sixfold

The first AI purpose-built for insurance underwriters.

Sixfold's platform brings AI agents directly into the underwriter's workflow, agents that triage cases, surface cited risk insights, score every submission 1 to 5 against carrier guidelines, write referral emails, and audit cases for compliance issues before binding. The latest advancement, Institutional Intelligence, encodes a carrier's risk appetite, underwriting guidelines, and historical decisions so the platform answers the way the carrier expects.

Sixfold operates with insurance-grade data governance and isolation, integrates via API into existing workbenches, CRMs, and policy admin systems, and is deployed in production at carriers including Zurich North America, Generali Global Corporate & Commercial, Skyward Specialty, and Mosaic. Backed by a Series B from Brewer Lane Ventures, with strategic investment from Guidewire Software, Bessemer, and Salesforce Ventures.

What AgentStatus is

We measure whether users can reach the agent, then whether it still passes its checks.

Reachability. Controlled validate traffic from 2,500+ residential devices across 70 countries measures whether users can open the agent the way they do — past CDN, WAF, and bot walls that treat datacenter synthetics differently. Multi-geo is observer vantage for access and last-mile latency. It does not change agent tool egress or localize answers by probe IP.

Reliability. Once reachable, we run gold/contract checks when truth exists, plus rephrase, drift, and policy consistency probes when it doesn't. Dual LLM-as-judge scores open answers with a known ceiling — stably wrong but consistent still needs a domain expert.

Outside-in validation is two separate jobs

Reachability

Claims agent · residential
Status dashboard with reachability verdict and regional coverage

Outcome

Residential path

Your monitor hits the VIP lane. Users hit the WAF. Datacenter checks get blocked, throttled, or allowlisted. Residential observers take the inbound path customers take — so “up” means reachable from home networks, not from AWS.

Reliability

Claims agent · eval
Answer quality dashboard with evaluation prompts and pass fail results

Outcome

Answer quality

Reachable and self-contradicting is still broken. Rephrase flips, drift, and policy breaks need no ground truth. Gold and dual judges cover the rest when truth exists. Uptime grades none of that.

Where we fit

We sit beside the platform. We do not replace it.

01

Eval-time accuracy vs production drift

Sixfold's models are evaluated against carrier guidelines at training and deployment time, the foundation. AgentStatus answers a different question: a quarter into deployment, after model updates and changes to a carrier's appetite, is the agent still scoring submissions the way Institutional Intelligence said it should?
02

Inside-out scoring vs outside-in evidence

A 1 to 5 appetite-fit score reported inside the workbench is a strong inside-out signal at the moment of decision. Distributed validate traffic catches the cases where the same submission, is unreachable from residential ISPs while cloud checks stay green — or, once reached, drifts against gold/contract and consistency checks after a model refresh, before a quote goes out misaligned with appetite. Observer geo is inbound reach, not score localization by probe IP.
03

Global execution footprint

2,500+ nodes across 70 countries is the proof we are not synthetic from a single cloud region. For Sixfold's carrier customers operating across multiple geographies, lines, and partner integrations, it matters that residential observers validate inbound reach across those markets — including Lloyd's specialty lines via Cohort 12 — catching access failures cloud synthetics miss. Multi-geo is vantage for reachability, not score localization by probe IP.
04

Partner-friendly integration posture

We do not assume we can discover Sixfold customers the way some web-widget vendors can be scraped. Credential-based surfaces (sandbox API access, customer-approved monitoring, joint customer scenarios) are the right model, aligned with the SOC 2 / data-isolation posture Sixfold already maintains for its carrier customers.

The split

How the work divides

How the work divides

Their platform

Sixfold, inside-out
  • Institutional Intelligence
  • 1-5 appetite-fit scoring
  • Cited risk insights
  • Referral & compliance agents
  • Workbench / CRM / PA integration

Outcome

System of record

Dashboards, exports, lifecycle tools, and orchestration remain theirs. We do not replace that surface.

AgentStatus

AgentStatus — reach + reliability
  • Continuous validate traffic
  • Gold libraries & drift detection
  • Multi-turn / multi-agent journeys
  • Real-network execution evidence
  • 2,500+ nodes across 70 countries

Outcome

Outside-in layer

Residential inbound path past CDN/WAF, then gold, consistency, and scoped judges once the agent is reachable.

Proof of scale

Auditable scale metrics

In about two months, we have executed on the order of 18 million validate runs across the network. We also maintain on the order of 6,000 agent records in our system, meaning rows and configurations we track, including evaluation and pipeline agents, not "6,000 paying customers."

We have also caught node operators trying to game the network with datacenter VMs instead of real consumer egress. Detection of adversarial behaviour is built into the product. If helpful, we can share stricter production-only definitions under NDA.

What we are not claiming

We are an independent layer that runs alongside your stack.

We are not a replacement for Sixfold's Institutional Intelligence, agent library, or workbench integrations. We are an independent layer that can coexist with them, and where useful, help carriers correlate outside-in validate outcomes with inside-out scoring accuracy, so underwriting leaders have continuous evidence that the deployed agent is still behaving the way the carrier's appetite requires.

What we'd like from this conversation

These three asks would move a pilot forward.

01

Validate the fit

Where do Sixfold's carrier customers want independent assurance for agent output, and where does Sixfold prefer everything native to the platform?

02

A practical next step

A sandbox agent we can validate with gold prompts representative of an underwriting workflow (submission triage, appetite-fit scoring, referral generation), so Institutional Intelligence and AgentStatus drift detection tell one story together.

03

Partner path

If there is a partner path, we'd like to understand supported integration patterns for carriers operating across multiple geographies and lines, particularly through Sixfold's strategic alliances with Guidewire and across the Lloyd's Market.

Closing

Sixfold × AgentStatus.

Sixfold helps carriers build and operate AI agents that bring joy back to underwriting. AgentStatus helps those same carriers prove, continuously, that those agents behave the way appetite, regulators, and brokers require, globally, with evidence that holds up under scrutiny.

Contact·dulra@carmel.so·roman@carmel.so

Metrics are stated with explicit definitions: validate runs are scheduled executions over ~two months; agent records are database rows, not revenue customers. Public Sixfold references above reflect Sixfold's public product pages, customer disclosures, and Series B announcement as of the date of this note.