Back to website
2-min read

AgentStatus × Hertz

User-side monitoring for Hertz's AI-powered customer support agent.

We ran a short-window validation sweep against the Hertz AI chat agent from real consumer devices. Two early findings worth flagging, plus a proposal to extend.

22M
Validations run
8,000+
Agents tracked
2,500+
Devices
70
Countries
agentstatusagentstatus.dev | partner brief

User-side validation is two separate jobs

Reachability

Claims agent · residential
Status dashboard with reachability verdict and regional coverage

Outcome

Can we talk to it?

Residential observers take the inbound path customers take — past CDN, WAF, and bot walls that treat datacenter synthetics differently. Monitoring asks: is it healthy right now? Reliability asks: does it keep working over time? “Up” means reachable from home networks, not from AWS.

Outcome verification

Claims agent · eval
Outcome verification dashboard with evaluation prompts and pass fail results

Outcome

Did it do the right thing?

Reachable and wrong is still broken. Scenario — did it finish the job? Compositional — do the pieces hold together? Safety — must-not-say, policy, attack probes. Stability — same ask, same story? Scores: Consistency and Drift. Not prose matching. Optional sample review is corroboration only.

Intro

The agent is up. That's not the same as working.

Hertz built and deployed an AI agent to handle customer support at scale. Platform-level monitoring confirms the agent is running and responding. What it doesn't tell you is what a real customer experiences when they open the chat widget on day one of a rental.

We ran a small, deliberate sweep from outside the platform to ground the conversation in real data.

What we ran

Here is what we already ran against this surface.

We sent three basic customer questions to the Hertz AI agent from real consumer devices across two regions over a 7-day window. 75 total probes. The questions are the exact ones any Hertz customer might ask on day one of a rental:

  • -"How do I check my reservation?"
  • -"What is your cancellation policy?"
  • -"How do I extend my rental?"

Finding 1

Customers are abandoning before the answer arrives.

p50 response time

9,408ms

~9.4 seconds

p95 response time

10,082ms

~10 seconds

sub-2-second responses

0 of 75

probes responded in under 2 seconds

latency comparison

Hertz AI (p50)9,408ms
Industry standard2,000ms
Feels instant500ms

Industry standard for AI chat support is under 2 seconds. At 9 seconds, customers are abandoning the conversation before the answer arrives.

The agent passes internal uptime checks. But from the outside, a customer asking "how do I extend my rental?" waits nearly 10 seconds for an answer. That's the gap between platform-level monitoring and user-side monitoring.

Finding 2

There is no visibility into the regions where customers actually are.

Probes ran from Hong Kong and Canada. Hertz's core customer base is US-based. There is no visibility into how the agent performs from US residential IPs or US mobile networks - the actual conditions under which Hertz customers open the chat widget.

Platform-level monitoring watches infrastructure. It doesn't tell you what a customer in Dallas or Chicago actually experiences.

This finding is a framing angle, not a measurement - a 2-week extension on US residential nodes would convert it into hard data.

The offer

Here is how AgentStatus would show up for Hertz.

Hertz built and deployed an AI agent to handle customer support at scale. What Decagon tells you is whether the agent is running. What AgentStatus tells you is whether the agent is working - from the regions where your customers actually are, asking the questions your customers actually ask, on the devices they actually use.

One view is inside-out. The other is user-side. You need both.

Honest framing

Here is what this proposal is, and what it is not.

This is a 7-day, 75-probe snapshot. Early signals, not a longitudinal study. The latency finding is strong and verifiable. The coverage finding is a framing angle worth converting into measurement.

This isn't a claim about outcome verification, correctness, or wrong answers. We measured response time and reachability, not whether the agent did the job. That's a separate workstream - and one we'd scope into the pilot.

The ask

Here is the concrete next step we are proposing.

A two-week pilot. US residential nodes. Extended prompt coverage across reservation, cancellation, extension, and loyalty flows. Weekly reports. Honest finding at the end.

Closing

Seven days. Seventy-five probes. Two findings.

We'd love to hop on a call and walk through what a US-residential, two-week pilot looks like for Hertz.

Contact·dulra@carmel.so·roman@carmel.so

A 7-day study concluding May 2026, on the publicly-reachable Hertz AI chat surface. Validations ran at conservative rate limits with no auth bypass; no customer data collected beyond verdict metadata, latency aggregates, and prompt outcomes. AgentStatus is independent user-side production monitoring for AI agents and is not affiliated with Hertz or Decagon.