All articles

Agent Status / Field Notes

The Rise of the Agent Economy: Who's Building the Infrastructure?

Something fundamental is shifting in how software gets built and consumed.

We're moving from an economy of applications to an economy of agents.

And like every infrastructure shift before it, the real opportunity isn't in building agents-it's in building the infrastructure that agents need to thrive.


01Section

The Agent Explosion

The numbers tell the story of rapid adoption-and rapid failure:

  • 95% of AI agents failed in production in 2025 (vaza.ai)
  • 42.9 percentage point reliability gap between benchmark performance and real-world conditions (HB-Eval)
  • 8.8% performance drop when agents face perturbations vs ideal conditions (arxiv)

We've crossed a threshold. Agents aren't experiments anymore. They're infrastructure-infrastructure that fails at alarming rates without proper observability.

And with that transition comes a familiar challenge: how do you operate this stuff reliably at scale?


02Section

Lessons from Previous Shifts

Every infrastructure era follows a pattern:

The Cloud Transition (2006-2015)

First: Raw compute (EC2)

Then: Platform services (RDS, S3)

Finally: Operational tooling (CloudWatch, DataDog, PagerDuty)

The monitoring and operations layer came last, but it became the most critical. You can't run infrastructure you can't observe.

The Container Transition (2014-2020)

First: Container runtimes (Docker)

Then: Orchestration (Kubernetes)

Finally: Observability (Prometheus, Jaeger, service mesh)

Same pattern. The tooling to operate containers reliably lagged the ability to create them.

The Agent Transition (2023-?)

First: Model APIs (OpenAI, Anthropic)

Then: Agent frameworks (LangChain, AutoGPT, CrewAI)

Now: Operational infrastructure (monitoring, validation, orchestration)

We're in the "finally" phase. The ability to build agents has outpaced the ability to operate them reliably.


03Section

The Infrastructure Gap

Building an agent is easy. Building a reliable agent is hard.

Here's what's missing:

Observability

Traditional APM sees:

  • Request/response times
  • Error rates
  • Throughput

Agents need:

  • Semantic correctness validation
  • Behavioral consistency tracking
  • Multi-turn conversation analysis
  • Tool use monitoring
  • Memory/state inspection

Your DataDog dashboard can't tell you if your agent is hallucinating. That's a problem.

Validation

Traditional testing:

  • Unit tests
  • Integration tests
  • E2E tests

Agents need:

  • Continuous behavioral validation
  • Golden dataset verification
  • Prompt regression testing
  • Semantic equivalence checking
  • Red team/adversarial testing

You can't write a unit test for "responds helpfully." That's a problem.

Deployment

Traditional CI/CD:

  • Code changes → tests → deploy
  • Rollback on errors

Agents need:

  • Prompt versioning
  • Model version pinning
  • A/B testing for agent behavior
  • Gradual rollouts with semantic gates
  • Rollback on quality degradation (not just errors)

Git doesn't version prompts well. That's a problem.

Reliability

Traditional SRE:

  • Uptime monitoring
  • Error budget management
  • Incident response

Agents need:

  • "Up but broken" detection
  • Geographic validation
  • Latency SLA enforcement
  • Correctness SLA enforcement
  • Automated recovery from semantic failures

99.9% uptime means nothing when research shows agents can drop from 86.9% to 44.0% success rates under real-world stress (HB-Eval).


04Section

Who's Building What

The agent infrastructure landscape is emerging:

Model Providers (OpenAI, Anthropic, Google)

Building: APIs, safety systems, some observability

Gap: Third-party validation, geographic testing, multi-model orchestration

Orchestration Frameworks (LangChain, LlamaIndex, CrewAI)

Building: Agent construction, tool integration, memory systems

Gap: Production operations, monitoring, reliability

Traditional Observability (DataDog, New Relic)

Building: Adding LLM-aware features

Gap: Semantic validation, agent-specific failure modes, geographic distribution

Agent-Native Infrastructure (Emerging)

Building: Purpose-built monitoring, validation, orchestration

This is where the opportunity lives.


05Section

The Trust Problem

There's a deeper issue: trust.

Agents are black boxes that make decisions. Unlike traditional software where you can trace exactly what happened, agents operate probabilistically.

This creates organizational friction:

  • Engineers can't explain exactly why the agent did something
  • Product can't guarantee consistent experiences
  • Legal can't prove compliance
  • Customers can't verify claims

The companies that solve trust will capture the market.

Trust requires:

  1. Visibility - What is the agent doing?
  2. Verification - Is it doing the right thing?
  3. Accountability - Can we prove it?

These are infrastructure problems masquerading as AI problems.


06Section

Market Structure

The agent infrastructure market is segmenting:

Horizontal Infrastructure

Serves all agent types:

  • Monitoring and observability
  • Validation and testing
  • Deployment and versioning
  • Security and compliance

Agent Status sits here. We validate any agent, regardless of framework or model.

Vertical Infrastructure

Serves specific domains:

  • Healthcare agent compliance
  • Financial agent audit trails
  • Legal agent discovery tools

Domain expertise + infrastructure.

Embedded Infrastructure

Baked into platforms:

  • OpenAI's built-in monitoring
  • AWS Bedrock observability
  • Model provider safety systems

Convenience, but vendor lock-in.


07Section

The Reliability Moat

Here's our thesis:

In the agent economy, reliability is the moat.

When every company can build an agent (and they can), differentiation shifts to:

  1. Quality - Does it actually work well?
  2. Reliability - Does it work consistently?
  3. Trust - Can customers depend on it?

These are operational problems, not AI research problems.

The companies that invest in agent reliability infrastructure will:

  • Ship higher-quality agents
  • Catch issues faster
  • Iterate more confidently
  • Win customer trust

The companies that don't will:

  • Learn about problems from customers
  • Ship slowly and cautiously
  • Accumulate technical debt
  • Lose to competitors who move faster

08Section

What We're Building

Agent Status is infrastructure for the agent economy.

We started with a simple question: "Is this agent actually working?"

That led to:

  • Gold prompt validation - Semantic correctness checking
  • Geographic distribution - Real-device tests worldwide
  • Continuous monitoring - Not just at deploy, all the time
  • Threshold-based verdicts - UP, DEGRADED, DOWN (not just percentages)

We're building the monitoring layer that agents deserve.


09Section

What's Next

Near-Term (2026)

  • Semantic validation becomes standard
  • Agent SLAs include correctness metrics
  • Geographic reliability becomes table stakes
  • "Agent-aware" monitoring emerges as category

Medium-Term (2027)

  • Multi-agent orchestration monitoring
  • Cross-agent trust scoring
  • Agent marketplace certification
  • Automated remediation and failover

Long-Term (2028+)

  • Agent reliability standards and certification
  • Insurance products for agent failures
  • Regulatory requirements for critical agents
  • Mature, commoditized infrastructure

10Section

The Opportunity

Every infrastructure transition creates new winners.

The cloud transition created AWS, then DataDog, PagerDuty, and a constellation of operational tools.

The container transition created Docker, Kubernetes, and the cloud-native ecosystem.

The agent transition will create new infrastructure giants. The only question is who.

The opportunity is clear:

  • 95% of agents failing in production without proper monitoring (vaza.ai)
  • 42.9% reliability gap between ideal and real conditions (HB-Eval)
  • Traditional APM wasn't built for semantic validation

Someone will monitor those agents. Someone will validate them. Someone will make them reliable.

That's what we're building.

Independent monitoring

See your agent the way the world sees it.

Outside-in validations from real residential nodes, evaluation prompts that catch silent-200 failures.