A comprehensive guide to defining, measuring, and enforcing service level agreements for AI agents.
Why Agent SLAs Are Different
Traditional SLAs measure uptime: "Is the service responding?" Agent SLAs must also measure correctness: "Is the agent responding correctly?"
An agent can achieve 99.99% uptime while being semantically broken - returning HTTP 200 with useless or wrong responses. Traditional SLAs miss this entirely.
SLA Components
1. Availability
Definition: Percentage of time the agent endpoint is reachable and responding.
| Tier | Target | Allowed Downtime/Month |
|---|---|---|
| Basic | 99% | 7.2 hours |
| Standard | 99.5% | 3.6 hours |
| Premium | 99.9% | 43 minutes |
| Critical | 99.95% | 22 minutes |
2. Semantic Correctness
Definition: Percentage of responses that are actually correct and useful.
| Tier | Eval Pass Rate |
|---|---|
| Basic | 90% |
| Standard | 95% |
| Premium | 98% |
An agent returning "I cannot help" to every query has 100% availability but 0% correctness.
3. Performance
| Metric | Basic | Standard | Premium |
|---|---|---|---|
| TTFB P50 | <5s | <3s | <2s |
| TTFB P95 | <10s | <6s | <4s |
| Total P50 | <30s | <15s | <10s |
4. Geographic Consistency
| Metric | Target |
|---|---|
| All regions availability | >95% |
| Latency variance | <2x baseline |
| Regional outage detection | <15 minutes |
Defining Your SLA
Step 1: Understand Your Users
- Where are they located?
- What latency is acceptable for your use case?
- How critical is the agent to their workflow?
- What's the business impact of failures?
Step 2: Establish Baselines
Before committing to targets, measure current availability, eval pass rate, latency percentiles, and regional performance. Set targets you can actually achieve, then improve.
Step 3: Choose Components
| Use Case | Availability | Correctness | Latency | Geographic |
|---|---|---|---|---|
| Internal tool | ✓ | ✓ | - | - |
| Customer-facing | ✓ | ✓ | ✓ | ✓ |
| Enterprise B2B | ✓ | ✓ | ✓ | ✓ |
| API product | ✓ | ✓ | ✓ | ✓ |
Step 4: Define Measurement
For each component, specify the metric, method, frequency, and exclusions (maintenance windows, etc.).
Step 5: Set Consequences
SLAs without teeth are just targets. Consider service credits, termination rights, and reporting requirements.
Measurement Best Practices
- Use third-party monitoring: Provider-reported metrics have inherent conflicts of interest.
- Measure continuously: Spot checks miss issues.
- Include semantic checks: HTTP health checks are necessary but insufficient.
- Track percentiles: Averages hide tail latency.
- Distinguish error types: Separate agent failures from infrastructure failures.
Credit Calculations
Tiered Credits
| Actual Availability | Credit |
|---|---|
| 99.0% - 99.9% | 10% |
| 98.0% - 99.0% | 25% |
| 95.0% - 98.0% | 50% |
| <95.0% | 100% |
Always cap maximum credits (per month: 50-100% of monthly fee). Uncapped credits create unbounded liability.
Common Pitfalls
Uptime-Only SLAs: Measuring only HTTP availability misses semantic failures. Include eval pass rate.
Average-Only Latency: Averages hide P99 latency spikes. Use percentiles.
Single-Location Measurement: Checking from one location misses regional failures. Require geographic distribution.
No Baseline: Setting targets without historical data leads to unachievable or meaningless SLAs.
Unclear Exclusions: Vague exclusion language creates disputes. Enumerate specific exclusions precisely.
Example SLA
SERVICE LEVEL AGREEMENT
AI Customer Support Agent
1. AVAILABILITY
Target: 99.9%
Measurement: HTTP 2xx responses
Period: Monthly
2. SEMANTIC CORRECTNESS
Target: 95% gold prompt pass rate
Evaluation prompts: Health and contract tier
3. PERFORMANCE
TTFB P50: <2 seconds
TTFB P95: <5 seconds
Measurement: Third-party monitoring (Agent Status)
4. GEOGRAPHIC COVERAGE
Regions: US, EU, APAC
Per-region availability: >95%
5. CREDITS
Availability miss: 10% per 0.1% below target
Correctness miss: 10% per 1% below target
Maximum: 50% of monthly feeImplementation
Setting Up Monitoring
- Configure agent in monitoring system
- Enable all relevant regions
- Set check frequency matching SLA period
- Configure gold prompt validation
- Enable alerting for SLA violations
Handling Violations
- Detect violation via monitoring
- Document incident details
- Calculate credit owed
- Communicate with customer
- Conduct post-incident review
