How Production Teams Use Agent Status in Their CI/CD Pipeline

"We run tests before deploy. Isn't that enough?"

agentstatusagentstatus.dev | July 2026

"We run tests before deploy. Isn't that enough?"

No. And here's why.

Tests verify that code works in a test environment. Production validation verifies that your agent works in production, right now, from real locations.

The best teams do both. Here's how.


01Section

The Testing Gap

What Tests Catch

  • Logic errors in your code
  • Regressions from changes
  • Integration issues with mocked dependencies
  • Format and schema problems

What Tests Miss

  • Production environment issues
  • Real dependency behavior
  • Network/geographic problems
  • Model version changes
  • Rate limiting in production
  • Performance under real conditions

Tests answer: "Does the code work?" Production validation answers: "Does the deployed agent work?"

Different questions. Both matter.


02Section

Integration Patterns

Pattern 1: Post-Deploy Validation

When: After deploying, before considering deploy "complete"

What: Trigger Agent Status validation, fail deploy if agent is unhealthy

snippet
# GitHub Actions
deploy:
  runs-on: ubuntu-latest
  steps:
    - name: Deploy
      run: ./deploy.sh

    - name: Wait for deployment
      run: sleep 30

    - name: Validate Agent
      env:
        AGENTSTATUS_API_KEY: ${{ secrets.AGENTSTATUS_API_KEY }}
        AGENT_ID: ${{ secrets.AGENTSTATUS_AGENT_ID }}
      run: |
        # Trigger validation
        RESPONSE=$(curl -s -X POST \
          "https://api.agentstatus.dev/agents/$AGENT_ID/run" \
          -H "Authorization: Bearer $AGENTSTATUS_API_KEY")

        DECISION_ID=$(echo $RESPONSE | jq -r '.decision_id')

        # Poll for results
        for i in {1..30}; do
          sleep 5
          RESULT=$(curl -s \
            "https://api.fabric.carmel.so/api/agentstatus/runs/$DECISION_ID" \
            -H "Authorization: Bearer $AGENTSTATUS_API_KEY")

          STATUS=$(echo $RESULT | jq -r '.status // "pending"')

          if [ "$STATUS" = "completed" ]; then
            VERDICT=$(echo $RESULT | jq -r '.overall_verdict')

            if [ "$VERDICT" = "DOWN" ]; then
              echo "❌ Agent validation FAILED"
              exit 1
            fi

            echo "✅ Agent validation PASSED: $VERDICT"
            exit 0
          fi
        done

        echo "⏱️ Validation timed out"
        exit 1

    - name: Rollback on failure
      if: failure()
      run: ./rollback.sh

Benefits: Catch deploy issues immediately. Auto-rollback on failure. Confidence before marking deploy complete.

Pattern 2: Staging Gate

When: Before promoting from staging to production

What: Validate staging agent meets quality bar

snippet
promote-to-production:
  runs-on: ubuntu-latest
  steps:
    - name: Validate Staging
      env:
        AGENTSTATUS_API_KEY: ${{ secrets.AGENTSTATUS_API_KEY }}
        STAGING_AGENT_ID: ${{ secrets.STAGING_AGENT_ID }}
      run: |
        ./scripts/validate-agent.sh $STAGING_AGENT_ID

    - name: Promote to Production
      if: success()
      run: |
        ./promote-to-production.sh

    - name: Validate Production
      env:
        AGENTSTATUS_API_KEY: ${{ secrets.AGENTSTATUS_API_KEY }}
        PROD_AGENT_ID: ${{ secrets.PROD_AGENT_ID }}
      run: |
        ./scripts/validate-agent.sh $PROD_AGENT_ID

Benefits: Catch issues before they reach production. Gate promotion on quality. Validate both environments.

Pattern 3: Canary Validation

When: During gradual rollouts

What: Validate canary before expanding

snippet
canary-deploy:
  runs-on: ubuntu-latest
  steps:
    - name: Deploy to Canary (5%)
      run: ./deploy-canary.sh

    - name: Wait for traffic
      run: sleep 300  # 5 minutes of canary traffic

    - name: Validate Canary
      run: |
        RESULT=$(./scripts/validate-agent.sh $CANARY_AGENT_ID)

        LATENCY=$(echo $RESULT | jq '.latency_p50_ms')
        GOLD_RATE=$(echo $RESULT | jq '.gold_pass_rate')

        PROD_LATENCY=$(./scripts/get-prod-baseline.sh latency)

        if [ $LATENCY -gt $(($PROD_LATENCY * 150 / 100)) ]; then
          echo "Canary latency too high: $LATENCY vs $PROD_LATENCY"
          exit 1
        fi

        if [ $GOLD_RATE -lt 90 ]; then
          echo "Canary gold rate too low: $GOLD_RATE"
          exit 1
        fi

        echo "Canary healthy"

    - name: Expand Rollout
      if: success()
      run: ./expand-rollout.sh

    - name: Rollback Canary
      if: failure()
      run: ./rollback-canary.sh

Benefits: Quantified canary health. Automatic comparison to baseline. Expansion gated on quality.

Pattern 4: Scheduled Validation

When: On a schedule (nightly, weekly)

What: Comprehensive validation without blocking deploys

snippet
# .github/workflows/nightly-validation.yml
name: Nightly Agent Validation

on:
  schedule:
    - cron: '0 2 * * *'  # 2 AM daily

jobs:
  validate:
    runs-on: ubuntu-latest
    steps:
      - name: Full Validation
        run: |
          for AGENT_ID in $PRODUCTION_AGENTS; do
            ./scripts/full-validation.sh $AGENT_ID
          done
      - name: Generate Report
        run: ./scripts/generate-report.sh
      - name: Send Report
        run: ./scripts/send-report.sh

Benefits: Comprehensive coverage. Historical trend data. Doesn't block deploys.


03Section

The Validation Script

A reusable script for all patterns:

snippet
#!/bin/bash
# scripts/validate-agent.sh

AGENT_ID=$1
TIMEOUT=${2:-120}  # seconds
INTERVAL=${3:-5}   # seconds

# Trigger validation
echo "Triggering validation for $AGENT_ID..."
RESPONSE=$(curl -s -X POST \
  "https://api.fabric.carmel.so/api/agentstatus/agents/$AGENT_ID/run" \
  -H "Authorization: Bearer $AGENTSTATUS_API_KEY")

DECISION_ID=$(echo $RESPONSE | jq -r '.decision_id')

if [ "$DECISION_ID" = "null" ] || [ -z "$DECISION_ID" ]; then
  echo "Failed to trigger validation"
  exit 1
fi

echo "Decision ID: $DECISION_ID"

# Poll for results
ELAPSED=0
while [ $ELAPSED -lt $TIMEOUT ]; do
  sleep $INTERVAL
  ELAPSED=$((ELAPSED + INTERVAL))

  RESULT=$(curl -s \
    "https://api.fabric.carmel.so/api/agentstatus/runs/$DECISION_ID" \
    -H "Authorization: Bearer $AGENTSTATUS_API_KEY")

  STATUS=$(echo $RESULT | jq -r '.status // "pending"')

  if [ "$STATUS" = "completed" ]; then
    VERDICT=$(echo $RESULT | jq -r '.overall_verdict')
    LATENCY=$(echo $RESULT | jq '.summary.latency_p50_ms')
    GOLD=$(echo $RESULT | jq '.summary.gold_pass_rate')

    echo "Verdict: $VERDICT"
    echo "Latency P50: ${LATENCY}ms"
    echo "Eval Pass Rate: ${GOLD}%"

    if [ "$VERDICT" = "DOWN" ]; then
      echo "❌ FAILED"
      exit 1
    elif [ "$VERDICT" = "DEGRADED" ]; then
      echo "⚠️ DEGRADED"
      exit 0  # or exit 1 if you want to fail on degraded
    else
      echo "✅ PASSED"
      exit 0
    fi
  fi

  echo "Status: $STATUS (${ELAPSED}s elapsed)"
done

echo "⏱️ Timeout after ${TIMEOUT}s"
exit 1

04Section

Advanced Patterns

Quality Gates

Define quality thresholds and enforce them:

snippet
quality-gates:
  min_gold_pass_rate: 95
  max_latency_p50_ms: 2000
  max_latency_p95_ms: 5000
  required_regions:
    - us
    - eu

Baseline Comparison

Compare to historical baseline:

snippet
LATENCY_THRESHOLD = 1.5  # 50% regression
GOLD_THRESHOLD = 0.95    # 5% drop

if current['latency_p50_ms'] > baseline['latency_p50_ms'] * LATENCY_THRESHOLD:
    print("Latency regression detected")
    sys.exit(1)

if current['gold_pass_rate'] < baseline['gold_pass_rate'] * GOLD_THRESHOLD:
    print("Quality regression detected")
    sys.exit(1)

Multi-Agent Coordination

Validate multiple agents together using matrix strategies:

snippet
strategy:
  matrix:
    agent:
      - name: support-bot
        id: agent_001
      - name: sales-bot
        id: agent_002
      - name: docs-bot
        id: agent_003

05Section

Best Practices

  1. Don't Block on Transient Failures - Use thresholds and retry logic. A single failed test shouldn't fail your deploy.
  2. Cache Results - If you need to reference validation results multiple times, cache them to a temp file.
  3. Set Reasonable Timeouts - Agent validation takes 30-60 seconds typically. Set timeouts accordingly (2-3 minutes is safe).
  4. Handle Degraded Appropriately - DEGRADED doesn't always mean "block deploy." Block during critical launches, warn during normal deploys.
  5. Keep Credentials Secure - Use secret management. Never hardcode API keys.

MCP validation

Your biggest blind spots sit outside your logs, inside hosts and paths you don't control.

We show what breaks before it hits your wire, and what to fix first.