1.0 Detect
TracingSee exactly what your agents are doing.Every LLM call, agent step, and tool invocation captured as a structured trace — with cost, latency, and tokens per span.< 5 min setup · OTel native OTLP/HTTP · 100 events per batchAgent MonitoringKnow which agents are healthy, and which aren't.Composite health scores, delegation graphs, and per-agent cost attribution — built for systems with many cooperating AI agents.A–F health grades · 0–100 composite score · 3 signals error · cost · evalAlerts & IncidentsGet paged before your users notice.Alert rules on error rate, latency, cost, and eval quality. Multi-channel notifications. Full incident lifecycle with AI-generated postmortems.5 incident states · 3 alert metric targets · 4 channelsCost ManagementCut your LLM bill before it cuts your runway.Attribution across 5 dimensions, 30-day forecasting with confidence bands, and AI-generated optimization recommendations with savings estimates.30 days forecast horizon · 5 attribution dimensions · Top 200 per dimension
2.0 Prevent
GuardrailsStop bad outputs before they reach users.7 guardrail types run inline on every LLM request — block, warn, redact, or log. PII, toxicity, topic drift, format, cost ceiling, and custom rules.7 types · 4 actions · 50ms min latency capEvaluationsMeasure output quality on every trace, automatically.12 built-in LLM-as-judge templates run on every new trace with no setup. Track quality trends, catch regressions, and run manual eval campaigns.12 built-in templates · 200 metric keys · 0–1 sample rateSimulationsTest your AI app against real data before deploying.Run up to 100 scenarios per batch against named datasets. Turn production failures into regression tests in one click.100 scenarios per batch · 500 items per dataset · 3 scenario typesPrompt ManagementShip prompt changes without breaking production.Version history, production promotion, automatic regression detection after every deploy, and AI-powered optimization suggestions.14 days regression lookback · 10% regression threshold · SHA-256 content hash
3.0 Fix
ZespanPilotAsk your production data anything, in plain English.Conversational AI copilot that queries your agent data in plain English, takes actions on your behalf, and knows what you're looking at right now.NLQ natural language · CSV export & reports · ✓ approval workflowPlaygroundFind prompt failures in the sandbox, not in production.Test prompts across 4 providers with real streaming, tool calls, structured output, and your actual guardrails — before any code ships.4 providers · 100+ models · Live streaming