Your agent didn’t crash.
It just started being wrong.

Every request returned 200.

A user found out before you did. Zespan traces every span, catches the regression, and hands you a fix tested against the failure that caused it.

NewZespanPilot: insight and action in one chat →
Zespan system health dashboard: error rate, p95 latency, spend and request count with trend sparklines, a what-to-fix-first panel ranked by cost of inaction, and a ZespanPilot digest attributing latency to a single agent

Auto-detects the stack you already run

OpenAIOpenAIAnthropicAnthropicGoogle GeminiGoogle GeminiAWS BedrockAWS BedrockLangChainLangChainCrewAICrewAILlamaIndexLlamaIndexVercel AI SDKVercel AI SDKOpenTelemetryOpenTelemetryMistralMistralGroqGroqAutoGenAutoGen

1.0  Detect

Know the moment quality slips,
not when a user complains.

Full-stack understanding

Span waterfallDelegation graphsSession groupingCost attributionOTLP ingestAI root cause

Every agent run is captured the moment it happens, with no sampling and no manual instrumentation. Verdict-based clustering groups failures by what actually went wrong, and AI correlation ties the alert to the trace with a root cause already attached.

traces12,480 traces · 1.8% error rate
Zespan Tracing — 12,480 traces · 1.8% error rate

2.0  Prevent

Turn the trace that broke something
into a rule that stops it twice.

Gated by default

7 guardrail types12 judge templatesPRE + POST stagesNear-miss captureRegression datasetsTrust ledger

Flag a trace and it becomes an enforced guardrail policy in one action, no separate config step. Twelve built-in LLM-as-judge templates score every trace, and near-misses are queued as suggested rules for a human to approve.

guardrails9 protections · 82 checks
Zespan Guardrails — 9 protections · 82 checks

3.0  Fix

Ask it anything, then let it
ship the fix you approve.

Human in the loop

Dry-run previewRole-gated actionsReplay against failuresFull audit logCost-quality scoringOne-tap undo

ZespanPilot answers in plain English and takes action across ~40 registered operations, every mutation role-gated and audit-logged. When an incident closes, Zespan proposes the fix and replays it against the failures that incident captured.

fix-proposal · incident_412gated
01

Incident #412 detected

severity: high · agent: RefundAgent

02

Fix proposed

prompt v9 → v10 · from the incident’s own captured failures

03

Gate: replayed against 41 captured failures

39/41 now pass · 2 regressions flagged

04

Awaiting approval

nothing ships until a human clicks approve

StagingProductionRejectillustrative

The closed loop

When an incident closes, Zespan generates a fix candidate and replays it against every failure that incident actually captured, so you see the pass rate before you decide anything.

Every approved or rejected candidate becomes a permanent regression test, pulled from a real production failure, never a synthetic one. Nothing ships automatically.

Closed-loop fix proposal
Regression tests from real failures
Cost-quality frontier
zespanpilot~40 actions · 9 categories