Radar
← All projectsSaturday, August 22, 2026
Tracely
Trace-native CI for agents: production failures freeze into hermetic cases that block the PR for $0 replay.
Why it matters
Tracely is an open-source, self-hostable stack for agent observability that closes the loop into CI. Traces land over OTLP with agent, conversation, and turn columns. Online evaluators grade runs as they arrive. Failures cluster into issues. One click freezes a bad run into a hermetic regression case with recorded tool and LLM fixtures, then tracely gate replays the suite offline on every PR with no live model spend. Docker Compose and a Railway one-click deploy bring up API, worker, UI, Postgres, ClickHouse, Redis, and MinIO. There is an MCP endpoint so coding agents can drive the workspace, plus alerts that can hit Slack, email, or webhooks.
Most agent eval stacks make you invent datasets. Real breakage already shows up in production traces with the exact tools and model turns. Founders shipping agents need those failures to become gates, not another dashboard tab you check after the customer complains.
How it works
Clone and run docker compose --profile demo up to open a seeded UI on localhost:3001. Point your agent with pip install tracely-ai and tracely.init against your endpoint and ingest key, then wrap runs in tracely.trace. Wire structural checks and LLM judges as columns on the trace table. Promote a failing cluster into a case, attach fail-to-pass contracts, and run tracely gate or the GitHub Action so a regression fails the PR with a commit status and comment.
This is not another coding harness and not a hand-written eval spreadsheet. The recorded production run is the test. Replay is deterministic and free of API keys in CI, while live judges still watch prod. You get observe, cluster, freeze, gate, and alert in one product instead of stitching three SaaS tools.
Capabilities
- Public demo available
- API / SDK surface
- MCP server / client support
Similar tools