Baserun captures LLM traces via a lightweight decorator-based SDK and provides a dashboard for debugging prompt chains, testing variants, and measuring quality.
Fiddler AI Observability platform monitors LLM applications for hallucinations, toxicity, bias, and drift, with explainability and alerting for production AI.
Open-source, OpenTelemetry-compliant LLM observability tool by Scale3Labs — traces calls to all major LLM providers and frameworks with a self-hostable UI.
Literal AI is an observability and evaluation platform for conversational AI — captures multi-step threads, scores responses, and integrates with Chainlit.
Maxim AI is an evaluation and observability platform for AI agents — supports multi-step trace analysis, prompt testing, and production quality monitoring.
New Relic AI Monitoring instruments LLM calls end-to-end — traces model invocations, measures token costs, and surfaces anomalies via the New Relic platform.