Armory
Source

Leaderboard

Scored on public signals; components with none are listed as Unranked · Formula

RankScoreComponentDescriptionEvidenceLast commitInstall
174.422baserunobservability · observabilityBaserun captures LLM traces via a lightweight decorator-based SDK and provides a dashboard for debugging prompt chains, testing variants, and…16 stars · 5 forksStaleNo one-command install · Source
275.153honeycombobservability · observabilityHoneycomb's OpenTelemetry-native observability platform — high-cardinality event store ideal for tracing LLM pipelines and debugging slow agent…16 stars · 6 forksNo one-command install · Source
394.644athina-aiobservability · observabilityAthina AI provides developer-focused LLM monitoring and eval framework — real-time inference logging, automated evals, and regression detection in CI.301 stars · 23 forksNo one-command install · Source
495.863phosphoobservability · observabilityPhospho is a text analytics and evaluation platform for LLM apps — logs sessions, runs clustering, detects failures, and surfaces actionable insights.439 stars · 35 forksNo one-command install · Source
595.990claude-code-statuslineobservability · observabilityEnhanced 4-line statusline for Claude Code with themes, cost tracking, and MCP server monitoring476 stars · 35 forksNo one-command install · Source
698.121openinferenceobservability · observabilityOpenInference is an open standard and Python/JS instrumentation library for capturing LLM and agent traces in OpenTelemetry format, built by Arize AI.1,192 stars · 302 forksNo one-command install · Source
798.667openlitobservability · observabilityOpenLIT is an OpenTelemetry-native LLM observability toolkit with GPU monitoring, cost tracking, and a prompt hub — one-line setup for 20+ providers.2,736 stars · 367 forksNo one-command install · Source
898.792langwatchobservability · observabilityLangWatch provides real-time LLM analytics, guardrails, and evaluation pipelines with a visual studio for monitoring multi-step agent conversations.3,522 stars · 362 forksNo one-command install · Source
998.822pydantic-logfireobservability · observabilityLogfire by Pydantic — OpenTelemetry-based structured logging and tracing for Python applications with built-in support for FastAPI, SQLAlchemy, and…4,450 stars · 283 forksNo one-command install · Source
1099.046grafana-tempoobservability · observabilityGrafana Tempo is a cost-efficient distributed tracing backend (OpenTelemetry-native) that pairs with Loki for logs and Prometheus for metrics in LLM…5,461 stars · 746 forksNo one-command install · Source
1199.089heliconeobservability · observabilityOpen-source LLM observability platform — proxy-based logging, cost tracking, caching, and rate limiting for OpenAI-compatible APIs.6,123 stars · 662 forksNo one-command install · Source
1299.249ccstatuslineobservability · observabilityA highly customizable status line formatter for Claude Code CLI that displays model info, git branch, token usage, and other metrics in your terminal.13,035 stars · 579 forksNo one-command install · Source
1399.334portkey-ai-gatewayobservability · observabilityOpen-source AI gateway providing a unified API across 100+ LLM providers with built-in observability, request logging, fallbacks, caching, and load…12,873 stars · 1,280 forksNo one-command install · Source
1499.472opentelemetry-genaiobservability · observabilityOpenTelemetry semantic conventions and instrumentation for GenAI/LLM spans, traces, and metrics via the GenAI semconv working group.4,897 stars · 3,851 forksNo one-command install · Source
1599.490comet-opikobservability · observabilityOpik by Comet is an open-source LLM evaluation and tracing platform — log traces, run automated evals, create datasets, and track prompt improvements…22,249 stars · 1,827 forks · 1 mentionNo one-command install · Source
1699.677langfuseobservability · observabilityOpen-source LLM engineering platform with traces, evals, prompt management, and datasets for debugging and improving LLM applications.34,067 stars · 3,678 forks · 10 mentionsNo one-command install · Source
1799.743sentry-llm-monitoringobservability · observabilitySentry's error and performance monitoring extended to LLM applications — captures exceptions, latency, and AI token usage with OpenTelemetry…44,709 stars · 4,832 forks · 15 mentionsNo one-command install · Source
18Unrankeddatadog-llm-observabilityobservability · observabilityDatadog's managed LLM Observability product — traces LLM calls, monitors prompt/completion quality, detects anomalies, and integrates with existing…No signals yetNo commit datelisted No one-command install · Source
19Unrankedfiddler-aiobservability · observabilityFiddler AI Observability platform monitors LLM applications for hallucinations, toxicity, bias, and drift, with explainability and alerting for…No signals yetNo commit datelisted No one-command install · Source
20Unrankedhoneyhiveobservability · observabilityHoneyHive is an AI evaluation and observability platform for tracing agent pipelines, running evaluations, and debugging regressions in production.No signals yetNo commit datelisted No one-command install · Source
21Unrankedlangtraceobservability · observabilityOpen-source, OpenTelemetry-compliant LLM observability tool by Scale3Labs — traces calls to all major LLM providers and frameworks with a…No signals yetNo commit datelisted No one-command install · Source
22Unrankedliteral-aiobservability · observabilityLiteral AI is an observability and evaluation platform for conversational AI — captures multi-step threads, scores responses, and integrates with…No signals yetNo commit datelisted No one-command install · Source
23Unrankedlunaryobservability · observabilityOpen-source LLM observability and prompt management platform — tracks conversations, errors, costs, and user feedback for production AI applications.No signals yetNo commit datelisted No one-command install · Source
24Unrankedmaxim-aiobservability · observabilityMaxim AI is an evaluation and observability platform for AI agents — supports multi-step trace analysis, prompt testing, and production quality…No signals yetNo commit datelisted No one-command install · Source
25Unrankednew-relic-ai-monitoringobservability · observabilityNew Relic AI Monitoring instruments LLM calls end-to-end — traces model invocations, measures token costs, and surfaces anomalies via the New Relic…No signals yetNo commit datelisted No one-command install · Source

Score colour shows how many signals stand behind it, never how good it is: amber, three or more; dimmer amber, two; grey, one. The Evidence column names them.

Stars, forks and last commit are as GitHub reported them when Armory last read each repository: for most, or later. A repository may have changed since.

Leaderboard · Armory