Armory
Source

Observability

Traces, cost and errors from each run

Listed
34
Ranked
76.5%
Top Score
99.7

Pick

The pick and its runners-up, in score order

  • langfuse99.63 signals

    The pick · #3 on this shelf

    Traces, prompt management, datasets and evals in one self-hostable platform

    The two rows above it are tracing add-ons to Sentry and MLflow; this is a whole platform on its own.

    34,067 stars · 3,678 forks · 10 mentions
    No one-command install · Source
  • comet-opik99.43 signals

    Runner-up · #5 on this shelf

    Open-source tracing and evaluation: log traces, run evals, track prompt changes

    Built like the pick; above it are two tracing add-ons, the pick and a Claude Code status line.

    22,249 stars · 1,827 forks · 1 mention
    No one-command install · Source

Top Ranked

Needs the armory CLI · not on npm yet, build it from cli/ in the repository

RankScoreComponentDescriptionEvidenceLast commitInstall
199.743sentry-llm-monitoringobservability · observabilitySentry's error and performance monitoring extended to LLM applications — captures exceptions, latency, and AI token usage with OpenTelemetry…44,709 stars · 4,832 forks · 15 mentionsNo one-command install · Source
299.684mlflow-tracingobservability · ai-agentsMLflow's LLM tracing module instruments model calls, agent steps, and tool invocations, storing them alongside experiment runs for reproducibility.27,768 stars · 6,246 forksNo one-command install · Source
399.677langfuseobservability · observabilityOpen-source LLM engineering platform with traces, evals, prompt management, and datasets for debugging and improving LLM applications.34,067 stars · 3,678 forks · 10 mentionsNo one-command install · Source
499.526claude-hudobservability · ai-agentsA status line for Claude Code that shows context usage, tools, agents, to-dos and more. Highly configurable, and maintained when it was listed.27,778 stars · 1,286 forksNo one-command install · Source
599.490comet-opikobservability · observabilityOpik by Comet is an open-source LLM evaluation and tracing platform — log traces, run automated evals, create datasets, and track prompt improvements…22,249 stars · 1,827 forks · 1 mentionNo one-command install · Source
699.472opentelemetry-genaiobservability · observabilityOpenTelemetry semantic conventions and instrumentation for GenAI/LLM spans, traces, and metrics via the GenAI semconv working group.4,897 stars · 3,851 forksNo one-command install · Source
799.334portkey-ai-gatewayobservability · observabilityOpen-source AI gateway providing a unified API across 100+ LLM providers with built-in observability, request logging, fallbacks, caching, and load…12,873 stars · 1,280 forksNo one-command install · Source
899.249ccstatuslineobservability · observabilityA highly customizable status line formatter for Claude Code CLI that displays model info, git branch, token usage, and other metrics in your terminal.13,035 stars · 579 forksNo one-command install · Source
999.194openllmetryobservability · ai-agentsOpenTelemetry-based observability for LLM applications — auto-instruments OpenAI, Anthropic, LangChain, and 20+ providers with zero code changes.7,452 stars · 1,099 forksNo one-command install · Source
1099.089heliconeobservability · observabilityOpen-source LLM observability platform — proxy-based logging, cost tracking, caching, and rate limiting for OpenAI-compatible APIs.6,123 stars · 662 forksNo one-command install · Source
1199.046grafana-tempoobservability · observabilityGrafana Tempo is a cost-efficient distributed tracing backend (OpenTelemetry-native) that pairs with Loki for logs and Prometheus for metrics in LLM…5,461 stars · 746 forksNo one-command install · Source
1298.822pydantic-logfireobservability · observabilityLogfire by Pydantic — OpenTelemetry-based structured logging and tracing for Python applications with built-in support for FastAPI, SQLAlchemy, and…4,450 stars · 283 forksNo one-command install · Source
1398.792langwatchobservability · observabilityLangWatch provides real-time LLM analytics, guardrails, and evaluation pipelines with a visual studio for monitoring multi-step agent conversations.3,522 stars · 362 forksNo one-command install · Source
1498.669ccometixline-claude-code-statuslineobservability · ai-agentsA high-performance Claude Code statusline tool written in Rust with Git integration, usage tracking, interactive TUI configuration, and Claude Code…3,456 stars · 215 forksNo one-command install · Source
1598.667openlitobservability · observabilityOpenLIT is an OpenTelemetry-native LLM observability toolkit with GPU monitoring, cost tracking, and a prompt hub — one-line setup for 20+ providers.2,736 stars · 367 forksNo one-command install · Source
1698.624laminarobservability · ai-agentsLaminar is an open-source platform for tracing, evaluating, and labeling LLM and agent pipelines with a TypeScript/Python SDK and a self-hostable…3,218 stars · 229 forks · 1 mentionNo one-command install · Source
1798.121openinferenceobservability · observabilityOpenInference is an open standard and Python/JS instrumentation library for capturing LLM and agent traces in OpenTelemetry format, built by Arize AI.1,192 stars · 302 forksNo one-command install · Source
1898.005langsmithobservability · back-endLangChain's platform for tracing, evaluating, and monitoring LLM applications — deep integration with LangChain/LangGraph plus a REST API for any…1,043 stars · 288 forksNo one-command install · Source
1997.619claude-powerlineobservability · ai-agentsA vim-style powerline statusline for Claude Code with real-time usage tracking, git integration, custom themes, and more1,163 stars · 82 forksNo one-command install · Source
2095.990claude-code-statuslineobservability · observabilityEnhanced 4-line statusline for Claude Code with themes, cost tracking, and MCP server monitoring476 stars · 35 forksNo one-command install · Source

Score colour shows how many signals stand behind it, never how good it is: amber, three or more; dimmer amber, two; grey, one. The Evidence column names them.

Stars, forks and last commit are as GitHub reported them when Armory last read each repository: for most, or later. A repository may have changed since.

Leaderboard ranks all 26 scored rows. Formula shows the calculation.