Armory
Source

Leaderboard

Scored on public signals; components with none are listed as Unranked · Formula

RankScoreComponentDescriptionEvidenceLast commitInstall
199.743sentry-llm-monitoringobservability · observabilitySentry's error and performance monitoring extended to LLM applications — captures exceptions, latency, and AI token usage with OpenTelemetry…44,709 stars · 4,832 forks · 15 mentionsNo one-command install · Source
299.677langfuseobservability · observabilityOpen-source LLM engineering platform with traces, evals, prompt management, and datasets for debugging and improving LLM applications.34,067 stars · 3,678 forks · 10 mentionsNo one-command install · Source
399.490comet-opikobservability · observabilityOpik by Comet is an open-source LLM evaluation and tracing platform — log traces, run automated evals, create datasets, and track prompt improvements…22,249 stars · 1,827 forks · 1 mentionNo one-command install · Source
499.472opentelemetry-genaiobservability · observabilityOpenTelemetry semantic conventions and instrumentation for GenAI/LLM spans, traces, and metrics via the GenAI semconv working group.4,897 stars · 3,851 forksNo one-command install · Source
599.455deepevaleval · observabilityOpen-source LLM evaluation framework with 14+ metrics (hallucination, faithfulness, answer relevancy) and CI support.18,041 stars · 1,890 forksNo one-command install · Source
699.334portkey-ai-gatewayobservability · observabilityOpen-source AI gateway providing a unified API across 100+ LLM providers with built-in observability, request logging, fallbacks, caching, and load…12,873 stars · 1,280 forksNo one-command install · Source
799.089heliconeobservability · observabilityOpen-source LLM observability platform — proxy-based logging, cost tracking, caching, and rate limiting for OpenAI-compatible APIs.6,123 stars · 662 forksNo one-command install · Source
899.046grafana-tempoobservability · observabilityGrafana Tempo is a cost-efficient distributed tracing backend (OpenTelemetry-native) that pairs with Loki for logs and Prometheus for metrics in LLM…5,461 stars · 746 forksNo one-command install · Source
998.667openlitobservability · observabilityOpenLIT is an OpenTelemetry-native LLM observability toolkit with GPU monitoring, cost tracking, and a prompt hub — one-line setup for 20+ providers.2,736 stars · 367 forksNo one-command install · Source
1098.121openinferenceobservability · observabilityOpenInference is an open standard and Python/JS instrumentation library for capturing LLM and agent traces in OpenTelemetry format, built by Arize AI.1,192 stars · 302 forksNo one-command install · Source
1197.880langtraceeval · observabilityOpen-source observability tool for LLMs with OpenTelemetry-based tracing, automated evals, and annotation workflows.1,228 stars · 127 forksNo one-command install · Source
1294.644athina-aiobservability · observabilityAthina AI provides developer-focused LLM monitoring and eval framework — real-time inference logging, automated evals, and regression detection in CI.301 stars · 23 forksNo one-command install · Source
1393.031signoz-mcp-servermcp · observabilityEnables AI assistants and LLMs to query SigNoz observability data (metrics, traces, logs, alerts, dashboards) using natural language.117 stars · 43 forksNo one-command install · Source
1491.841avivsinai-langfuse-mcpmcp · observabilityQuery Langfuse traces, debug exceptions, analyze sessions, and manage prompts. Full observability toolkit for LLM applications.105 stars · 24 forksNo one-command install · Source
1590.950thinkmcp · observabilityProvides a lightweight 'think' tool for structured reasoning, enabling LLMs to pause, log thoughts, and improve multi-step problem solving without…106 stars · 15 forksNo one-command install · Source
1690.750multi-model-advisormcp · observabilityQueries multiple Ollama models in parallel with distinct system prompts focused on empathy, logic, and creativity to provide diverse perspectives on…86 stars · 20 forksNo one-command install · Source
1790.554vellum-evalseval · observabilityVellum evaluation SDK for running LLM test suites with custom metrics, dataset pinning, and CI workflow integration.82 stars · 20 forksNo one-command install · Source
1882.655composer-trade-mcpmcp · observabilityEnables MCP-enabled LLMs to create, backtest, and trade automated investing strategies (symphonies) on Composer, with tools for monitoring performance…2 stars · 45 forksNo one-command install · Source
1982.319braintrusteval · observabilityDeveloper platform for logging, evaluating, and comparing LLM experiments with dataset versioning and scoring functions.27 stars · 12 forksNo one-command install · Source
2080.342galileo-evaluateeval · observabilityGalileo evaluation and observability SDK for detecting hallucinations, data errors, and model weaknesses in LLM pipelines.22 stars · 11 forksNo one-command install · Source
2175.153honeycombobservability · observabilityHoneycomb's OpenTelemetry-native observability platform — high-cardinality event store ideal for tracing LLM pipelines and debugging slow agent…16 stars · 6 forksNo one-command install · Source
2274.422baserunobservability · observabilityBaserun captures LLM traces via a lightweight decorator-based SDK and provides a dashboard for debugging prompt chains, testing variants, and…16 stars · 5 forksStaleNo one-command install · Source
2371.115tencent-cloud-log-service-cls-mcp-servermcp · observabilityEnables large language models to directly access Tencent Cloud Log Service for log search, metric queries, and alarm management without code.11 stars · 6 forksNo one-command install · Source
2371.115background-process-mcpmcp · observabilityalso listed as waylaidwanderer-background-processEnables LLMs to start, stop, and monitor long-running command-line processes in the background.11 stars · 6 forksNo one-command install · Source
2670.060galileomcp · observabilityIntegrates with Galileo's evaluation and observability platform to enable dataset creation, prompt template management, experiment setup, log…6 stars · 7 forksNo one-command install · Source
2759.040spanlens-mcpmcp · observabilityMCP-native LLM observability. Query your Spanlens traces, stats, cost anomalies, and savings from Cursor, Claude Desktop, or any MCP client. Open…13 starsNo one-command install · Source
2859.029aplaceforallmystuff-piholemcp · observabilityManage DNS blocking, monitor traffic stats, and control whitelist/blacklist for Pi-hole v68 stars · 1 forkNo one-command install · Source
2958.969tilt-mcp-servermcp · observabilityEnables LLMs and AI assistants to interact with Tilt development environments, providing tools to list resources, fetch logs, and monitor status.4 stars · 4 forksNo one-command install · Source
2958.969ansible-aapmcp · observabilityEnables LLMs to discover and launch Ansible job templates on Ansible Automation Platform, monitor job status, and retrieve outputs.4 stars · 4 forksNo one-command install · Source
3157.245aplaceforallmystuff-tailscalemcp · observabilityProvides read-only access to Tailscale network management for monitoring device status, tracking client updates, and generating network statistics…7 stars · 1 forkNo one-command install · Source
3255.515rhoai-observability-mcpmcp · observabilityProvides AI assistants with direct access to Red Hat OpenShift AI observability data, enabling querying of Prometheus metrics, Alertmanager alerts…5 stars · 2 forksNo one-command install · Source
3347.273log-logn-langfuse-mcp-javamcp · observabilityQuery Langfuse traces, debug exceptions, analyze sessions, scores, datasets, schema, observations and manage prompts. Full observability toolkit for…3 stars · 2 forksNo one-command install · Source
3443.275jungle-grid-mcp-servermcp · observabilityMCP server for Jungle Grid, an agentic GPU execution layer that lets AI agents estimate, submit, monitor, and fetch logs for inference, training…4 starsNo one-command install · Source
3542.840scorecardmcp · observabilityEvaluate and optimize LLM systems with comprehensive testing and metrics3 forksNo one-command install · Source
3632.223berserk-mcpmcp · observabilityEnables LLMs to answer Berserk observability questions by calling verified KQL tools instead of hand-authoring queries, with role-based tool filtering…2 starsNo one-command install · Source
3632.223odigo-elastic-s2l-mcpmcp · observabilityConnects LLMs to Elasticsearch with a Semantic-to-Lexical layer that translates technical field names into business knowledge, enabling autonomous…2 starsNo one-command install · Source
3822.557logicmcp-servermcp · observabilityEnables traceable requirement discovery, technical alignment, and ISO-aligned process checking through deterministic MCP tools and resources, without…1 starNo one-command install · Source
3822.557ebpf-mcp-tracermcp · observabilityEnables LLMs to safely write and run bpftrace scripts against the Linux kernel for observability, with explicit probe allowlists and execution…1 starNo one-command install · Source
3822.557llm-brand-monitormcp · observabilityTrack how 350+ AI models mention your brand — monitor visibility, sentiment, and competitor mentions.1 starNo one-command install · Source
3822.557mcp-delonghi-ecammcp · observabilityEnables LLMs to control DeLonghi ECAM espresso machines over a local network, including brewing beverages, monitoring status, and managing machine…1 starNo one-command install · Source
3822.557mi25-tuning-mcpmcp · observabilityManages multi_llm-client operations for MI25/gfx900 GPUs, enabling configuration updates, inference execution, benchmarking, and performance log…1 starNo one-command install · Source
3822.557mcp-log-analyzermcp · observabilityAnalyzes log files locally using Ollama and files structured GitHub Issues automatically, with all processing kept on your machine.1 starNo one-command install · Source
3822.557oalles-agentic-system-monitoring-ragmcp · observabilitySpring Boot-based server that connects system monitoring tools with a RAG service, enabling real-time access to both system metrics and corporate…1 starNo one-command install · Source
45Unrankeddynatrace-saas-mcp-servermcp · observabilityEnables LLM agents to query Dynatrace SaaS for observability data (logs, metrics, traces, entities, problems, vulnerabilities) and manage…No signals yetNo one-command install · Source
46Unrankedhome-network-mcpmcp · observabilityLets an LLM client monitor devices, services, disk usage, and uptime in a home network and home lab.No signals yetNo one-command install · Source
47Unrankedml-lab-mcpmcp · observabilityEnables LLMs to manage and run machine learning training jobs on a remote server, including syncing code, submitting experiments, monitoring progress…No signals yetNo one-command install · Source
48Unrankedinference-aiopsmcp · observabilityEnables governance-grade AIops for GPU inference clusters with root-cause analysis, metrics, and policy-governed operations for vLLM and Ray.No signals yetNo one-command install · Source
49Unrankedworkloadtruth-mcp-servermcp · observabilityEnables classification of GPU workloads as training, inference, or idle from telemetry data, with tools for one-shot classification, benchmarking, and…No signals yetNo one-command install · Source
50Unrankedlocal-worker-mcpmcp · observabilityDelegates heavy, repetitive, and verifiable tasks like PDF extraction, code analysis, and log processing to a local LLM to reduce token consumption…No signals yetNo one-command install · Source
51Unrankedunitree-go2-mcp-servermcp · observabilityEnables LLM agents to control and monitor a Unitree Go2 robot through MCP tools, including live telemetry, navigation, camera feeds, and waypoint…No signals yetNo one-command install · Source
52Unrankedmcp-devops-dashboardmcp · observabilityMonitors host CPU/RAM, local ports, and Docker containers, streaming live telemetry to a React dashboard over SSE. Exposes MCP tools to query system…No signals yetNo one-command install · Source
53Unrankedio-mcpmcp · observabilityEnables local research workflows (paper discovery, relevance scoring, digests) and homelab monitoring (Prometheus, logs) through an MCP server, using…No signals yetNo one-command install · Source
54Unrankedapple-health-semantic-mcpmcp · observabilityQuery Apple Health export data with an LLM via a semantic layer that avoids common data traps, enabling accurate natural-language queries about health…No signals yetNo one-command install · Source
55Unrankedcodex-worker-runtimemcp · observabilityEnables Codex to delegate bounded work to external LLMs through role-based MCP tools, with worker health checks and audit logging.No signals yetNo one-command install · Source
56Unrankedros-llm-integration-bridgemcp · observabilityEnables large language models to interact with ROS robots seamlessly, allowing natural language control, real-time sensor monitoring, and autonomous…No signals yetNo one-command install · Source
57Unrankedmcp-hayabusamcp · observabilityEnables an LLM client to scan Windows event log files (EVTX) for suspicious activity using Hayabusa, and browse its detection rules directly in…No signals yetNo one-command install · Source
58Unrankedcloudwatch-mcp-agentmcp · observabilityEnables natural language queries for AWS CloudWatch logs, metrics, and alarms via an LLM agent with MCP tools.No signals yetNo one-command install · Source
59Unrankedemotion-mcpmcp · observabilityProvides a dynamic emotion simulation system for AI character roleplay, based on Freudian psychodynamics, that analyzes dialogue content via LLM to…No signals yetNo one-command install · Source
60Unrankedes-error-lensmcp · observabilityEnables LLM agents to search and analyze Elasticsearch logs for errors, detect recurring patterns, analyze error-rate trends, and retrieve full trace…No signals yetNo one-command install · Source
61Unrankedtracepii-shield-mcpmcp · observabilityRedacts PII from LLM traces and tool payloads before they leave review, enabling PII scanning, payload redaction, sensitive field classification…No signals yetNo one-command install · Source
62Unrankedmdl-train-mcpmcp · observabilityAn MCP server for monitoring and managing training jobs on Modal. Built for LLMs that need to check on long-running GPU training without drowning in…No signals yetNo one-command install · Source
63Unrankedsession-logger-mcpmcp · observabilityEnables saving LLM chat conversations to JSONL log files and querying them by session ID, user ID, keyword, or date range.No signals yetNo one-command install · Source
64Unrankedagent-evaluationskill · observabilityTesting and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top…No signals yetNo commit datelisted armory install agent-evaluation --cli claude
65Unrankedagentlens-mcp-servermcp · observabilityEnables AI agents to access observability and evaluation data, including run history, span traces, LLM-as-judge evaluation results, and regression…No signals yetNo commit datelisted No one-command install · Source
66Unrankeddatadog-llm-observabilityobservability · observabilityDatadog's managed LLM Observability product — traces LLM calls, monitors prompt/completion quality, detects anomalies, and integrates with existing…No signals yetNo commit datelisted No one-command install · Source
67Unrankedfast-award-screener-mcpmcp · observabilityScreens DIBBS RFQs for micro-purchase viability using deterministic logic without LLM calls.No signals yetNo commit datelisted No one-command install · Source
68Unrankedfiddler-aiobservability · observabilityFiddler AI Observability platform monitors LLM applications for hallucinations, toxicity, bias, and drift, with explainability and alerting for…No signals yetNo commit datelisted No one-command install · Source
69Unrankedhoneyhiveeval · observabilityLLM evaluation and experimentation platform with session tracing, dataset management, and metric-based run comparison.No signals yetNo commit datelisted No one-command install · Source
70Unrankedlangsmith-observabilityskill · observabilityLLM observability platform for tracing, evaluation, and monitoring. Use when debugging LLM applications, evaluating model outputs against datasets…No signals yetNo commit datelisted armory install langsmith-observability --cli claude
71Unrankedlangsmith-tracing-3rules · observabilityConfigure LangSmith tracing environment variables for Claude Code observability. Sends conversation traces to LangSmith for monitoring and analysis…No signals yetNo commit datelisted armory install langsmith-tracing-3 --cli claude
72Unrankedlangtraceobservability · observabilityOpen-source, OpenTelemetry-compliant LLM observability tool by Scale3Labs — traces calls to all major LLM providers and frameworks with a…No signals yetNo commit datelisted No one-command install · Source
73Unrankedllm-evaluationskill · observabilityMaster comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.No signals yetNo commit datelisted armory install llm-evaluation --cli claude
74Unrankedlunaryobservability · observabilityOpen-source LLM observability and prompt management platform — tracks conversations, errors, costs, and user feedback for production AI applications.No signals yetNo commit datelisted No one-command install · Source
75Unrankedml-engineer-2skill · observabilityBuild production ML systems with PyTorch 2.x, TensorFlow, and modern ML frameworks. Implements model serving, feature engineering, A/B testing, and…No signals yetNo commit datelisted armory install ml-engineer-2 --cli claude
76Unrankedmle-reviewersubagent · observabilityProduction machine-learning engineering reviewer for data contracts, feature pipelines, training reproducibility, offline/online evaluation, model…No signals yetNo commit datelisted armory install mle-reviewer --cli claude
77Unrankednew-relic-ai-monitoringobservability · observabilityNew Relic AI Monitoring instruments LLM calls end-to-end — traces model invocations, measures token costs, and surfaces anomalies via the New Relic…No signals yetNo commit datelisted No one-command install · Source
78Unrankedphoenix-observabilityskill · observabilityOpen-source AI observability platform for LLM tracing, evaluation, and monitoring. Use when debugging LLM applications with detailed traces, running…No signals yetNo commit datelisted armory install phoenix-observability --cli claude
79Unrankedstatsmodelsskill · observabilityStatistical modeling toolkit. OLS, GLM, logistic, ARIMA, time series, hypothesis tests, diagnostics, AIC/BIC, for rigorous statistical inference and…No signals yetNo commit datelisted armory install statsmodels --cli claude

Score colour shows how many signals stand behind it, never how good it is: amber, three or more; dimmer amber, two; grey, one. The Evidence column names them.

Stars, forks and last commit are as GitHub reported them when Armory last read each repository: for most, or later. A repository may have changed since.

Leaderboard · Armory