Leaderboard
Scored on public signals; components with none are listed as Unranked · Formula
Component
Domain
Vertical
| Rank | Score | Component | Description | Evidence | Last commit | Install |
|---|---|---|---|---|---|---|
| 1 | 99.743 | sentry-llm-monitoringobservability · observability | Sentry's error and performance monitoring extended to LLM applications — captures exceptions, latency, and AI token usage with OpenTelemetry… | 44,709 stars · 4,832 forks · 15 mentions | No one-command install · Source | |
| 2 | 99.677 | langfuseobservability · observability | Open-source LLM engineering platform with traces, evals, prompt management, and datasets for debugging and improving LLM applications. | 34,067 stars · 3,678 forks · 10 mentions | No one-command install · Source | |
| 3 | 99.490 | comet-opikobservability · observability | Opik by Comet is an open-source LLM evaluation and tracing platform — log traces, run automated evals, create datasets, and track prompt improvements… | 22,249 stars · 1,827 forks · 1 mention | No one-command install · Source | |
| 4 | 99.472 | opentelemetry-genaiobservability · observability | OpenTelemetry semantic conventions and instrumentation for GenAI/LLM spans, traces, and metrics via the GenAI semconv working group. | 4,897 stars · 3,851 forks | No one-command install · Source | |
| 5 | 99.455 | deepevaleval · observability | Open-source LLM evaluation framework with 14+ metrics (hallucination, faithfulness, answer relevancy) and CI support. | 18,041 stars · 1,890 forks | No one-command install · Source | |
| 6 | 99.334 | portkey-ai-gatewayobservability · observability | Open-source AI gateway providing a unified API across 100+ LLM providers with built-in observability, request logging, fallbacks, caching, and load… | 12,873 stars · 1,280 forks | No one-command install · Source | |
| 7 | 99.089 | heliconeobservability · observability | Open-source LLM observability platform — proxy-based logging, cost tracking, caching, and rate limiting for OpenAI-compatible APIs. | 6,123 stars · 662 forks | No one-command install · Source | |
| 8 | 99.046 | grafana-tempoobservability · observability | Grafana Tempo is a cost-efficient distributed tracing backend (OpenTelemetry-native) that pairs with Loki for logs and Prometheus for metrics in LLM… | 5,461 stars · 746 forks | No one-command install · Source | |
| 9 | 98.667 | openlitobservability · observability | OpenLIT is an OpenTelemetry-native LLM observability toolkit with GPU monitoring, cost tracking, and a prompt hub — one-line setup for 20+ providers. | 2,736 stars · 367 forks | No one-command install · Source | |
| 10 | 98.121 | openinferenceobservability · observability | OpenInference is an open standard and Python/JS instrumentation library for capturing LLM and agent traces in OpenTelemetry format, built by Arize AI. | 1,192 stars · 302 forks | No one-command install · Source | |
| 11 | 97.880 | langtraceeval · observability | Open-source observability tool for LLMs with OpenTelemetry-based tracing, automated evals, and annotation workflows. | 1,228 stars · 127 forks | No one-command install · Source | |
| 12 | 94.644 | athina-aiobservability · observability | Athina AI provides developer-focused LLM monitoring and eval framework — real-time inference logging, automated evals, and regression detection in CI. | 301 stars · 23 forks | No one-command install · Source | |
| 13 | 93.031 | signoz-mcp-servermcp · observability | Enables AI assistants and LLMs to query SigNoz observability data (metrics, traces, logs, alerts, dashboards) using natural language. | 117 stars · 43 forks | No one-command install · Source | |
| 14 | 91.841 | avivsinai-langfuse-mcpmcp · observability | Query Langfuse traces, debug exceptions, analyze sessions, and manage prompts. Full observability toolkit for LLM applications. | 105 stars · 24 forks | No one-command install · Source | |
| 15 | 90.950 | thinkmcp · observability | Provides a lightweight 'think' tool for structured reasoning, enabling LLMs to pause, log thoughts, and improve multi-step problem solving without… | 106 stars · 15 forks | No one-command install · Source | |
| 16 | 90.750 | multi-model-advisormcp · observability | Queries multiple Ollama models in parallel with distinct system prompts focused on empathy, logic, and creativity to provide diverse perspectives on… | 86 stars · 20 forks | No one-command install · Source | |
| 17 | 90.554 | vellum-evalseval · observability | Vellum evaluation SDK for running LLM test suites with custom metrics, dataset pinning, and CI workflow integration. | 82 stars · 20 forks | No one-command install · Source | |
| 18 | 82.655 | composer-trade-mcpmcp · observability | Enables MCP-enabled LLMs to create, backtest, and trade automated investing strategies (symphonies) on Composer, with tools for monitoring performance… | 2 stars · 45 forks | No one-command install · Source | |
| 19 | 82.319 | braintrusteval · observability | Developer platform for logging, evaluating, and comparing LLM experiments with dataset versioning and scoring functions. | 27 stars · 12 forks | No one-command install · Source | |
| 20 | 80.342 | galileo-evaluateeval · observability | Galileo evaluation and observability SDK for detecting hallucinations, data errors, and model weaknesses in LLM pipelines. | 22 stars · 11 forks | No one-command install · Source | |
| 21 | 75.153 | honeycombobservability · observability | Honeycomb's OpenTelemetry-native observability platform — high-cardinality event store ideal for tracing LLM pipelines and debugging slow agent… | 16 stars · 6 forks | No one-command install · Source | |
| 22 | 74.422 | baserunobservability · observability | Baserun captures LLM traces via a lightweight decorator-based SDK and provides a dashboard for debugging prompt chains, testing variants, and… | 16 stars · 5 forks | Stale | No one-command install · Source |
| 23 | 71.115 | tencent-cloud-log-service-cls-mcp-servermcp · observability | Enables large language models to directly access Tencent Cloud Log Service for log search, metric queries, and alarm management without code. | 11 stars · 6 forks | No one-command install · Source | |
| 23 | 71.115 | background-process-mcpmcp · observabilityalso listed as waylaidwanderer-background-process | Enables LLMs to start, stop, and monitor long-running command-line processes in the background. | 11 stars · 6 forks | No one-command install · Source | |
| 26 | 70.060 | galileomcp · observability | Integrates with Galileo's evaluation and observability platform to enable dataset creation, prompt template management, experiment setup, log… | 6 stars · 7 forks | No one-command install · Source | |
| 27 | 59.040 | spanlens-mcpmcp · observability | MCP-native LLM observability. Query your Spanlens traces, stats, cost anomalies, and savings from Cursor, Claude Desktop, or any MCP client. Open… | 13 stars | No one-command install · Source | |
| 28 | 59.029 | aplaceforallmystuff-piholemcp · observability | Manage DNS blocking, monitor traffic stats, and control whitelist/blacklist for Pi-hole v6 | 8 stars · 1 fork | No one-command install · Source | |
| 29 | 58.969 | tilt-mcp-servermcp · observability | Enables LLMs and AI assistants to interact with Tilt development environments, providing tools to list resources, fetch logs, and monitor status. | 4 stars · 4 forks | No one-command install · Source | |
| 29 | 58.969 | ansible-aapmcp · observability | Enables LLMs to discover and launch Ansible job templates on Ansible Automation Platform, monitor job status, and retrieve outputs. | 4 stars · 4 forks | No one-command install · Source | |
| 31 | 57.245 | aplaceforallmystuff-tailscalemcp · observability | Provides read-only access to Tailscale network management for monitoring device status, tracking client updates, and generating network statistics… | 7 stars · 1 fork | No one-command install · Source | |
| 32 | 55.515 | rhoai-observability-mcpmcp · observability | Provides AI assistants with direct access to Red Hat OpenShift AI observability data, enabling querying of Prometheus metrics, Alertmanager alerts… | 5 stars · 2 forks | No one-command install · Source | |
| 33 | 47.273 | log-logn-langfuse-mcp-javamcp · observability | Query Langfuse traces, debug exceptions, analyze sessions, scores, datasets, schema, observations and manage prompts. Full observability toolkit for… | 3 stars · 2 forks | No one-command install · Source | |
| 34 | 43.275 | jungle-grid-mcp-servermcp · observability | MCP server for Jungle Grid, an agentic GPU execution layer that lets AI agents estimate, submit, monitor, and fetch logs for inference, training… | 4 stars | No one-command install · Source | |
| 35 | 42.840 | scorecardmcp · observability | Evaluate and optimize LLM systems with comprehensive testing and metrics | 3 forks | No one-command install · Source | |
| 36 | 32.223 | berserk-mcpmcp · observability | Enables LLMs to answer Berserk observability questions by calling verified KQL tools instead of hand-authoring queries, with role-based tool filtering… | 2 stars | No one-command install · Source | |
| 36 | 32.223 | odigo-elastic-s2l-mcpmcp · observability | Connects LLMs to Elasticsearch with a Semantic-to-Lexical layer that translates technical field names into business knowledge, enabling autonomous… | 2 stars | No one-command install · Source | |
| 38 | 22.557 | logicmcp-servermcp · observability | Enables traceable requirement discovery, technical alignment, and ISO-aligned process checking through deterministic MCP tools and resources, without… | 1 star | No one-command install · Source | |
| 38 | 22.557 | ebpf-mcp-tracermcp · observability | Enables LLMs to safely write and run bpftrace scripts against the Linux kernel for observability, with explicit probe allowlists and execution… | 1 star | No one-command install · Source | |
| 38 | 22.557 | llm-brand-monitormcp · observability | Track how 350+ AI models mention your brand — monitor visibility, sentiment, and competitor mentions. | 1 star | No one-command install · Source | |
| 38 | 22.557 | mcp-delonghi-ecammcp · observability | Enables LLMs to control DeLonghi ECAM espresso machines over a local network, including brewing beverages, monitoring status, and managing machine… | 1 star | No one-command install · Source | |
| 38 | 22.557 | mi25-tuning-mcpmcp · observability | Manages multi_llm-client operations for MI25/gfx900 GPUs, enabling configuration updates, inference execution, benchmarking, and performance log… | 1 star | No one-command install · Source | |
| 38 | 22.557 | mcp-log-analyzermcp · observability | Analyzes log files locally using Ollama and files structured GitHub Issues automatically, with all processing kept on your machine. | 1 star | No one-command install · Source | |
| 38 | 22.557 | oalles-agentic-system-monitoring-ragmcp · observability | Spring Boot-based server that connects system monitoring tools with a RAG service, enabling real-time access to both system metrics and corporate… | 1 star | No one-command install · Source | |
| 45 | Unranked | dynatrace-saas-mcp-servermcp · observability | Enables LLM agents to query Dynatrace SaaS for observability data (logs, metrics, traces, entities, problems, vulnerabilities) and manage… | No signals yet | No one-command install · Source | |
| 46 | Unranked | home-network-mcpmcp · observability | Lets an LLM client monitor devices, services, disk usage, and uptime in a home network and home lab. | No signals yet | No one-command install · Source | |
| 47 | Unranked | ml-lab-mcpmcp · observability | Enables LLMs to manage and run machine learning training jobs on a remote server, including syncing code, submitting experiments, monitoring progress… | No signals yet | No one-command install · Source | |
| 48 | Unranked | inference-aiopsmcp · observability | Enables governance-grade AIops for GPU inference clusters with root-cause analysis, metrics, and policy-governed operations for vLLM and Ray. | No signals yet | No one-command install · Source | |
| 49 | Unranked | workloadtruth-mcp-servermcp · observability | Enables classification of GPU workloads as training, inference, or idle from telemetry data, with tools for one-shot classification, benchmarking, and… | No signals yet | No one-command install · Source | |
| 50 | Unranked | local-worker-mcpmcp · observability | Delegates heavy, repetitive, and verifiable tasks like PDF extraction, code analysis, and log processing to a local LLM to reduce token consumption… | No signals yet | No one-command install · Source | |
| 51 | Unranked | unitree-go2-mcp-servermcp · observability | Enables LLM agents to control and monitor a Unitree Go2 robot through MCP tools, including live telemetry, navigation, camera feeds, and waypoint… | No signals yet | No one-command install · Source | |
| 52 | Unranked | mcp-devops-dashboardmcp · observability | Monitors host CPU/RAM, local ports, and Docker containers, streaming live telemetry to a React dashboard over SSE. Exposes MCP tools to query system… | No signals yet | No one-command install · Source | |
| 53 | Unranked | io-mcpmcp · observability | Enables local research workflows (paper discovery, relevance scoring, digests) and homelab monitoring (Prometheus, logs) through an MCP server, using… | No signals yet | No one-command install · Source | |
| 54 | Unranked | apple-health-semantic-mcpmcp · observability | Query Apple Health export data with an LLM via a semantic layer that avoids common data traps, enabling accurate natural-language queries about health… | No signals yet | No one-command install · Source | |
| 55 | Unranked | codex-worker-runtimemcp · observability | Enables Codex to delegate bounded work to external LLMs through role-based MCP tools, with worker health checks and audit logging. | No signals yet | No one-command install · Source | |
| 56 | Unranked | ros-llm-integration-bridgemcp · observability | Enables large language models to interact with ROS robots seamlessly, allowing natural language control, real-time sensor monitoring, and autonomous… | No signals yet | No one-command install · Source | |
| 57 | Unranked | mcp-hayabusamcp · observability | Enables an LLM client to scan Windows event log files (EVTX) for suspicious activity using Hayabusa, and browse its detection rules directly in… | No signals yet | No one-command install · Source | |
| 58 | Unranked | cloudwatch-mcp-agentmcp · observability | Enables natural language queries for AWS CloudWatch logs, metrics, and alarms via an LLM agent with MCP tools. | No signals yet | No one-command install · Source | |
| 59 | Unranked | emotion-mcpmcp · observability | Provides a dynamic emotion simulation system for AI character roleplay, based on Freudian psychodynamics, that analyzes dialogue content via LLM to… | No signals yet | No one-command install · Source | |
| 60 | Unranked | es-error-lensmcp · observability | Enables LLM agents to search and analyze Elasticsearch logs for errors, detect recurring patterns, analyze error-rate trends, and retrieve full trace… | No signals yet | No one-command install · Source | |
| 61 | Unranked | tracepii-shield-mcpmcp · observability | Redacts PII from LLM traces and tool payloads before they leave review, enabling PII scanning, payload redaction, sensitive field classification… | No signals yet | No one-command install · Source | |
| 62 | Unranked | mdl-train-mcpmcp · observability | An MCP server for monitoring and managing training jobs on Modal. Built for LLMs that need to check on long-running GPU training without drowning in… | No signals yet | No one-command install · Source | |
| 63 | Unranked | session-logger-mcpmcp · observability | Enables saving LLM chat conversations to JSONL log files and querying them by session ID, user ID, keyword, or date range. | No signals yet | No one-command install · Source | |
| 64 | Unranked | agent-evaluationskill · observability | Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top… | No signals yet | No commit datelisted | armory install agent-evaluation --cli claude |
| 65 | Unranked | agentlens-mcp-servermcp · observability | Enables AI agents to access observability and evaluation data, including run history, span traces, LLM-as-judge evaluation results, and regression… | No signals yet | No commit datelisted | No one-command install · Source |
| 66 | Unranked | datadog-llm-observabilityobservability · observability | Datadog's managed LLM Observability product — traces LLM calls, monitors prompt/completion quality, detects anomalies, and integrates with existing… | No signals yet | No commit datelisted | No one-command install · Source |
| 67 | Unranked | fast-award-screener-mcpmcp · observability | Screens DIBBS RFQs for micro-purchase viability using deterministic logic without LLM calls. | No signals yet | No commit datelisted | No one-command install · Source |
| 68 | Unranked | fiddler-aiobservability · observability | Fiddler AI Observability platform monitors LLM applications for hallucinations, toxicity, bias, and drift, with explainability and alerting for… | No signals yet | No commit datelisted | No one-command install · Source |
| 69 | Unranked | honeyhiveeval · observability | LLM evaluation and experimentation platform with session tracing, dataset management, and metric-based run comparison. | No signals yet | No commit datelisted | No one-command install · Source |
| 70 | Unranked | langsmith-observabilityskill · observability | LLM observability platform for tracing, evaluation, and monitoring. Use when debugging LLM applications, evaluating model outputs against datasets… | No signals yet | No commit datelisted | armory install langsmith-observability --cli claude |
| 71 | Unranked | langsmith-tracing-3rules · observability | Configure LangSmith tracing environment variables for Claude Code observability. Sends conversation traces to LangSmith for monitoring and analysis… | No signals yet | No commit datelisted | armory install langsmith-tracing-3 --cli claude |
| 72 | Unranked | langtraceobservability · observability | Open-source, OpenTelemetry-compliant LLM observability tool by Scale3Labs — traces calls to all major LLM providers and frameworks with a… | No signals yet | No commit datelisted | No one-command install · Source |
| 73 | Unranked | llm-evaluationskill · observability | Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing. | No signals yet | No commit datelisted | armory install llm-evaluation --cli claude |
| 74 | Unranked | lunaryobservability · observability | Open-source LLM observability and prompt management platform — tracks conversations, errors, costs, and user feedback for production AI applications. | No signals yet | No commit datelisted | No one-command install · Source |
| 75 | Unranked | ml-engineer-2skill · observability | Build production ML systems with PyTorch 2.x, TensorFlow, and modern ML frameworks. Implements model serving, feature engineering, A/B testing, and… | No signals yet | No commit datelisted | armory install ml-engineer-2 --cli claude |
| 76 | Unranked | mle-reviewersubagent · observability | Production machine-learning engineering reviewer for data contracts, feature pipelines, training reproducibility, offline/online evaluation, model… | No signals yet | No commit datelisted | armory install mle-reviewer --cli claude |
| 77 | Unranked | new-relic-ai-monitoringobservability · observability | New Relic AI Monitoring instruments LLM calls end-to-end — traces model invocations, measures token costs, and surfaces anomalies via the New Relic… | No signals yet | No commit datelisted | No one-command install · Source |
| 78 | Unranked | phoenix-observabilityskill · observability | Open-source AI observability platform for LLM tracing, evaluation, and monitoring. Use when debugging LLM applications with detailed traces, running… | No signals yet | No commit datelisted | armory install phoenix-observability --cli claude |
| 79 | Unranked | statsmodelsskill · observability | Statistical modeling toolkit. OLS, GLM, logistic, ARIMA, time series, hypothesis tests, diagnostics, AIC/BIC, for rigorous statistical inference and… | No signals yet | No commit datelisted | armory install statsmodels --cli claude |
Score colour shows how many signals stand behind it, never how good it is: amber, three or more; dimmer amber, two; grey, one. The Evidence column names them.
Stars, forks and last commit are as GitHub reported them when Armory last read each repository: for most, or later. A repository may have changed since.