Armory
Source

Leaderboard

Scored on public signals; components with none are listed as Unranked · Formula

RankScoreComponentDescriptionEvidenceLast commitInstall
169.000humanloop-evalseval · ai-agentsHumanloop Python SDK with integrated evals, dataset versioning, and human + LLM judge scoring for production pipelines.12 stars · 3 forksNo one-command install · Source
280.342galileo-evaluateeval · observabilityGalileo evaluation and observability SDK for detecting hallucinations, data errors, and model weaknesses in LLM pipelines.22 stars · 11 forksNo one-command install · Source
382.319braintrusteval · observabilityDeveloper platform for logging, evaluating, and comparing LLM experiments with dataset versioning and scoring functions.27 stars · 12 forksNo one-command install · Source
490.554vellum-evalseval · observabilityVellum evaluation SDK for running LLM test suites with custom metrics, dataset pinning, and CI workflow integration.82 stars · 20 forksNo one-command install · Source
596.207continuous-evaleval · ai-agentsRelari's modular evaluation library for LLM pipelines with deterministic + LLM-based metrics for RAG and agent workflows.517 stars · 38 forksNo one-command install · Source
697.880langtraceeval · observabilityOpen-source observability tool for LLMs with OpenTelemetry-based tracing, automated evals, and annotation workflows.1,228 stars · 127 forksNo one-command install · Source
797.923wandb-weave-evalseval · ai-agentsWeights & Biases Weave evaluation framework for tracking LLM experiments, scoring model outputs, and comparing runs.1,130 stars · 170 forks · 3 mentionsNo one-command install · Source
898.688lightevaleval · ai-agentsHugging Face lightweight evaluation library for LLMs across academic benchmarks, with fast local and remote inference support.2,533 stars · 553 forksNo one-command install · Source
998.783trulenseval · back-endEvaluation and tracking for LLM and RAG applications with a feedback-function API and experiment dashboard.3,530 stars · 335 forks · 1 mentionNo one-command install · Source
1098.991agentaeval · ai-agentsOpen-source LLM developer platform with prompt playground, evaluation pipelines, and A/B testing for iterating on LLM apps.4,670 stars · 661 forksNo one-command install · Source
1199.103swe-bencheval · ai-agentsSWE-bench: benchmark for evaluating LLMs on real-world GitHub issue resolution across 12 popular Python repositories.5,762 stars · 957 forks · 9 mentionsNo one-command install · Source
1299.273phoenixeval · ai-agentsArize Phoenix: open-source LLM observability with built-in evals, span tracing, and dataset curation for RAG and agents.11,286 stars · 1,086 forks · 2 mentionsNo one-command install · Source
1399.455deepevaleval · observabilityOpen-source LLM evaluation framework with 14+ metrics (hallucination, faithfulness, answer relevancy) and CI support.18,041 stars · 1,890 forksNo one-command install · Source
1499.512openai-evalseval · ai-agentsOpenAI's official framework for evaluating LLMs and LLM-powered systems, with a registry of community eval sets.19,509 stars · 3,093 forksNo one-command install · Source
1599.536promptfooeval · ai-agentsCLI and library for testing, comparing, and red-teaming LLM prompts and agents with assertions and CI integration.24,737 stars · 2,255 forksNo one-command install · Source
16Unrankedhoneyhiveeval · observabilityLLM evaluation and experimentation platform with session tracing, dataset management, and metric-based run comparison.No signals yetNo commit datelisted No one-command install · Source
17Unrankedpatronus-aieval · ai-agentsAutomated LLM evaluation and hallucination detection platform with a Python SDK and judge-model scoring.No signals yetNo commit datelisted No one-command install · Source

Score colour shows how many signals stand behind it, never how good it is: amber, three or more; dimmer amber, two; grey, one. The Evidence column names them.

Stars, forks and last commit are as GitHub reported them when Armory last read each repository: for most, or later. A repository may have changed since.

Leaderboard · Armory