Armory
Source

Leaderboard

Scored on public signals; components with none are listed as Unranked · Formula

RankScoreComponentDescriptionEvidenceLast commitInstall
199.455deepevaleval · observabilityOpen-source LLM evaluation framework with 14+ metrics (hallucination, faithfulness, answer relevancy) and CI support.18,041 stars · 1,890 forksNo one-command install · Source
297.880langtraceeval · observabilityOpen-source observability tool for LLMs with OpenTelemetry-based tracing, automated evals, and annotation workflows.1,228 stars · 127 forksNo one-command install · Source
390.554vellum-evalseval · observabilityVellum evaluation SDK for running LLM test suites with custom metrics, dataset pinning, and CI workflow integration.82 stars · 20 forksNo one-command install · Source
482.319braintrusteval · observabilityDeveloper platform for logging, evaluating, and comparing LLM experiments with dataset versioning and scoring functions.27 stars · 12 forksNo one-command install · Source
580.342galileo-evaluateeval · observabilityGalileo evaluation and observability SDK for detecting hallucinations, data errors, and model weaknesses in LLM pipelines.22 stars · 11 forksNo one-command install · Source
6UnrankedagentbenchOurseval · observabilityUse to put a number on harness quality: run an agent harness against a task set and get a score, so harness changes are validated by evidence. It is…No signals yetNo one-command install · Source
7Unrankedhoneyhiveeval · observabilityLLM evaluation and experimentation platform with session tracing, dataset management, and metric-based run comparison.No signals yetNo commit datelisted No one-command install · Source

Score colour shows how many signals stand behind it, never how good it is: amber, three or more; dimmer amber, two; grey, one. The Evidence column names them.

Stars, forks and last commit are as GitHub reported them when Armory last read each repository: for most, or later. A repository may have changed since.

Leaderboard · Armory