Leaderboard
Scored on public signals; components with none are listed as Unranked · Formula
Component
Domain
Vertical
| Rank | Score | Component | Description | Evidence | Last commit | Install |
|---|---|---|---|---|---|---|
| 1 | 99.455 | deepevaleval · observability | Open-source LLM evaluation framework with 14+ metrics (hallucination, faithfulness, answer relevancy) and CI support. | 18,041 stars · 1,890 forks | No one-command install · Source | |
| 2 | 97.880 | langtraceeval · observability | Open-source observability tool for LLMs with OpenTelemetry-based tracing, automated evals, and annotation workflows. | 1,228 stars · 127 forks | No one-command install · Source | |
| 3 | 90.554 | vellum-evalseval · observability | Vellum evaluation SDK for running LLM test suites with custom metrics, dataset pinning, and CI workflow integration. | 82 stars · 20 forks | No one-command install · Source | |
| 4 | 82.319 | braintrusteval · observability | Developer platform for logging, evaluating, and comparing LLM experiments with dataset versioning and scoring functions. | 27 stars · 12 forks | No one-command install · Source | |
| 5 | 80.342 | galileo-evaluateeval · observability | Galileo evaluation and observability SDK for detecting hallucinations, data errors, and model weaknesses in LLM pipelines. | 22 stars · 11 forks | No one-command install · Source | |
| 6 | Unranked | agentbenchOurseval · observability | Use to put a number on harness quality: run an agent harness against a task set and get a score, so harness changes are validated by evidence. It is… | No signals yet | No one-command install · Source | |
| 7 | Unranked | honeyhiveeval · observability | LLM evaluation and experimentation platform with session tracing, dataset management, and metric-based run comparison. | No signals yet | No commit datelisted | No one-command install · Source |
Score colour shows how many signals stand behind it, never how good it is: amber, three or more; dimmer amber, two; grey, one. The Evidence column names them.
Stars, forks and last commit are as GitHub reported them when Armory last read each repository: for most, or later. A repository may have changed since.