Leaderboard
Scored on public signals; components with none are listed as Unranked · Formula
Component
Domain
Vertical
| Rank | Score | Component | Description | Evidence | Last commit | Install |
|---|---|---|---|---|---|---|
| 1 | 99.505 | lm-evaluation-harnesseval · other | EleutherAI's unified framework for evaluating language models on hundreds of academic benchmarks. | 13,860 stars · 3,533 forks | No one-command install · Source | |
| 2 | 98.825 | big-bencheval · front-end | Google's Beyond the Imitation Game benchmark: 200+ diverse tasks designed to probe capabilities beyond standard NLP benchmarks. | 3,247 stars · 618 forks | Stale | No one-command install · Source |
| 3 | 98.706 | helmeval · other | Stanford CRFM Holistic Evaluation of Language Models: standardized benchmark suite covering accuracy, calibration, robustness, and fairness. | 2,898 stars · 412 forks | No one-command install · Source |
Score colour shows how many signals stand behind it, never how good it is: amber, three or more; dimmer amber, two; grey, one. The Evidence column names them.
Stars, forks and last commit are as GitHub reported them when Armory last read each repository: for most, or later. A repository may have changed since.