Armory
Source

Leaderboard

Scored on public signals; components with none are listed as Unranked · Formula

RankScoreComponentDescriptionEvidenceLast commitInstall
199.505lm-evaluation-harnesseval · otherEleutherAI's unified framework for evaluating language models on hundreds of academic benchmarks.13,860 stars · 3,533 forksNo one-command install · Source
298.825big-bencheval · front-endGoogle's Beyond the Imitation Game benchmark: 200+ diverse tasks designed to probe capabilities beyond standard NLP benchmarks.3,247 stars · 618 forksStaleNo one-command install · Source
398.706helmeval · otherStanford CRFM Holistic Evaluation of Language Models: standardized benchmark suite covering accuracy, calibration, robustness, and fairness.2,898 stars · 412 forksNo one-command install · Source

Score colour shows how many signals stand behind it, never how good it is: amber, three or more; dimmer amber, two; grey, one. The Evidence column names them.

Stars, forks and last commit are as GitHub reported them when Armory last read each repository: for most, or later. A repository may have changed since.

Leaderboard · Armory