Armory
Source

Leaderboard

Scored on public signals; components with none are listed as Unranked · Formula

RankScoreComponentDescriptionEvidenceLast commitInstall
199.505lm-evaluation-harnesseval · otherEleutherAI's unified framework for evaluating language models on hundreds of academic benchmarks.13,860 stars · 3,533 forksNo one-command install · Source
298.743inspect-aieval · otherUK AISI's framework for safety evaluations of large language models, with task-based scaffolding and solver pipelines.2,683 stars · 685 forksNo one-command install · Source
398.706helmeval · otherStanford CRFM Holistic Evaluation of Language Models: standardized benchmark suite covering accuracy, calibration, robustness, and fairness.2,898 stars · 412 forksNo one-command install · Source
498.324evalpluseval · otherRigorous code generation evaluation framework built on top of HumanEval and MBPP with 80x more test cases.1,819 stars · 208 forksNo one-command install · Source
594.309harbor-framework-terminal-bench-2-1eval · otherContributed by SentinelTerminal-Bench 2.1119 stars · 62 forks · 3 mentionsNo one-command install · Source

Score colour shows how many signals stand behind it, never how good it is: amber, three or more; dimmer amber, two; grey, one. The Evidence column names them.

Stars, forks and last commit are as GitHub reported them when Armory last read each repository: for most, or later. A repository may have changed since.

Leaderboard · Armory