Leaderboard
Scored on public signals; components with none are listed as Unranked · Formula
Component
Domain
Vertical
| Rank | Score | Component | Description | Evidence | Last commit | Install |
|---|---|---|---|---|---|---|
| 1 | 94.309 | harbor-framework-terminal-bench-2-1eval · otherContributed by Sentinel | Terminal-Bench 2.1 | 119 stars · 62 forks · 3 mentions | No one-command install · Source | |
| 2 | 98.324 | evalpluseval · other | Rigorous code generation evaluation framework built on top of HumanEval and MBPP with 80x more test cases. | 1,819 stars · 208 forks | No one-command install · Source | |
| 3 | 98.706 | helmeval · other | Stanford CRFM Holistic Evaluation of Language Models: standardized benchmark suite covering accuracy, calibration, robustness, and fairness. | 2,898 stars · 412 forks | No one-command install · Source | |
| 4 | 98.743 | inspect-aieval · other | UK AISI's framework for safety evaluations of large language models, with task-based scaffolding and solver pipelines. | 2,683 stars · 685 forks | No one-command install · Source | |
| 5 | 99.505 | lm-evaluation-harnesseval · other | EleutherAI's unified framework for evaluating language models on hundreds of academic benchmarks. | 13,860 stars · 3,533 forks | No one-command install · Source |
Score colour shows how many signals stand behind it, never how good it is: amber, three or more; dimmer amber, two; grey, one. The Evidence column names them.
Stars, forks and last commit are as GitHub reported them when Armory last read each repository: for most, or later. A repository may have changed since.