Armory
Source
Browse
Evals

helm

Stanford CRFM Holistic Evaluation of Language Models: standardized benchmark suite covering accuracy, calibration, robustness, and fairness.

Score
98.7062 signals
Evidence
2,898 stars · 412 forks
Last commit
as last read from GitHub; most reads are from 2 Sep 2026 or later
Listed

Install

No one-command install. Set it up from its source.

Alternatives · Evals

  1. lm-evaluation-harness13,860 stars · 3,533 forks99.505

What it is

Stanford CRFM Holistic Evaluation of Language Models: standardized benchmark suite covering accuracy, calibration, robustness, and fairness.

When to use it

Stanford CRFM Holistic Evaluation of Language Models: standardized benchmark suite covering accuracy, calibration, robustness, and fairness.

Notes

Curated evals entry. Verified 2026-05-27.