Armory
Source

Evals

Graded tasks that prove a change helped

Listed
37
Ranked
86.5%
Top Score
99.5

Pick

The pick and its runners-up, in score order

  • promptfoo99.52 signals

    The pick · #1 on this shelf

    Assertions and red-teaming for prompts and agents, runs in CI

    24,737 stars · 2,255 forks
    No one-command install · Source
  • deepeval99.42 signals

    Runner-up · #4 on this shelf

    Evaluation framework with 14+ metrics such as hallucination and faithfulness, runs in CI

    OpenAI Evals and the LM Evaluation Harness, above it, mainly grade models; this grades your own app.

    18,041 stars · 1,890 forks
    No one-command install · Source
  • swe-bench99.13 signals

    Runner-up · #7 on this shelf

    Real GitHub issues drawn from 12 Python repositories, scored on the resulting patch

    Listed as the benchmark coding agents are compared on; the rows above it test prompts, apps and models instead.

    5,762 stars · 957 forks · 9 mentions
    No one-command install · Source

Top Ranked

Needs the armory CLI · not on npm yet, build it from cli/ in the repository

RankScoreComponentDescriptionEvidenceLast commitInstall
199.536promptfooeval · ai-agentsCLI and library for testing, comparing, and red-teaming LLM prompts and agents with assertions and CI integration.24,737 stars · 2,255 forksNo one-command install · Source
299.512openai-evalseval · ai-agentsOpenAI's official framework for evaluating LLMs and LLM-powered systems, with a registry of community eval sets.19,509 stars · 3,093 forksNo one-command install · Source
399.505lm-evaluation-harnesseval · otherEleutherAI's unified framework for evaluating language models on hundreds of academic benchmarks.13,860 stars · 3,533 forksNo one-command install · Source
499.455deepevaleval · observabilityOpen-source LLM evaluation framework with 14+ metrics (hallucination, faithfulness, answer relevancy) and CI support.18,041 stars · 1,890 forksNo one-command install · Source
599.421ragaseval · searchReference-free evaluation of retrieval-augmented generation pipelines; measures faithfulness, answer relevance, and context precision.15,853 stars · 1,727 forks · 1 mention · failed install testNo one-command install · Source
699.273phoenixeval · ai-agentsArize Phoenix: open-source LLM observability with built-in evals, span tracing, and dataset curation for RAG and agents.11,286 stars · 1,086 forks · 2 mentionsNo one-command install · Source
799.103swe-bencheval · ai-agentsSWE-bench: benchmark for evaluating LLMs on real-world GitHub issue resolution across 12 popular Python repositories.5,762 stars · 957 forks · 9 mentionsNo one-command install · Source
899.052giskardeval · ai-agentsOpen-source LLM testing framework for detecting vulnerabilities (prompt injection, hallucinations, bias) via automated scan.5,838 stars · 542 forksNo one-command install · Source
998.991agentaeval · ai-agentsOpen-source LLM developer platform with prompt playground, evaluation pipelines, and A/B testing for iterating on LLM apps.4,670 stars · 661 forksNo one-command install · Source
1098.956openai-simple-evalseval · ai-agentsOpenAI's lightweight benchmark suite (MMLU, HumanEval, MATH, GPQA, MGSM) for fast model capability comparisons.4,621 stars · 509 forksNo one-command install · Source
1198.825big-bencheval · front-endGoogle's Beyond the Imitation Game benchmark: 200+ diverse tasks designed to probe capabilities beyond standard NLP benchmarks.3,247 stars · 618 forksStaleNo one-command install · Source
1298.783trulenseval · back-endEvaluation and tracking for LLM and RAG applications with a feedback-function API and experiment dashboard.3,530 stars · 335 forks · 1 mentionNo one-command install · Source
1398.776xlang-ai-osworldeval · ai-agentsContributed by Sentinel[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments3,117 stars · 530 forks · 7 mentionsNo one-command install · Source
1498.743inspect-aieval · otherUK AISI's framework for safety evaluations of large language models, with task-based scaffolding and solver pipelines.2,683 stars · 685 forksNo one-command install · Source
1598.706helmeval · otherStanford CRFM Holistic Evaluation of Language Models: standardized benchmark suite covering accuracy, calibration, robustness, and fairness.2,898 stars · 412 forksNo one-command install · Source
1698.688lightevaleval · ai-agentsHugging Face lightweight evaluation library for LLMs across academic benchmarks, with fast local and remote inference support.2,533 stars · 553 forksNo one-command install · Source
1798.324evalpluseval · otherRigorous code generation evaluation framework built on top of HumanEval and MBPP with 80x more test cases.1,819 stars · 208 forksNo one-command install · Source
1898.283webarenaeval · browserWebArena: realistic web-based environment for evaluating autonomous agents on long-horizon browser interaction tasks.1,592 stars · 249 forks · 1 mentionNo one-command install · Source
1998.177tau-bencheval · ai-agentsTau-bench: agent benchmark for tool-agent-user interactions in retail and airline domains with policy-grounded evaluation.1,416 stars · 215 forks · 1 mentionNo one-command install · Source
2098.141harbor-framework-terminal-bencheval · ai-agentsContributed by SentinelMeasuring and evolving with the frontier of agent work588 stars · 434 forks · 13 mentionsNo one-command install · Source

Score colour shows how many signals stand behind it, never how good it is: amber, three or more; dimmer amber, two; grey, one. The Evidence column names them.

Stars, forks and last commit are as GitHub reported them when Armory last read each repository: for most, or later. A repository may have changed since.

Leaderboard ranks all 32 scored rows. Formula shows the calculation.