Evals
Graded tasks that prove a change helped
- Listed
- 37
- Ranked
- 86.5%
- Top Score
- 99.5
Pick
The pick and its runners-up, in score order
- promptfoo99.52 signals
The pick · #1 on this shelf
Assertions and red-teaming for prompts and agents, runs in CI
24,737 stars · 2,255 forksNo one-command install · Source - deepeval99.42 signals
Runner-up · #4 on this shelf
Evaluation framework with 14+ metrics such as hallucination and faithfulness, runs in CI
OpenAI Evals and the LM Evaluation Harness, above it, mainly grade models; this grades your own app.
18,041 stars · 1,890 forksNo one-command install · Source - swe-bench99.13 signals
Runner-up · #7 on this shelf
Real GitHub issues drawn from 12 Python repositories, scored on the resulting patch
Listed as the benchmark coding agents are compared on; the rows above it test prompts, apps and models instead.
5,762 stars · 957 forks · 9 mentionsNo one-command install · Source
Top Ranked
Needs the armory CLI · not on npm yet, build it from cli/ in the repository
| Rank | Score | Component | Description | Evidence | Last commit | Install |
|---|---|---|---|---|---|---|
| 1 | 99.536 | promptfooeval · ai-agents | CLI and library for testing, comparing, and red-teaming LLM prompts and agents with assertions and CI integration. | 24,737 stars · 2,255 forks | No one-command install · Source | |
| 2 | 99.512 | openai-evalseval · ai-agents | OpenAI's official framework for evaluating LLMs and LLM-powered systems, with a registry of community eval sets. | 19,509 stars · 3,093 forks | No one-command install · Source | |
| 3 | 99.505 | lm-evaluation-harnesseval · other | EleutherAI's unified framework for evaluating language models on hundreds of academic benchmarks. | 13,860 stars · 3,533 forks | No one-command install · Source | |
| 4 | 99.455 | deepevaleval · observability | Open-source LLM evaluation framework with 14+ metrics (hallucination, faithfulness, answer relevancy) and CI support. | 18,041 stars · 1,890 forks | No one-command install · Source | |
| 5 | 99.421 | ragaseval · search | Reference-free evaluation of retrieval-augmented generation pipelines; measures faithfulness, answer relevance, and context precision. | 15,853 stars · 1,727 forks · 1 mention · failed install test | No one-command install · Source | |
| 6 | 99.273 | phoenixeval · ai-agents | Arize Phoenix: open-source LLM observability with built-in evals, span tracing, and dataset curation for RAG and agents. | 11,286 stars · 1,086 forks · 2 mentions | No one-command install · Source | |
| 7 | 99.103 | swe-bencheval · ai-agents | SWE-bench: benchmark for evaluating LLMs on real-world GitHub issue resolution across 12 popular Python repositories. | 5,762 stars · 957 forks · 9 mentions | No one-command install · Source | |
| 8 | 99.052 | giskardeval · ai-agents | Open-source LLM testing framework for detecting vulnerabilities (prompt injection, hallucinations, bias) via automated scan. | 5,838 stars · 542 forks | No one-command install · Source | |
| 9 | 98.991 | agentaeval · ai-agents | Open-source LLM developer platform with prompt playground, evaluation pipelines, and A/B testing for iterating on LLM apps. | 4,670 stars · 661 forks | No one-command install · Source | |
| 10 | 98.956 | openai-simple-evalseval · ai-agents | OpenAI's lightweight benchmark suite (MMLU, HumanEval, MATH, GPQA, MGSM) for fast model capability comparisons. | 4,621 stars · 509 forks | No one-command install · Source | |
| 11 | 98.825 | big-bencheval · front-end | Google's Beyond the Imitation Game benchmark: 200+ diverse tasks designed to probe capabilities beyond standard NLP benchmarks. | 3,247 stars · 618 forks | Stale | No one-command install · Source |
| 12 | 98.783 | trulenseval · back-end | Evaluation and tracking for LLM and RAG applications with a feedback-function API and experiment dashboard. | 3,530 stars · 335 forks · 1 mention | No one-command install · Source | |
| 13 | 98.776 | xlang-ai-osworldeval · ai-agentsContributed by Sentinel | [NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments | 3,117 stars · 530 forks · 7 mentions | No one-command install · Source | |
| 14 | 98.743 | inspect-aieval · other | UK AISI's framework for safety evaluations of large language models, with task-based scaffolding and solver pipelines. | 2,683 stars · 685 forks | No one-command install · Source | |
| 15 | 98.706 | helmeval · other | Stanford CRFM Holistic Evaluation of Language Models: standardized benchmark suite covering accuracy, calibration, robustness, and fairness. | 2,898 stars · 412 forks | No one-command install · Source | |
| 16 | 98.688 | lightevaleval · ai-agents | Hugging Face lightweight evaluation library for LLMs across academic benchmarks, with fast local and remote inference support. | 2,533 stars · 553 forks | No one-command install · Source | |
| 17 | 98.324 | evalpluseval · other | Rigorous code generation evaluation framework built on top of HumanEval and MBPP with 80x more test cases. | 1,819 stars · 208 forks | No one-command install · Source | |
| 18 | 98.283 | webarenaeval · browser | WebArena: realistic web-based environment for evaluating autonomous agents on long-horizon browser interaction tasks. | 1,592 stars · 249 forks · 1 mention | No one-command install · Source | |
| 19 | 98.177 | tau-bencheval · ai-agents | Tau-bench: agent benchmark for tool-agent-user interactions in retail and airline domains with policy-grounded evaluation. | 1,416 stars · 215 forks · 1 mention | No one-command install · Source | |
| 20 | 98.141 | harbor-framework-terminal-bencheval · ai-agentsContributed by Sentinel | Measuring and evolving with the frontier of agent work | 588 stars · 434 forks · 13 mentions | No one-command install · Source |
Score colour shows how many signals stand behind it, never how good it is: amber, three or more; dimmer amber, two; grey, one. The Evidence column names them.
Stars, forks and last commit are as GitHub reported them when Armory last read each repository: for most, or later. A repository may have changed since.
Leaderboard ranks all 32 scored rows. Formula shows the calculation.