Armory
Source

Leaderboard

Scored on public signals; components with none are listed as Unranked · Formula

RankScoreComponentDescriptionEvidenceLast commitInstall
199.536promptfooeval · ai-agentsCLI and library for testing, comparing, and red-teaming LLM prompts and agents with assertions and CI integration.24,737 stars · 2,255 forksNo one-command install · Source
299.512openai-evalseval · ai-agentsOpenAI's official framework for evaluating LLMs and LLM-powered systems, with a registry of community eval sets.19,509 stars · 3,093 forksNo one-command install · Source
399.505lm-evaluation-harnesseval · otherEleutherAI's unified framework for evaluating language models on hundreds of academic benchmarks.13,860 stars · 3,533 forksNo one-command install · Source
499.455deepevaleval · observabilityOpen-source LLM evaluation framework with 14+ metrics (hallucination, faithfulness, answer relevancy) and CI support.18,041 stars · 1,890 forksNo one-command install · Source
599.421ragaseval · searchReference-free evaluation of retrieval-augmented generation pipelines; measures faithfulness, answer relevance, and context precision.15,853 stars · 1,727 forks · 1 mention · failed install testNo one-command install · Source
699.273phoenixeval · ai-agentsArize Phoenix: open-source LLM observability with built-in evals, span tracing, and dataset curation for RAG and agents.11,286 stars · 1,086 forks · 2 mentionsNo one-command install · Source
799.103swe-bencheval · ai-agentsSWE-bench: benchmark for evaluating LLMs on real-world GitHub issue resolution across 12 popular Python repositories.5,762 stars · 957 forks · 9 mentionsNo one-command install · Source
899.052giskardeval · ai-agentsOpen-source LLM testing framework for detecting vulnerabilities (prompt injection, hallucinations, bias) via automated scan.5,838 stars · 542 forksNo one-command install · Source
998.991agentaeval · ai-agentsOpen-source LLM developer platform with prompt playground, evaluation pipelines, and A/B testing for iterating on LLM apps.4,670 stars · 661 forksNo one-command install · Source
1098.956openai-simple-evalseval · ai-agentsOpenAI's lightweight benchmark suite (MMLU, HumanEval, MATH, GPQA, MGSM) for fast model capability comparisons.4,621 stars · 509 forksNo one-command install · Source
1198.825big-bencheval · front-endGoogle's Beyond the Imitation Game benchmark: 200+ diverse tasks designed to probe capabilities beyond standard NLP benchmarks.3,247 stars · 618 forksStaleNo one-command install · Source
1298.783trulenseval · back-endEvaluation and tracking for LLM and RAG applications with a feedback-function API and experiment dashboard.3,530 stars · 335 forks · 1 mentionNo one-command install · Source
1398.776xlang-ai-osworldeval · ai-agentsContributed by Sentinel[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments3,117 stars · 530 forks · 7 mentionsNo one-command install · Source
1498.743inspect-aieval · otherUK AISI's framework for safety evaluations of large language models, with task-based scaffolding and solver pipelines.2,683 stars · 685 forksNo one-command install · Source
1598.706helmeval · otherStanford CRFM Holistic Evaluation of Language Models: standardized benchmark suite covering accuracy, calibration, robustness, and fairness.2,898 stars · 412 forksNo one-command install · Source
1698.688lightevaleval · ai-agentsHugging Face lightweight evaluation library for LLMs across academic benchmarks, with fast local and remote inference support.2,533 stars · 553 forksNo one-command install · Source
1798.324evalpluseval · otherRigorous code generation evaluation framework built on top of HumanEval and MBPP with 80x more test cases.1,819 stars · 208 forksNo one-command install · Source
1898.283webarenaeval · browserWebArena: realistic web-based environment for evaluating autonomous agents on long-horizon browser interaction tasks.1,592 stars · 249 forks · 1 mentionNo one-command install · Source
1998.177tau-bencheval · ai-agentsTau-bench: agent benchmark for tool-agent-user interactions in retail and airline domains with policy-grounded evaluation.1,416 stars · 215 forks · 1 mentionNo one-command install · Source
2098.141harbor-framework-terminal-bencheval · ai-agentsContributed by SentinelMeasuring and evolving with the frontier of agent work588 stars · 434 forks · 13 mentionsNo one-command install · Source
2197.923wandb-weave-evalseval · ai-agentsWeights & Biases Weave evaluation framework for tracking LLM experiments, scoring model outputs, and comparing runs.1,130 stars · 170 forks · 3 mentionsNo one-command install · Source
2297.880langtraceeval · observabilityOpen-source observability tool for LLMs with OpenTelemetry-based tracing, automated evals, and annotation workflows.1,228 stars · 127 forksNo one-command install · Source
2397.541xiaowu0162-longmemevaleval · commsContributed by SentinelBenchmarking Chat Assistants on Long-Term Interactive Memory (ICLR 2025)1,049 stars · 81 forks · 10 mentionsNo one-command install · Source
2497.111princeton-nlp-webshopeval · ai-agentsContributed by Sentinel[NeurIPS 2022] 🛒WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents589 stars · 107 forks · 3 mentionsStaleNo one-command install · Source
2596.786stonybrooknlp-appworldeval · ai-agentsContributed by Sentinel🌍 AppWorld: A Controllable World of Apps and People for Benchmarking Function Calling and Interactive Coding Agent, ACL'24 Best Resource Paper.500 stars · 78 forks · 4 mentionsNo one-command install · Source
2696.207continuous-evaleval · ai-agentsRelari's modular evaluation library for LLM pipelines with deterministic + LLM-based metrics for RAG and agent workflows.517 stars · 38 forksNo one-command install · Source
2794.339metr-task-standardeval · ai-agentsMETR's Task Standard: a specification and scaffold for creating agentic tasks used in autonomous agent capability evaluations.192 stars · 37 forksNo one-command install · Source
2894.309harbor-framework-terminal-bench-2-1eval · otherContributed by SentinelTerminal-Bench 2.1119 stars · 62 forks · 3 mentionsNo one-command install · Source
2990.554vellum-evalseval · observabilityVellum evaluation SDK for running LLM test suites with custom metrics, dataset pinning, and CI workflow integration.82 stars · 20 forksNo one-command install · Source
3082.319braintrusteval · observabilityDeveloper platform for logging, evaluating, and comparing LLM experiments with dataset versioning and scoring functions.27 stars · 12 forksNo one-command install · Source
3180.342galileo-evaluateeval · observabilityGalileo evaluation and observability SDK for detecting hallucinations, data errors, and model weaknesses in LLM pipelines.22 stars · 11 forksNo one-command install · Source
3269.000humanloop-evalseval · ai-agentsHumanloop Python SDK with integrated evals, dataset versioning, and human + LLM judge scoring for production pipelines.12 stars · 3 forksNo one-command install · Source
33UnrankedagentbenchOurseval · observabilityUse to put a number on harness quality — run an agent harness against a task set and get a score — so harness changes are validated by evidence, the…No signals yetNo one-command install · Source
34Unrankedevals-cookbookseval · ai-agentsOpenAI Cookbook eval recipes: task-specific templates for summarization, QA, and classification evaluation using the Evals framework.No signals yetNo commit datelisted No one-command install · Source
35Unrankedgaia-benchmarkeval · ai-agentsGAIA: benchmark of 466 real-world questions requiring multi-step reasoning, web browsing, and tool use for general AI assistants.No signals yetNo commit datelisted No one-command install · Source
36Unrankedhoneyhiveeval · authLLM evaluation and experimentation platform with session tracing, dataset management, and metric-based run comparison.No signals yetNo commit datelisted No one-command install · Source
37Unrankedpatronus-aieval · ai-agentsAutomated LLM evaluation and hallucination detection platform with a Python SDK and judge-model scoring.No signals yetNo commit datelisted No one-command install · Source

Score colour shows how many signals stand behind it, never how good it is: amber, three or more; dimmer amber, two; grey, one. The Evidence column names them.

Stars, forks and last commit are as GitHub reported them when Armory last read each repository: for most, or later. A repository may have changed since.

Leaderboard · Armory