Armory
Source

Leaderboard

Scored on public signals; components with none are listed as Unranked · Formula

RankScoreComponentDescriptionEvidenceLast commitInstall
199.536promptfooeval · ai-agentsCLI and library for testing, comparing, and red-teaming LLM prompts and agents with assertions and CI integration.24,737 stars · 2,255 forksNo one-command install · Source
299.512openai-evalseval · ai-agentsOpenAI's official framework for evaluating LLMs and LLM-powered systems, with a registry of community eval sets.19,509 stars · 3,093 forksNo one-command install · Source
399.273phoenixeval · ai-agentsArize Phoenix: open-source LLM observability with built-in evals, span tracing, and dataset curation for RAG and agents.11,286 stars · 1,086 forks · 2 mentionsNo one-command install · Source
499.103swe-bencheval · ai-agentsSWE-bench: benchmark for evaluating LLMs on real-world GitHub issue resolution across 12 popular Python repositories.5,762 stars · 957 forks · 9 mentionsNo one-command install · Source
599.052giskardeval · ai-agentsOpen-source LLM testing framework for detecting vulnerabilities (prompt injection, hallucinations, bias) via automated scan.5,838 stars · 542 forksNo one-command install · Source
698.991agentaeval · ai-agentsOpen-source LLM developer platform with prompt playground, evaluation pipelines, and A/B testing for iterating on LLM apps.4,670 stars · 661 forksNo one-command install · Source
798.956openai-simple-evalseval · ai-agentsOpenAI's lightweight benchmark suite (MMLU, HumanEval, MATH, GPQA, MGSM) for fast model capability comparisons.4,621 stars · 509 forksNo one-command install · Source
898.776xlang-ai-osworldeval · ai-agentsContributed by Sentinel[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments3,117 stars · 530 forks · 7 mentionsNo one-command install · Source
998.688lightevaleval · ai-agentsHugging Face lightweight evaluation library for LLMs across academic benchmarks, with fast local and remote inference support.2,533 stars · 553 forksNo one-command install · Source
1098.177tau-bencheval · ai-agentsTau-bench: agent benchmark for tool-agent-user interactions in retail and airline domains with policy-grounded evaluation.1,416 stars · 215 forks · 1 mentionNo one-command install · Source
1198.141harbor-framework-terminal-bencheval · ai-agentsContributed by SentinelMeasuring and evolving with the frontier of agent work588 stars · 434 forks · 13 mentionsNo one-command install · Source
1297.923wandb-weave-evalseval · ai-agentsWeights & Biases Weave evaluation framework for tracking LLM experiments, scoring model outputs, and comparing runs.1,130 stars · 170 forks · 3 mentionsNo one-command install · Source
1397.111princeton-nlp-webshopeval · ai-agentsContributed by Sentinel[NeurIPS 2022] 🛒WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents589 stars · 107 forks · 3 mentionsStaleNo one-command install · Source
1496.786stonybrooknlp-appworldeval · ai-agentsContributed by Sentinel🌍 AppWorld: A Controllable World of Apps and People for Benchmarking Function Calling and Interactive Coding Agent, ACL'24 Best Resource Paper.500 stars · 78 forks · 4 mentionsNo one-command install · Source
1596.207continuous-evaleval · ai-agentsRelari's modular evaluation library for LLM pipelines with deterministic + LLM-based metrics for RAG and agent workflows.517 stars · 38 forksNo one-command install · Source
1694.339metr-task-standardeval · ai-agentsMETR's Task Standard: a specification and scaffold for creating agentic tasks used in autonomous agent capability evaluations.192 stars · 37 forksNo one-command install · Source
1769.000humanloop-evalseval · ai-agentsHumanloop Python SDK with integrated evals, dataset versioning, and human + LLM judge scoring for production pipelines.12 stars · 3 forksNo one-command install · Source
18Unrankedevals-cookbookseval · ai-agentsOpenAI Cookbook eval recipes: task-specific templates for summarization, QA, and classification evaluation using the Evals framework.No signals yetNo commit datelisted No one-command install · Source
19Unrankedgaia-benchmarkeval · ai-agentsGAIA: benchmark of 466 real-world questions requiring multi-step reasoning, web browsing, and tool use for general AI assistants.No signals yetNo commit datelisted No one-command install · Source
20Unrankedpatronus-aieval · ai-agentsAutomated LLM evaluation and hallucination detection platform with a Python SDK and judge-model scoring.No signals yetNo commit datelisted No one-command install · Source

Score colour shows how many signals stand behind it, never how good it is: amber, three or more; dimmer amber, two; grey, one. The Evidence column names them.

Stars, forks and last commit are as GitHub reported them when Armory last read each repository: for most, or later. A repository may have changed since.

Leaderboard · Armory