Armory
Source

Leaderboard

Scored on public signals; components with none are listed as Unranked · Formula

RankScoreComponentDescriptionEvidenceLast commitInstall
199.536promptfooeval · ai-agentsCLI and library for testing, comparing, and red-teaming LLM prompts and agents with assertions and CI integration.24,737 stars · 2,255 forksNo one-command install · Source
299.512openai-evalseval · ai-agentsOpenAI's official framework for evaluating LLMs and LLM-powered systems, with a registry of community eval sets.19,509 stars · 3,093 forksNo one-command install · Source
399.273phoenixeval · ai-agentsArize Phoenix: open-source LLM observability with built-in evals, span tracing, and dataset curation for RAG and agents.11,286 stars · 1,086 forks · 2 mentionsNo one-command install · Source
499.103swe-bencheval · ai-agentsSWE-bench: benchmark for evaluating LLMs on real-world GitHub issue resolution across 12 popular Python repositories.5,762 stars · 957 forks · 9 mentionsNo one-command install · Source
598.991agentaeval · ai-agentsOpen-source LLM developer platform with prompt playground, evaluation pipelines, and A/B testing for iterating on LLM apps.4,670 stars · 661 forksNo one-command install · Source
698.688lightevaleval · ai-agentsHugging Face lightweight evaluation library for LLMs across academic benchmarks, with fast local and remote inference support.2,533 stars · 553 forksNo one-command install · Source
797.923wandb-weave-evalseval · ai-agentsWeights & Biases Weave evaluation framework for tracking LLM experiments, scoring model outputs, and comparing runs.1,130 stars · 170 forks · 3 mentionsNo one-command install · Source
896.207continuous-evaleval · ai-agentsRelari's modular evaluation library for LLM pipelines with deterministic + LLM-based metrics for RAG and agent workflows.517 stars · 38 forksNo one-command install · Source
969.000humanloop-evalseval · ai-agentsHumanloop Python SDK with integrated evals, dataset versioning, and human + LLM judge scoring for production pipelines.12 stars · 3 forksNo one-command install · Source
10Unrankedpatronus-aieval · ai-agentsAutomated LLM evaluation and hallucination detection platform with a Python SDK and judge-model scoring.No signals yetNo commit datelisted No one-command install · Source

Score colour shows how many signals stand behind it, never how good it is: amber, three or more; dimmer amber, two; grey, one. The Evidence column names them.

Stars, forks and last commit are as GitHub reported them when Armory last read each repository: for most, or later. A repository may have changed since.

Leaderboard · Armory