Leaderboard
Scored on public signals; components with none are listed as Unranked · Formula
Component
Domain
Vertical
| Rank | Score | Component | Description | Evidence | Last commit | Install |
|---|---|---|---|---|---|---|
| 1 | 69.000 | humanloop-evalseval · ai-agents | Humanloop Python SDK with integrated evals, dataset versioning, and human + LLM judge scoring for production pipelines. | 12 stars · 3 forks | No one-command install · Source | |
| 2 | 94.339 | metr-task-standardeval · ai-agents | METR's Task Standard: a specification and scaffold for creating agentic tasks used in autonomous agent capability evaluations. | 192 stars · 37 forks | No one-command install · Source | |
| 3 | 96.207 | continuous-evaleval · ai-agents | Relari's modular evaluation library for LLM pipelines with deterministic + LLM-based metrics for RAG and agent workflows. | 517 stars · 38 forks | No one-command install · Source | |
| 4 | 96.786 | stonybrooknlp-appworldeval · ai-agentsContributed by Sentinel | 🌍 AppWorld: A Controllable World of Apps and People for Benchmarking Function Calling and Interactive Coding Agent, ACL'24 Best Resource Paper. | 500 stars · 78 forks · 4 mentions | No one-command install · Source | |
| 5 | 97.111 | princeton-nlp-webshopeval · ai-agentsContributed by Sentinel | [NeurIPS 2022] 🛒WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents | 589 stars · 107 forks · 3 mentions | Stale | No one-command install · Source |
| 6 | 97.923 | wandb-weave-evalseval · ai-agents | Weights & Biases Weave evaluation framework for tracking LLM experiments, scoring model outputs, and comparing runs. | 1,130 stars · 170 forks · 3 mentions | No one-command install · Source | |
| 7 | 98.141 | harbor-framework-terminal-bencheval · ai-agentsContributed by Sentinel | Measuring and evolving with the frontier of agent work | 588 stars · 434 forks · 13 mentions | No one-command install · Source | |
| 8 | 98.177 | tau-bencheval · ai-agents | Tau-bench: agent benchmark for tool-agent-user interactions in retail and airline domains with policy-grounded evaluation. | 1,416 stars · 215 forks · 1 mention | No one-command install · Source | |
| 9 | 98.688 | lightevaleval · ai-agents | Hugging Face lightweight evaluation library for LLMs across academic benchmarks, with fast local and remote inference support. | 2,533 stars · 553 forks | No one-command install · Source | |
| 10 | 98.776 | xlang-ai-osworldeval · ai-agentsContributed by Sentinel | [NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments | 3,117 stars · 530 forks · 7 mentions | No one-command install · Source | |
| 11 | 98.956 | openai-simple-evalseval · ai-agents | OpenAI's lightweight benchmark suite (MMLU, HumanEval, MATH, GPQA, MGSM) for fast model capability comparisons. | 4,621 stars · 509 forks | No one-command install · Source | |
| 12 | 98.991 | agentaeval · ai-agents | Open-source LLM developer platform with prompt playground, evaluation pipelines, and A/B testing for iterating on LLM apps. | 4,670 stars · 661 forks | No one-command install · Source | |
| 13 | 99.052 | giskardeval · ai-agents | Open-source LLM testing framework for detecting vulnerabilities (prompt injection, hallucinations, bias) via automated scan. | 5,838 stars · 542 forks | No one-command install · Source | |
| 14 | 99.103 | swe-bencheval · ai-agents | SWE-bench: benchmark for evaluating LLMs on real-world GitHub issue resolution across 12 popular Python repositories. | 5,762 stars · 957 forks · 9 mentions | No one-command install · Source | |
| 15 | 99.273 | phoenixeval · ai-agents | Arize Phoenix: open-source LLM observability with built-in evals, span tracing, and dataset curation for RAG and agents. | 11,286 stars · 1,086 forks · 2 mentions | No one-command install · Source | |
| 16 | 99.512 | openai-evalseval · ai-agents | OpenAI's official framework for evaluating LLMs and LLM-powered systems, with a registry of community eval sets. | 19,509 stars · 3,093 forks | No one-command install · Source | |
| 17 | 99.536 | promptfooeval · ai-agents | CLI and library for testing, comparing, and red-teaming LLM prompts and agents with assertions and CI integration. | 24,737 stars · 2,255 forks | No one-command install · Source | |
| 18 | Unranked | evals-cookbookseval · ai-agents | OpenAI Cookbook eval recipes: task-specific templates for summarization, QA, and classification evaluation using the Evals framework. | No signals yet | No commit datelisted | No one-command install · Source |
| 19 | Unranked | gaia-benchmarkeval · ai-agents | GAIA: benchmark of 466 real-world questions requiring multi-step reasoning, web browsing, and tool use for general AI assistants. | No signals yet | No commit datelisted | No one-command install · Source |
| 20 | Unranked | patronus-aieval · ai-agents | Automated LLM evaluation and hallucination detection platform with a Python SDK and judge-model scoring. | No signals yet | No commit datelisted | No one-command install · Source |
Score colour shows how many signals stand behind it, never how good it is: amber, three or more; dimmer amber, two; grey, one. The Evidence column names them.
Stars, forks and last commit are as GitHub reported them when Armory last read each repository: for most, or later. A repository may have changed since.