Armory
Source

Browse

Search and filter by type across the catalog

37 results in Evals 路 page 2 of 2

EvalsExperimental

stonybrooknlp-appworld

馃實 AppWorld: A Controllable World of Apps and People for Benchmarking Function Calling and Interactive Coding Agent, ACL'24 Best Resource Paper.

Contributed by Sentinel

96.7500 stars 路 78 forks 路 4 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install 路 SourceDetails
EvalsPreview

continuous-eval

Relari's modular evaluation library for LLM pipelines with deterministic + LLM-based metrics for RAG and agent workflows.

96.2517 stars 路 38 forks
Harnessclaudecursorcodexopencodegemini
evalsragagentsmetrics
No one-command install 路 SourceDetails
EvalsPreview

metr-task-standard

METR's Task Standard: a specification and scaffold for creating agentic tasks used in autonomous agent capability evaluations.

94.3192 stars 路 37 forks
Harnessclaudecursorcodexopencodegemini
evalsagentstask-standardsafety
No one-command install 路 SourceDetails
EvalsExperimental

harbor-framework-terminal-bench-2-1

Terminal-Bench 2.1

Contributed by Sentinel

94.3119 stars 路 62 forks 路 3 mentions
Harnessclaudecodexcursorgeminiopencode
evals
No one-command install 路 SourceDetails
EvalsPreview

vellum-evals

Vellum evaluation SDK for running LLM test suites with custom metrics, dataset pinning, and CI workflow integration.

90.582 stars 路 20 forks
Harnessclaudecursorcodexopencodegemini
evalssdkcidataset
No one-command install 路 SourceDetails
EvalsPreview

braintrust

Developer platform for logging, evaluating, and comparing LLM experiments with dataset versioning and scoring functions.

82.327 stars 路 12 forks
Harnessclaudecursorcodexopencodegemini
evalsexperiment-trackingsdk
No one-command install 路 SourceDetails
EvalsPreview

galileo-evaluate

Galileo evaluation and observability SDK for detecting hallucinations, data errors, and model weaknesses in LLM pipelines.

80.322 stars 路 11 forks
Harnessclaudecursorcodexopencodegemini
evalshallucinationobservabilitysdk
No one-command install 路 SourceDetails
EvalsPreview

humanloop-evals

Humanloop Python SDK with integrated evals, dataset versioning, and human + LLM judge scoring for production pipelines.

69.012 stars 路 3 forks
Harnessclaudecursorcodexopencodegemini
evalshuman-evaldatasetsdk
No one-command install 路 SourceDetails
EvalsPreview

agentbench

Use to put a number on harness quality: run an agent harness against a task set and get a score, so harness changes are validated by evidence. It is the eval backbone of a self-improving loop.

UnrankedNo signals yet
Harnessclaudecodex
evalbenchmarkscoringharness
No one-command install 路 SourceDetails
EvalsPreview

evals-cookbooks

OpenAI Cookbook eval recipes: task-specific templates for summarization, QA, and classification evaluation using the Evals framework.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
evalscookbooktemplatesopenai
No one-command install 路 SourceDetails
EvalsPreview

gaia-benchmark

GAIA: benchmark of 466 real-world questions requiring multi-step reasoning, web browsing, and tool use for general AI assistants.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkagentstool-use
No one-command install 路 SourceDetails
EvalsPreview

honeyhive

LLM evaluation and experimentation platform with session tracing, dataset management, and metric-based run comparison.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
evalsexperiment-trackingtracingdataset
No one-command install 路 SourceDetails
EvalsPreview

patronus-ai

Automated LLM evaluation and hallucination detection platform with a Python SDK and judge-model scoring.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
evalshallucinationjudgesdk
No one-command install 路 SourceDetails
Browse 路 Armory