Armory
Source

Browse

Search and filter by type across the catalog

75 results in Evals, Identity · page 4 of 4

EvalsPreview

gaia-benchmark

GAIA: benchmark of 466 real-world questions requiring multi-step reasoning, web browsing, and tool use for general AI assistants.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkagentstool-use
No one-command install · SourceDetails
EvalsPreview

honeyhive

LLM evaluation and experimentation platform with session tracing, dataset management, and metric-based run comparison.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
evalsexperiment-trackingtracingdataset
No one-command install · SourceDetails
EvalsPreview

patronus-ai

Automated LLM evaluation and hallucination detection platform with a Python SDK and judge-model scoring.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
evalshallucinationjudgesdk
No one-command install · SourceDetails
Browse · Armory