75 results in Identity, Evals · page 4 of 4
gaia-benchmark
GAIA: benchmark of 466 real-world questions requiring multi-step reasoning, web browsing, and tool use for general AI assistants.
UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkagentstool-use
honeyhive
LLM evaluation and experimentation platform with session tracing, dataset management, and metric-based run comparison.
UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
evalsexperiment-trackingtracingdataset
patronus-ai
Automated LLM evaluation and hallucination detection platform with a Python SDK and judge-model scoring.
UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
evalshallucinationjudgesdk
Browse · Armory