37 results in Evals 路 page 2 of 2
Harness Claude Code Cursor Codex Gemini OpenCode stonybrooknlp-appworld 馃實 AppWorld: A Controllable World of Apps and People for Benchmarking Function Calling and Interactive Coding Agent, ACL'24 Best Resource Paper.
Contributed by Sentinel
96.7 500 stars 路 78 forks 路 4 mentions
Harness claude codex cursor gemini opencode
continuous-eval Relari's modular evaluation library for LLM pipelines with deterministic + LLM-based metrics for RAG and agent workflows.
96.2 517 stars 路 38 forks
Harness claude cursor codex opencode gemini
evals rag agents metrics
metr-task-standard METR's Task Standard: a specification and scaffold for creating agentic tasks used in autonomous agent capability evaluations.
94.3 192 stars 路 37 forks
Harness claude cursor codex opencode gemini
evals agents task-standard safety
harbor-framework-terminal-bench-2-1 Terminal-Bench 2.1
Contributed by Sentinel
94.3 119 stars 路 62 forks 路 3 mentions
Harness claude codex cursor gemini opencode
evals
vellum-evals Vellum evaluation SDK for running LLM test suites with custom metrics, dataset pinning, and CI workflow integration.
90.5 82 stars 路 20 forks
Harness claude cursor codex opencode gemini
evals sdk ci dataset
braintrust Developer platform for logging, evaluating, and comparing LLM experiments with dataset versioning and scoring functions.
82.3 27 stars 路 12 forks
Harness claude cursor codex opencode gemini
evals experiment-tracking sdk
galileo-evaluate Galileo evaluation and observability SDK for detecting hallucinations, data errors, and model weaknesses in LLM pipelines.
80.3 22 stars 路 11 forks
Harness claude cursor codex opencode gemini
evals hallucination observability sdk
humanloop-evals Humanloop Python SDK with integrated evals, dataset versioning, and human + LLM judge scoring for production pipelines.
69.0 12 stars 路 3 forks
Harness claude cursor codex opencode gemini
evals human-eval dataset sdk
agentbench Use to put a number on harness quality: run an agent harness against a task set and get a score, so harness changes are validated by evidence. It is the eval backbone of a self-improving loop.
Unranked No signals yet
Harness claude codex
eval benchmark scoring harness
evals-cookbooks OpenAI Cookbook eval recipes: task-specific templates for summarization, QA, and classification evaluation using the Evals framework.
Unranked No signals yet
Harness claude cursor codex opencode gemini
evals cookbook templates openai
gaia-benchmark GAIA: benchmark of 466 real-world questions requiring multi-step reasoning, web browsing, and tool use for general AI assistants.
Unranked No signals yet
Harness claude cursor codex opencode gemini
evals benchmark agents tool-use
honeyhive LLM evaluation and experimentation platform with session tracing, dataset management, and metric-based run comparison.
Unranked No signals yet
Harness claude cursor codex opencode gemini
evals experiment-tracking tracing dataset
patronus-ai Automated LLM evaluation and hallucination detection platform with a Python SDK and judge-model scoring.
Unranked No signals yet
Harness claude cursor codex opencode gemini
evals hallucination judge sdk
Previous Page 2 of 2
Browse 路 Armory