75 results in Identity, Evals · page 1 of 4
promptfoo
CLI and library for testing, comparing, and red-teaming LLM prompts and agents with assertions and CI integration.
99.524,737 stars · 2,255 forks
Harnessclaudecursorcodexopencodegemini
evalsred-teamingcicli
openai-evals
OpenAI's official framework for evaluating LLMs and LLM-powered systems, with a registry of community eval sets.
99.519,509 stars · 3,093 forks
Harnessclaudecursorcodexopencodegemini
evalsregistrybenchmark
lm-evaluation-harness
EleutherAI's unified framework for evaluating language models on hundreds of academic benchmarks.
99.513,860 stars · 3,533 forks
Harnessclaudecursorcodexopencodegemini
evalsacademicbenchmarkharness
deepeval
Open-source LLM evaluation framework with 14+ metrics (hallucination, faithfulness, answer relevancy) and CI support.
99.418,041 stars · 1,890 forks
Harnessclaudecursorcodexopencodegemini
evalsmetricsragci
ragas
Reference-free evaluation of retrieval-augmented generation pipelines; measures faithfulness, answer relevance, and context precision.
99.415,853 stars · 1,727 forks · 1 mention · failed install test
Harnessclaudecursorcodexopencodegemini
evalsragmetrics
phoenix
Arize Phoenix: open-source LLM observability with built-in evals, span tracing, and dataset curation for RAG and agents.
99.211,286 stars · 1,086 forks · 2 mentions
Harnessclaudecursorcodexopencodegemini
evalsobservabilityragagents
swe-bench
SWE-bench: benchmark for evaluating LLMs on real-world GitHub issue resolution across 12 popular Python repositories.
99.15,762 stars · 957 forks · 9 mentions
Harnessclaudecursorcodexopencodegemini
evalscodebenchmarkagents
giskard
Open-source LLM testing framework for detecting vulnerabilities (prompt injection, hallucinations, bias) via automated scan.
99.05,838 stars · 542 forks
Harnessclaudecursorcodexopencodegemini
evalssafetyvulnerabilityscan
agenta
Open-source LLM developer platform with prompt playground, evaluation pipelines, and A/B testing for iterating on LLM apps.
98.94,670 stars · 661 forks
Harnessclaudecursorcodexopencodegemini
evalsplaygroundab-testingci
openai-simple-evals
OpenAI's lightweight benchmark suite (MMLU, HumanEval, MATH, GPQA, MGSM) for fast model capability comparisons.
98.94,621 stars · 509 forks
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkmmlusimple
big-bench
Google's Beyond the Imitation Game benchmark: 200+ diverse tasks designed to probe capabilities beyond standard NLP benchmarks.
98.83,247 stars · 618 forks
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkgoogleacademic
trulens
Evaluation and tracking for LLM and RAG applications with a feedback-function API and experiment dashboard.
98.73,530 stars · 335 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
evalsragtrackingdashboard
xlang-ai-osworld
[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Contributed by Sentinel
98.73,117 stars · 530 forks · 7 mentions
Harnessclaudecodexcursorgeminiopencode
inspect-ai
UK AISI's framework for safety evaluations of large language models, with task-based scaffolding and solver pipelines.
98.72,683 stars · 685 forks
Harnessclaudecursorcodexopencodegemini
evalssafetyaisigovernment
helm
Stanford CRFM Holistic Evaluation of Language Models: standardized benchmark suite covering accuracy, calibration, robustness, and fairness.
98.72,898 stars · 412 forks
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkacademicstanford
lighteval
Hugging Face lightweight evaluation library for LLMs across academic benchmarks, with fast local and remote inference support.
98.62,533 stars · 553 forks
Harnessclaudecursorcodexopencodegemini
evalshuggingfacebenchmarklightweight
metapriseai-orgkernel
Use when every agent action must be traceable to a specific agent instance with its own key, scoped token and tamper-evident log.
98.52,699 stars · 247 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
evalplus
Rigorous code generation evaluation framework built on top of HumanEval and MBPP with 80x more test cases.
98.31,819 stars · 208 forks
Harnessclaudecursorcodexopencodegemini
evalscode-generationhumanevalbenchmark
infisical-agent-vault
A HTTP credential proxy and vault for AI agents like Claude Code, OpenClaw, Hermes, custom agents + harnesses, and more.
Contributed by Sentinel
98.32,171 stars · 141 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
webarena
WebArena: realistic web-based environment for evaluating autonomous agents on long-horizon browser interaction tasks.
98.21,592 stars · 249 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
evalsagentsbrowserbenchmark
tau-bench
Tau-bench: agent benchmark for tool-agent-user interactions in retail and airline domains with policy-grounded evaluation.
98.11,416 stars · 215 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
evalsagentstool-usebenchmark
harbor-framework-terminal-bench
Measuring and evolving with the frontier of agent work
Contributed by Sentinel
98.1588 stars · 434 forks · 13 mentions
Harnessclaudecodexcursorgeminiopencode
wandb-weave-evals
Weights & Biases Weave evaluation framework for tracking LLM experiments, scoring model outputs, and comparing runs.
97.91,130 stars · 170 forks · 3 mentions
Harnessclaudecursorcodexopencodegemini
evalswandbexperiment-trackingscoring
langtrace
Open-source observability tool for LLMs with OpenTelemetry-based tracing, automated evals, and annotation workflows.
97.81,228 stars · 127 forks
Harnessclaudecursorcodexopencodegemini
evalsobservabilityopentelemetrytracing
Browse · Armory