openai-evals
OpenAI's official framework for evaluating LLMs and LLM-powered systems, with a registry of community eval sets.
- Score
- 99.5122 signals
- Evidence
- 19,509 stars · 3,093 forks
- Last commit
- as last read from GitHub; most reads are from 2 Sep 2026 or later
- Listed
Install
No one-command install. Set it up from its source.
What it is
OpenAI's official framework for evaluating LLMs and LLM-powered systems, with a registry of community eval sets.
When to use it
OpenAI's official framework for evaluating LLMs and LLM-powered systems, with a registry of community eval sets.
Notes
Curated evals entry. Verified 2026-05-27.