braintrust
Developer platform for logging, evaluating, and comparing LLM experiments with dataset versioning and scoring functions.
- Score
- 82.3192 signals
- Evidence
- 27 stars · 12 forks
- Last commit
- as last read from GitHub; most reads are from 2 Sep 2026 or later
- Listed
Install
No one-command install. Set it up from its source.
Alternatives · Evals
- agenta4,670 stars · 661 forks98.991
- trulens3,530 stars · 335 forks · 1 mention98.783
- wandb-weave-evals1,130 stars · 170 forks · 3 mentions97.923
What it is
Developer platform for logging, evaluating, and comparing LLM experiments with dataset versioning and scoring functions.
When to use it
Developer platform for logging, evaluating, and comparing LLM experiments with dataset versioning and scoring functions.
Notes
Curated evals entry. Verified 2026-05-27.