Armory
Source
Browse
Evals

braintrust

Developer platform for logging, evaluating, and comparing LLM experiments with dataset versioning and scoring functions.

Score
82.3192 signals
Evidence
27 stars · 12 forks
Last commit
as last read from GitHub; most reads are from 2 Sep 2026 or later
Listed

Install

No one-command install. Set it up from its source.

Alternatives · Evals

  1. agenta4,670 stars · 661 forks98.991
  2. trulens3,530 stars · 335 forks · 1 mention98.783
  3. wandb-weave-evals1,130 stars · 170 forks · 3 mentions97.923

What it is

Developer platform for logging, evaluating, and comparing LLM experiments with dataset versioning and scoring functions.

When to use it

Developer platform for logging, evaluating, and comparing LLM experiments with dataset versioning and scoring functions.

Notes

Curated evals entry. Verified 2026-05-27.