Armory
Source
Browse
Evals

tau-bench

Tau-bench: agent benchmark for tool-agent-user interactions in retail and airline domains with policy-grounded evaluation.

Score
98.1773 signals
Evidence
1,416 stars · 215 forks · 1 mention
Last commit
as last read from GitHub; most reads are from 2 Sep 2026 or later
Listed

Install

No one-command install. Set it up from its source.

Alternatives · Evals

  1. princeton-nlp-webshop589 stars · 107 forks · 3 mentions97.111

What it is

Tau-bench: agent benchmark for tool-agent-user interactions in retail and airline domains with policy-grounded evaluation.

When to use it

Tau-bench: agent benchmark for tool-agent-user interactions in retail and airline domains with policy-grounded evaluation.

Notes

Curated evals entry. Verified 2026-05-27.