tau-bench
Tau-bench: agent benchmark for tool-agent-user interactions in retail and airline domains with policy-grounded evaluation.
- Score
- 98.1773 signals
- Evidence
- 1,416 stars · 215 forks · 1 mention
- Last commit
- as last read from GitHub; most reads are from 2 Sep 2026 or later
- Listed
Install
No one-command install. Set it up from its source.
Alternatives · Evals
- princeton-nlp-webshop589 stars · 107 forks · 3 mentions97.111
What it is
Tau-bench: agent benchmark for tool-agent-user interactions in retail and airline domains with policy-grounded evaluation.
When to use it
Tau-bench: agent benchmark for tool-agent-user interactions in retail and airline domains with policy-grounded evaluation.
Notes
Curated evals entry. Verified 2026-05-27.