swe-bench
SWE-bench: benchmark for evaluating LLMs on real-world GitHub issue resolution across 12 popular Python repositories.
- Score
- 99.1033 signals
- Evidence
- 5,762 stars · 957 forks · 9 mentions
- Last commit
- as last read from GitHub; most reads are from 2 Sep 2026 or later
- Listed
Install
No one-command install. Set it up from its source.
What it is
SWE-bench: benchmark for evaluating LLMs on real-world GitHub issue resolution across 12 popular Python repositories.
When to use it
SWE-bench: benchmark for evaluating LLMs on real-world GitHub issue resolution across 12 popular Python repositories.
How to install / invoke
See the source repo README: https://github.com/SWE-bench/SWE-bench
Notes
Curated evals entry. Verified 2026-05-27. The repository moved from princeton-nlp/SWE-bench, which now redirects here; the row the Sentinel feed added under the new name folded into this one (brain/lookup/duplicates/evals/swe-bench-swe-bench.md).