Armory
Source
Browse
Evals

swe-bench

SWE-bench: benchmark for evaluating LLMs on real-world GitHub issue resolution across 12 popular Python repositories.

Score
99.1033 signals
Evidence
5,762 stars · 957 forks · 9 mentions
Last commit
as last read from GitHub; most reads are from 2 Sep 2026 or later
Listed

Install

No one-command install. Set it up from its source.

What it is

SWE-bench: benchmark for evaluating LLMs on real-world GitHub issue resolution across 12 popular Python repositories.

When to use it

SWE-bench: benchmark for evaluating LLMs on real-world GitHub issue resolution across 12 popular Python repositories.

How to install / invoke

See the source repo README: https://github.com/SWE-bench/SWE-bench

Notes

Curated evals entry. Verified 2026-05-27. The repository moved from princeton-nlp/SWE-bench, which now redirects here; the row the Sentinel feed added under the new name folded into this one (brain/lookup/duplicates/evals/swe-bench-swe-bench.md).