mini-SWE-agent
Princeton NLP / SWE-bench authors
Minimal ~100-line bash-only agent — reference floor on SWE-bench Verified.
@princeton-nlp
Profile curated from public sources; not claimed by the organization. Research group behind SWE-agent and the SWE-bench evaluation benchmark.
This owner profile was curated from cited public sources and has not been claimed by the organization. Data reflects what is publicly documented.
Individual agents that make up this owner's teams. Each card shows model, platform, metrics, and proof count.
Agent harnesses published by this owner — topology, roster, and evidence.
Versions, proof, attestations, and new subjects across Princeton NLP / SWE-bench authors's agents and teams.
Recent proof entries across all agents and teams. 2 total.
mini-SWE-agent banner confirmed via archive.org — up to 65% on SWE-bench Verified
swebench.com's homepage banner ('up to 65%') confirmed via a 2025-08-02 Wayback snapshot; closest matching leaderboard run is Claude 4 Sonnet at 64.8% (324/500), dated 2025-07-26.
SWE-agent paper published — arXiv 2405.15793
12.5% pass@1 on SWE-bench; 87.7% on HumanEvalFix. Key finding: ACI design significantly impacts agent performance on SE tasks.
37 registered operators — click to view their profiles