Single-agent software engineer achieving 49% on SWE-bench Verified.
Recent activity
Version cuts and proof, newest first โ the living track record.
Spec sheet
The benchmark fields โ designed for comparison across teams.
- Topology
- Solo + Tools
- Agent count
- 1
- Platform
- Claude API
- Runs on
- Claude API
- Industries
- software-delivery
- Task kinds
- bug-fixingcode-editingtest-execution
- Trust tier
- Self-Reported
- Proof entries
- 1
Topology & roster
Single agent (solo-plus-tools). Claude 3.5 Sonnet operates two tools: a persistent Bash shell and a custom file editor. The model determines its own workflow freely โ "the model is free to choose how it moves from step to step, rather than having strict and discrete transitions."
System wiring
Node details
Typical Solo + Tools layout โ schematic, not verified wiring
HumanHuman operatorHuman gate
- Tool
- Human operator
- Autonomy
- Human-gated
- directs โ Agent
BuilderAgent
- Tool
- Agent
- Autonomy
- Runs autonomously
- writes to โ Tool A
- writes to โ Tool B
- directs โ Human operator
ResourceTool A
- Tool
- Tool A
- Autonomy
- Runs autonomously
- writes to โ Agent
ResourceTool B
- Tool
- Tool B
- Autonomy
- Runs autonomously
- writes to โ Agent
How a typical Solo + Tools team handles a task
Typical Solo + Tools layout โ schematic, not verified wiring
Task arrives
Human operator directs Agent.
Agent builds the work
Agent builds the work.
The artifact lands
The artifact lands in Tool A: Agent contributes via "writes to". The artifact lands in Tool B: Agent contributes via "writes to".
Human holds the last word
Human operator holds final approval.
Replicate a typical Solo + Tools setup
Typical Solo + Tools layout โ schematic, not verified wiring
Ingredients
- HumanHuman operator
- BuilderAgent
- ResourceTool A
- ResourceTool B
Setup order
- 1.Provision the substrate: Tool A and Tool B.
- 2.Wire Agent: it receives "directs" from Human operator.
- 3.Declare the human gate: Human operator holds final approval.
Performance metrics
Windowed metrics with provenance. [unknown] means it was not tracked โ an honest hole beats an invented figure.
Score on SWE-bench Verified (500 GitHub issues). Previous SOTA: 45%. Source: https://www.anthropic.com/research/swe-bench-sonnet [evidence_linked]
Previous best on SWE-bench Verified before Claude 3.5 Sonnet result. Source: https://www.anthropic.com/research/swe-bench-sonnet [evidence_linked]
Token economics
Cost transparency is part of the honesty architecture. [unknown] means it was not tracked โ not that it is zero.
Blueprint
Operational DNA โ why it works, how it was built, and how it is overseen. Not files for sale; knowledge of the design.
Minimal scaffolding gives the model maximum flexibility to choose its own strategy. The persistent Bash shell maintains state across tool calls, allowing iterative debugging without losing context. According to the source, the upgraded model improved from 33% to 49% on the same benchmark.
Minimal scaffolding. Two tools only: Bash (persistent state across calls) and str_replace_editor. The LLM autonomously decides when to read files, run tests, or edit code. Claude 3.5 Sonnet (upgraded version) was the model used.
No human-in-the-loop described for this team design. Evaluated on the SWE-bench Verified benchmark (500 real GitHub issues).
Proof (1)
The team's shared track record โ tasks, incidents, lessons, milestones. Per-entry provenance tags are always visible.
- MilestoneOct 22, 2024evidence-linked
Claude 3.5 Sonnet achieves 49% on SWE-bench Verified
Beats previous SOTA of 45%. Prior Claude 3.5 Sonnet scored 33%; Claude 3 Opus 22%. Two tools only: Bash + str_replace_editor. Source: Anthropic research page.
https://www.anthropic.com/research/swe-bench-sonnet
Sign in to add a proof entry.
Sign inAttestations (0)
Named third-party statements from people with first-hand experience. Attestations are what separates Peer-Attested from Evidence-Linked.
No attestations yet. Worked with this configuration or agent? Attest to it using the form below โ attestations are named third-party statements and are what separates Peer-Attested from Evidence-Linked.
Sign in to attest to this team.
Sign in