๐Ÿ•น๏ธ

Magentic-One

Self-ReportedCurated

Microsoft Research generalist 5-agent system: GAIA 32.33%, WebArena 32.8%.

Microsoft Researchยท Operating since Nov 7, 2024ยท active
Curated from arXiv 2411.04468 โ€” Magentic-One โ€” not claimed by or endorsed by the organization. Metrics cited only as the source states. Absent metrics render as [unknown].

Recent activity

Version cuts and proof, newest first โ€” the living track record.

  1. Artifact ยท Magentic-One paper published (arXiv 2411.04468)1y ago

Spec sheet

The benchmark fields โ€” designed for comparison across teams.

Topology
Supervisor
Agent count
5
Platform
AutoGen
Runs on
AutoGen ร—5
Industries
researchsoftware-deliverydata-extraction
Task kinds
web-navigationfile-operationscode-executioncomplex-reasoning
Trust tier
Self-Reported
Proof entries
1

Topology & roster

Supervisor

Hierarchical. The Orchestrator (lead agent) plans, tracks progress, and re-plans to recover from errors, directing four specialist agents: WebSurfer (web browser), FileSurfer (file navigation), Coder (Python), and ComputerTerminal (code execution). Modular: "agents to be added or removed from the team without additional prompt tuning or training."

System wiring

Typical Supervisor layout โ€” schematic, not verified wiring
Typical role-level schematic โ€” not verified wiringdirectsdispatchesdispatchesreportsreportsHuman operatorHumanoperatorHUMANGATESupervisorSupervisorORCHESTRATORWorker agent AWorker agentABUILDERWorker agent BWorker agentBBUILDER
Node details

Typical Supervisor layout โ€” schematic, not verified wiring

HumanHuman operatorHuman gate
Tool
Human operator
Autonomy
Human-gated
Sends
  • directs โ†’ Supervisor
OrchestratorSupervisor
Tool
Supervisor
Autonomy
Runs autonomously
Sends
  • dispatches โ†’ Worker agent A
  • dispatches โ†’ Worker agent B
Receives
  • directs โ† Human operator
  • reports โ† Worker agent A
  • reports โ† Worker agent B
BuilderWorker agent A
Tool
Worker agent A
Autonomy
Runs autonomously
Sends
  • reports โ†’ Supervisor
Receives
  • dispatches โ† Supervisor
BuilderWorker agent B
Tool
Worker agent B
Autonomy
Runs autonomously
Sends
  • reports โ†’ Supervisor
Receives
  • dispatches โ† Supervisor

How a typical Supervisor team handles a task

Typical Supervisor layout โ€” schematic, not verified wiring

  1. Task arrives

    Human operator directs Supervisor.

  2. The orchestrator routes the work

    Supervisor dispatches build work to Worker agent A and dispatches build work to Worker agent B.

  3. The builders execute

    Worker agent A and Worker agent B build the work.

  4. Human holds the last word

    Human operator holds final approval.

Replicate a typical Supervisor setup

Typical Supervisor layout โ€” schematic, not verified wiring

Ingredients

  • HumanHuman operator
  • OrchestratorSupervisor
  • BuilderWorker agent A
  • BuilderWorker agent B

Setup order

  1. 1.Stand up the orchestrator: Supervisor.
  2. 2.Wire Worker agent A: it receives "dispatches" from Supervisor and sends "reports" to Supervisor. Wire Worker agent B: it receives "dispatches" from Supervisor and sends "reports" to Supervisor.
  3. 3.Declare the human gate: Human operator holds final approval.

Performance metrics

Windowed metrics with provenance. [unknown] means it was not tracked โ€” an honest hole beats an invented figure.

GAIA benchmark score
32.3%
evidence-linked

ยฑ5.3 confidence interval; default GPT-4o-2024-05-13 configuration. Source: arXiv 2411.04468 [evidence_linked]

as of Nov 7, 2024
WebArena score
32.8%
evidence-linked

ยฑ3.2 confidence interval; default GPT-4o configuration. Source: arXiv 2411.04468 [evidence_linked]

as of Nov 7, 2024
AssistantBench accuracy
25.3%
evidence-linked

ยฑ6.3; default GPT-4o-2024-05-13. Source: arXiv 2411.04468 [evidence_linked]

as of Nov 7, 2024

Token economics

Cost transparency is part of the honesty architecture. [unknown] means it was not tracked โ€” not that it is zero.

No cost metrics on record. Cost tracking is hard across runtimes; honest absence beats invented figures.

Blueprint

Operational DNA โ€” why it works, how it was built, and how it is overseen. Not files for sale; knowledge of the design.

Why it works

Specialist agents each own a specific skill (web, files, code) that the Orchestrator cannot perform directly. The Orchestrator re-plans on error rather than failing silently. Modularity allows extending the team without retraining. GAIA benchmark: 32.33% (ยฑ5.3) with GPT-4o; 38.00% (ยฑ5.5) with GPT-4o + o1-preview.

How it was built

Built on AutoGen (Microsoft). Default model: GPT-4o-2024-05-13, with optional integration of o1-preview for enhanced reasoning. Evaluation tool AutoGenBench provides built-in controls for repetition and isolation. Open-source.

Oversight model

No human-in-the-loop described in the paper; evaluated on automated benchmarks. Designed as a generalist agentic system for complex tasks requiring multi-step reasoning.

Proof (1)

The team's shared track record โ€” tasks, incidents, lessons, milestones. Per-entry provenance tags are always visible.

  1. ArtifactNov 7, 2024evidence-linked

    Magentic-One paper published (arXiv 2411.04468)

    Five-agent system achieves GAIA 32.33% (ยฑ5.3), WebArena 32.8% (ยฑ3.2), AssistantBench 25.3% accuracy (ยฑ6.3) with GPT-4o. With o1-preview: GAIA 38.00% (ยฑ5.5).

    https://arxiv.org/abs/2411.04468

Sign in to add a proof entry.

Sign in

Attestations (0)

Named third-party statements from people with first-hand experience. Attestations are what separates Peer-Attested from Evidence-Linked.

No attestations yet. Worked with this configuration or agent? Attest to it using the form below โ€” attestations are named third-party statements and are what separates Peer-Attested from Evidence-Linked.

Sign in to attest to this team.

Sign in