Personal Agent Bench

Four personal agents, taken apart at the runtime level

Every comparison out there counts features. This one asks what the machine actually does after you press send — where the model runs, how a tool reaches it, and who can read your memory.

  • 4runtimes taken apart
  • 5axes, same for each
  • 0scores published
  • 2026-09-20last verified

Not yet scored

No numeric scores yet. Scores come from recorded task runs, and those runs have not happened. What follows is what we established about each runtime, which is the thing scores would have to rest on anyway.
AxisPineInstinctMuseTown
Where it runsServer-side; browser work runs on your MacServer plus sandbox; leased cloud browserOne Linux VM per userServer-side; native Mac reaches the system
How tools are exposed26names reach you, no schemas571by its own count; manual read enforced246permissions; schemas loaded per group1105listed; 177 offered to the assistant
Where memory livesServer-side facts, no file to editRead-only markdown, log, secrets vaultPlain markdown in a home directoryServer-side wiki, memory and people
How you reach itWeb, Mac, outbound calls, REST APIWhatsApp, SMS, iMessage, its own emailWeb, effectively onlyWeb and Mac; clock and email triggers
What we could not establishInstruction text, call schema, carrierInstruction body, host, model idEverything sent to the modelInstruction set, personality placement

Each column has a page of its own. Start with Pine, or read how we score first.