How we score personal agents
Most agent comparisons count features. A feature list tells you what a product claims; it tells you nothing about what happens after you press send.
We went the other way and took four products apart at the runtime level, then built the axes out of what we found. The axes are below, each phrased as the question it answers.
Where it runs
Does the model run near your machine or far from it, and what can it touch?
How tools are exposed
When does a tool definition enter the context, and who decides which ones?
Where memory lives
Files you can open, or a server-side store you can only reach through the product?
How you reach it
What is the default surface, and what happens on the ones that are not?
What we could not establish
Stated plainly, because a bench that only reports wins is not a bench.
Why there are no numbers yet
No numeric scores yet. Scores come from recorded task runs, and those runs have not happened. What follows is what we established about each runtime, which is the thing scores would have to rest on anyway. A competing site currently shows 116 products against 15 dimensions. We can't tell you where any of those cells came from, and neither can it. We'd rather publish four columns we can defend.
What we publish, and what we don't
Everything here is restated in our own words. We do not reproduce instruction text, internal documents or schemas from the products we examine, and we do not describe how we obtained anything. Findings, not artefacts.
Where we could not establish something, it sits in the last row of the table rather than quietly missing.
Next: the four teardowns these axes came out of, or what the existing agent benchmarks measure instead.