toolsurface

MCP tool-surface review

See how well an agent can actually use your MCP server.

Your tool manifest is loaded into the model's context on every single request, and read with no memory and no way to ask a clarifying question. When it is ambiguous you do not get an error, you get a confident wrong call. Paste a URL below and we will score it.

Public endpoints only, and we send exactly one tools/list call. Behind auth? Paste your manifest to kurt@toolsurface.dev instead.

Why this exists

We scored every reachable server in the public registry

19,008 servers enumerated, 176 live manifests pulled and scored against documented agent failure modes. The median scores a B, which is the uncomfortable part: this is not a field of disasters, it is a field where good enough that nobody investigates is the default, and the cost is paid quietly on every request.

97%carry at least one finding
64%never say what happens on failure
53%ship no tool annotations at all
89,254tokens in the largest manifest, per request

Read the full field study →

What we do

Measure it, fix it, keep it fixed

None of this is anyone's fault. Generating a manifest from an existing API specification produces exactly these results, and that was the sensible thing to do at the time. The fixes are unglamorous and mostly mechanical. They have simply not been anybody's job yet.

Review Free

The scan above, plus a written report naming the specific tools and parameters involved. Useful whether or not you ever reply.

Repair from $2,500

We rewrite the surface: descriptions, schemas, enums, annotations, pagination contracts, and consolidation where it is too wide to select from reliably. Fixed price, agreed up front.

Monitor $250/mo

Re-scored on every release and gated in CI, so a good surface does not quietly regress the next time someone adds an endpoint.

Run the scan, then send the result to kurt@toolsurface.dev. We would rather be useful first and discuss the rest afterwards.