The 4 Tests a Tool Has to Pass Before It Touches Client Work

I run client SEO and AEO out of Claude Code, so a tool is not something I look at, it's something my agent calls. Here are the 4 tests I run before one gets near a client account, and what happened when I applied them to the AEO tool market.

I was in the middle of a keyword strategy for a client. I asked my agent for one specific cut of the data, and it didn't come back. An error on the endpoint, and the field I needed wasn't on the plan anyway.

Everything stopped at once. Me, the agent, and the client's timeline.

That's the moment my tool criteria changed. Not a demo, not a feature page. A Tuesday afternoon where one API call decided whether the work shipped that week.

The way I work changed what I need from a tool

I don't sit in dashboards. I brief an agent, it reaches into the tools, and I keep it pointed at the goal. The tool is not something I look at, it's something my agent calls.

That flips the evaluation. Engine coverage and chart quality stop mattering much. 4 other things start mattering a lot.

The tool can break my flow, and the agent's

A dashboard I can click around when something is off. A failed call inside an agent run kills the whole chain, and the agent can't route around it for me.

The direction can flip

When a tool only speaks dashboard, I adapt my process to somebody else's UI, fetch numbers by hand, and the agent becomes a passenger I paste screenshots to. I'm serving the tool instead of the client.

The price drifts

Credits, seats, per-brand fees. I quote a retainer in January and my cost base moves every month after. Per-brand pricing is the worst of it, since every client I sign makes the tool more expensive exactly when the margin is thinnest.

The manual work comes back

Exports, screenshots, charts pasted into a deck the client skims once. My job is to deliver organic growth through content, not to spend my time building dashboards.

So here are the 4 tests. I ran them on the AI visibility market this summer, because that is the freshest example I have. The tests are not about AEO though. They work on any tool you let near a client account.

Test 1: can my agent write to it, or only read from it?

"Has an API" is not the answer. The question is whether the agent can act.

A year ago this test eliminated almost everything. Not anymore. As of August 2026, Peec, Profound, Otterly, Rankscale, Semrush and Ahrefs all publish an MCP server. Support for agents stopped being a differentiator, which is worth saying out loud because I used to sell it as one.

What replaced it is depth and gating. Semrush's MCP methods are read-only, on a plan that starts at $199. Profound's action layer needs the $399 tier or an enterprise contract. Otterly gates MCP to the $189 plan, since the $29 one is tracking only. Ahrefs is remote-only and read-heavy from $398.

Read access saves me a browser tab. Write access saves me an afternoon. The difference is whether onboarding a client reads like this:

Set up a brand for this client's site, create 3 topics covering their buying journey, add 10 prompts under each, run the technical audit, then run everything across all 4 engines.

Or like 40 minutes of forms. If you want the mechanics of wiring one of these up, I wrote the walkthrough here: how to connect AEO Copilot with Claude.

Test 2: can the data leave?

A score tells me where a client stands. It doesn't tell me what to publish next.

For that I need the raw material: the answer text, the sources cited, the competitors that showed up instead of my client, the prompts that came back empty. Then I need to put that next to what the client already ranks for in search, because the overlap is where the work is.

That crossing is the actual job, and no dashboard is going to do it for me. It has to happen somewhere else, which means the data has to be able to leave the tool.

When it can, the analysis is one conversation: the agent reads the AI answers, pulls the same client's Search Console data, and hands back a plan for the next 2 weeks. I don't export anything. I read it and decide what to argue with.

Test 3: what does it cost per client?

Sticker price is the wrong number. Cost per client per month is the one that comes out of my margin.

Run it on 10 client brands and the spread is brutal. A flat plan at $99 for unlimited brands is $9.90 a client. The same job on a $189 plan is $18.90. On a $399 plan with a 100-prompt cap it's $39.90 a client, and 10 prompts each, which is barely a sample.

Under $10 a client, this is a profitable line on a retainer. Over $30, it eats the margin on the exact service I'm selling. The whole calculation takes 30 seconds, and it's the one I got tired of doing on other people's pricing pages. Which is why mine is a flat number with no per-brand fee.

Test 4: who answers when it breaks?

Back to that Tuesday afternoon.

I needed that endpoint fixed in hours, not by a queue that answers on Thursday. In practice that means being able to message the person who builds the tool. Time lost there is my client's time, and that's business, not a support ticket.

This is the test that quietly favors small vendors. You trade an SLA for a direct line to whoever writes the code. For work that ships weekly, I take that trade.

What I actually use, and where I wouldn't

Full disclosure, since it matters here: I built AEO Copilot, so I'm not a neutral party on the AI visibility part of my stack. What I can tell you is that the 4 tests are why it exists in the shape it does. Write actions in the MCP server rather than read-only. Every result and source readable, so the data can leave. A flat plan across client brands. My own Slack on the support line.

Which is a slightly uncomfortable thing to publish, because it means the failures are mine too. The endpoint that broke in the first paragraph was somebody else's tool, but I've shipped my share of missing fields.

The fastest way to see whether any of it holds up is to run something on a real site rather than read me describe it. The free audit takes a URL and gives you the technical readiness scan I use on client sites. The 21 agent skills that run the rest of my client workflow are open source, so you can read exactly what I hand my agent.

Where I'd pick something else, honestly: if you need 10 or more engines, Rankscale covers 17 or more from $20 a month. If you're buying for an enterprise brand with a procurement process, Profound goes deeper than anything at this price and Scrunch plays the same game. If you track 1 brand and want a clean monthly PDF, Peec and Otterly do that well and none of the per-client math applies to you.

One number that says how young this whole category is. In June I ran a 30-prompt baseline about AEO tools across ChatGPT, Claude, Perplexity and Google AI Overview and logged the 614 citations that came back. The most-cited names were general SEO suites, not AEO products: Semrush 42, Ahrefs 28, then Profound 25, Otterly 11, Peec 10. Nobody has won this market, including me. So whatever you pick this quarter, you can change your mind next quarter.

The tool is the smaller half

A tracker measures. It doesn't decide what to publish, how to structure a page so a model can quote a paragraph of it, which sources to get cited in, or what to put in front of a client at the end of the month.

That part is method, and I wrote mine down: Get Found Everywhere, the agency edition, 18 chapters and 9 client-ready deliverables. Free, because a method nobody reads is worth nothing to me.

The 4 tests are the part I'd steal if I were you. Run them on your analytics tool, your CMS, your rank tracker, your reporting layer. Anything that can't be called by an agent, won't let its data leave, prices per client, or takes 3 days to answer will cost you more than its subscription.

If you want the long version of how these 4 tests played out tool by tool, I wrote that up separately: Best AEO tool for AI-native agencies.