Key Takeaways
- A case study screenshot is not proof. Anyone can stage one in a minute, and AI answers change between sessions anyway. Re-run it yourself.
- Three queries settle it: the unbranded category query, the branded query, and the competitor alternatives query. Ten minutes total.
- A real citation names the business inside the answer body with a reason attached, not just in the source links at the bottom.
- Run each query three times in fresh chats with memory off. Two clean appearances out of three beats one perfect screenshot.
- If an agency will not name a single client you can check, that is the finding. Named references are the cheapest thing they can hand you.
Why a Screenshot Proves Nothing
Every AI search pitch deck looks the same. A ChatGPT answer, the client's name highlighted in yellow, a caption saying something like cited on page one of AI. It looks like evidence. It is not.
Staging one takes about ninety seconds. Turn on custom instructions telling the model you are researching that specific company. Or leave memory on after a month of looking at their website. Or write the prompt so it already contains the business name, the city, and the exact specialty, which leaves the model very little room to answer with anyone else. Nothing in that screenshot tells you which of those happened.
Even honest screenshots are weak. AI answers are not deterministic. The same question asked twice returns different businesses, different orderings, and different source lists. Account state, location, and model version all move the result. So a screenshot captures one screen, one moment, one account, usually with no date attached.
The good news is that this cuts both ways. If a citation is real, it survives your own testing. You do not need the agency's cooperation, a tool subscription, or a technical background. You need the client's name, their city, and ten minutes.
The Three Queries to Type
Set up first, because the setup is what makes the test honest. Open ChatGPT in a private browser window or a logged out session. If you are logged in, turn memory and personalization off in settings. Start a brand new chat for every single run. Do not reuse a thread.
1. The category query
Ask what an actual customer would ask, with no business name in it. Who should I call for emergency water damage restoration in Tucson. Best pediatric dentist in Scottsdale. My AC died and it is 104 degrees, who do I call in Mesa.
This is the query that matters, and it is the one most decks skip. Run it three times in three fresh chats. Does the agency's client appear without you naming them? That is the entire test.
2. The branded query
Now name them. Is this company any good, and what are they known for. You are checking whether the model has substantive, specific material about the business or is just paraphrasing the homepage. Specific service lines, neighbourhoods served, and review themes suggest real work. Generic filler suggests a thin footprint.
3. The competitor query
Name a known competitor in the same city and ask for alternatives to them. If the agency's client is genuinely well cited, it turns up here without being prompted. If it never appears as an alternative, the result from query one was probably thinner than it looked. This is also the query most agencies have never run, which is worth knowing on its own.

What a Real Citation Looks Like
Strong looks like this. The business is named in the answer body, inside the recommendation itself, with a reason attached: known for 24 hour emergency response, or frequently mentioned for same week appointments. It appears across separate sessions with slightly different wording. It survives a plain, lazy prompt.
Weak looks like this. The name only shows up in the source link tray at the bottom. Buyers read the answer, not the footnotes, so a link is not a recommendation. Or it appears only once you have stacked three qualifiers into the prompt. Or it shows once in five runs, which is noise.
Write down what you see. Query, run number, appeared yes or no, and where in the answer. Five minutes of notes beats a memory of roughly seeing them a couple of times, and it gives you something concrete to put in front of the agency. If you want the wider framing on what to expect from month one, we covered it in what should be visible by day 30.
When They Will Not Name a Client
Every case study anonymized. A dental practice in the Southwest. A restoration company in the Midwest. Ask why, and the answer is usually NDAs.
In local services, an NDA covering the fact that a relationship exists is unusual. Agencies publish client logos constantly. Anonymizing is almost always a choice, and there are two reasons for it. One is legitimate: a client asked not to be listed, which happens. The other is that the result does not hold up when someone checks.
So do not argue about NDAs. Offer a smaller ask instead. Two named clients you can run a query against right now, plus one phone number for a reference call. That request is nearly free for an honest agency and impossible for a dishonest one. It also tells you whether they are comfortable being audited, which matters more over twelve months than any deck.
If you have already been burned once, the sequencing in setting up the next agency so you can fire it with evidence is the companion piece to this one.
The 10-Minute Checklist
| Minute | Do this | Pass looks like |
|---|---|---|
| 0 to 1 | Private window, memory and personalization off | Clean session, no saved context |
| 1 to 5 | Category query, three fresh chats | Named in the answer body twice out of three |
| 5 to 7 | Branded query, one chat | Specific services and review themes, not homepage filler |
| 7 to 9 | Competitor alternatives query | Client surfaces unprompted as an option |
| 9 to 10 | Repeat for the second named client | Pattern holds across two businesses, not one |

Sign, Push Back, or Walk
Sign if at least two of the named clients hold up on the unbranded category query across three runs, the agency handed over names without friction, and the scope is written down before you pay. Agree the same three queries as your own baseline on day one. Month to month terms matter here too, which we broke down in retainer terms explained.
Push back if results are real but inconsistent, say one client passes and one does not. That is normal and not disqualifying. Ask for the lead log behind the passing client: date, source, what the caller wanted, whether it became a job. If they can produce that, the inconsistency is competitive difficulty, not fiction. If they cannot, you have found the real gap.
Walk if they will not name a single client after you offered the smaller ask, or the named clients never appear on an unbranded query across multiple runs, or the entire proof set is screenshots with no lead data behind it. None of those are close calls. Before restarting the search, check whether switching is worth it or just resets the clock, and use the one-page scope on whoever comes next.
Run the check on us first
Fair is fair. Ask us for two client names on the call and run the three queries while we are still talking. If they do not hold up, you have lost twenty minutes and gained a method you can use on every agency after us.
Book a call and audit us liveFrequently Asked Questions
How do I verify an AI search agency's case study without asking them for anything?
You need one name and one city. Open a fresh ChatGPT chat with memory and personalization turned off, then ask the category question a customer would ask: who should I call for emergency water damage in Tucson. Do not name the business. Run it three times in three separate chats. If the client from the case study appears in at least two of those three answers, the citation is real. If it only appears when you type the business name yourself, you have proved the model can read a website, which is not the same thing as being recommended.
What counts as a real AI citation versus a weak one?
A real citation names the business inside the body of the answer, as part of the recommendation, usually with a reason attached like known for same day emergency response. It repeats across separate sessions. A weak citation shows up only in the source link list at the bottom, or only after you stack three qualifiers into the prompt to force it, or once out of five runs. Source links are a footnote. Buyers read the answer, not the footnote. If the agency's proof is a link in a citation tray, ask them to show you the same business inside the answer text.
Why do I get a different answer than the agency's screenshot shows?
Because AI answers are not deterministic and are not identical for everyone. The same prompt returns different results across sessions, accounts, locations, and model versions. Custom instructions and saved memory change output. A logged in account that has been researching one company for a month gets skewed results. So a screenshot tells you what one screen looked like at one moment on one account, with no date and no account state attached. That is not evidence. Your own repeat runs are evidence, which is why running it yourself takes ten minutes and settles the question better than any deck can.
What if the agency says they cannot name clients because of NDAs?
In local services, a signed NDA covering the mere existence of a relationship is rare. Agencies name clients on their own websites constantly. If every case study is anonymized as a dental practice in the Southwest, treat that as a choice, not a legal constraint. Ask for a compromise instead of accepting the blanket no: two named clients you can check yourself with a search query, plus one phone number you can call. A confident agency hands that over in a message. If the answer is still no after you offer the compromise, you have your answer about what the case studies are worth.
How much should AI search optimization cost, and does price tell me anything about legitimacy?
Full service retainers for local businesses generally run $800 to $3,000 per month, with setup fees of $500 to $1,500 where they exist. Monitoring only tools sit around $50 to $300 per month. Price alone tells you very little. A $3,000 retainer with unverifiable case studies is worse than an $800 retainer with two named clients you checked yourself in ten minutes. What price should track is scope: content production, citation and directory work, schema, and a lead log you can read. Ask what changes at each tier. If nothing concrete changes, the tiers are packaging.
How many test runs are enough before I trust the result?
Three runs per query, each in a brand new chat, is the working minimum. Five is better if the decision is expensive. Vary the wording slightly between runs the way a real customer would, because no two customers type the same sentence. Record what you see rather than trusting your memory of it. Two clean appearances out of three runs is a strong signal. One out of five is noise, and an agency presenting that single run as a case study is either not measuring or hoping you will not check. Do the same for every named client they give you.
What should I ask for instead of a case study deck?
Ask for a lead log. Names redacted is fine, but you want date, source, what the person asked for, and whether it turned into a job. That single artifact answers the question a deck avoids: did the phone ring and did work come from it. Then ask for the query they optimized for so you can run it yourself. Then ask for one reference call. A deck of screenshots costs an agency an afternoon in a design tool. A lead log costs them nothing if the work is real and is impossible to produce if it is not.
How soon can I run this same check on my own business after signing?
Immediately, and you should. Run the three queries before any work starts and save the answers with the date. That baseline is the only thing that makes month two meaningful, and almost nobody captures it. Then rerun monthly using the same prompts and the same conditions, memory off, fresh chat. Expect movement on branded queries within about four to six weeks and on unbranded category queries in roughly two to four months, depending on how much competition sits in that city. If nothing shifts on either by month three, you have a specific, checkable question to raise.
The reason this works is that AI search is the first channel where the buyer can audit the vendor without the vendor's permission. You cannot check a rankings screenshot. You can check whether a business gets recommended when you ask.
Bring your own three queries to the call. Book a slot here and we will run them together on your business before anyone talks about price.

About the author
Matthew Johnson is the founder of Pleiades Consultancy. He previously scaled his own marketing agency to multiple six figures before serving as CMO of an Amazon agency, where the client base tripled from 15 to 45 active clients during his tenure. He worked with some of the largest names in e-commerce, including Ridge Wallet, HexClad, BK Beauty, The Woobles, Walkize, Lonely Planet, and Obvi. He now works with local businesses to maximize their client acquisition and visibility through AI search with ChatGPT, Claude, Gemini, Perplexity, and Bing Copilot.
