We Asked Four AI Engines to Recommend a Business: The Protocol (and How to Run It)
Same question, four engines, repeated asks: who gets named, what gets cited, and does the answer hold still. That is the study worth running. Until a dated wave is published with sample size and method, use this protocol so you are not arguing from vibes.

Article Summary
- Almost everyone speculates about how answer engines choose. Few publish a repeatable test protocol.
- I use this multi-engine method with clients: fixed queries, named engines, honest coding, and a clear way to read results.
- Track concentration, overlap across engines, citation source categories, and stability across repetitions.
- Payoff: traits shared by consistently named businesses, and negative findings that did not correlate.
- You can run a category version in an afternoon and feed it into the 40-point checklist .
Method (run this before you argue)
- Engines: pick a fixed set (for example ChatGPT with search, Gemini, Perplexity, Copilot) and record versions/modes
- Queries: 10-25 buyer questions for one category and geography
- Repetitions: ask each query multiple times (stability check)
- Coding: business named, URL cited, source type (brand site, directory, reviews, editorial, forum, marketplace)
- Controls: same account settings where possible, same day window, note location assumptions
| What to measure | Why it matters |
|---|---|
| What to measureConcentration | Why it mattersDo a few brands dominate recommendations? |
| What to measureEngine overlap | Why it mattersWho appears across systems? |
| What to measureCitation sources | Why it mattersWhat evidence engines lean on? |
| What to measureStability | Why it mattersDoes the answer thrash between runs? |
| What to measureCommon traits | Why it mattersWhat do winners share that losers lack? |
Citation source categories
- Brand's own site
- Directories and profiles
- Review platforms
- Editorial / news
- Forums and communities
- Marketplaces
PRO TIP Rank them in your spreadsheet. Engine differences are the useful part. Tie findings to citations, reviews, and earned media.
PRO TIP Publish negative findings too ("ad spend did not predict inclusion in this sample"). That is what makes a study trustworthy and citable.
Run it for your category in an afternoon
- Write 15 queries a real buyer would ask
- Create a sheet with columns for engine, run, names, URLs, source types
- Execute two runs per query per engine
- Score overlap and source mix
- List traits of businesses named more than once
- Map gaps to the 40-point checklist and ship fixes
Putting it together
As of now, opinions about who AI recommends are pretty vague. Just because you see a result does not mean someone else sees the same result. It is not like classic SEO where somebody could reliably push out a PDF that says you are ranked number three for a keyword in the USA.
There is a lot more to it now. The answer engine has a context file for the person asking, and based on that context it will make recommendations you cannot reliably predict. Plenty of companies say they can predict it. The honest assessment is that it is still a best guess.
That is why I want a protocol, not a demo. Same question, four engines, repeated asks, written down. Re-run it after you change the site or citations. We will help you make the best guess honestly, because ultimately this is not about placement or impressions. It is about the end result: new clients, new sales, and new referrals.
Talk with Xeal about a recommendation test
Tell me your category and the hiring question that actually matters. We will run the protocol with you and turn the gaps into work that can produce real revenue.
FAQ
Which engines should I include?
Start with the four your buyers actually use. Keep the set fixed for the wave.
How often should I re-run?
Quarterly for competitive categories, or after major site/entity changes.
Does location change results?
Often yes. Record location context in the method notes.
Can I do this without paid tools?
Yes. Manual runs work. Spreadsheets beat vibes.
Will you run it for my industry?
Yes, as a scoped study. That is the CTA.
Where is the 500-answer dataset?
When a large dated study wave is published, Xeal will link it here with method and limits. Until then, run the protocol yourself and keep your scores.
Sources
What operators say after working with Xeal
Quotes from people who hired Tony and Xeal for websites, SEO, publicity, and strategy.
Tony is truly a visionary. He's always coming up with new ideas for his company and his clients. As an entrepreneur, Tony is excellent at business consultation and speaking on the radio.
Tony has loads of experience when it comes to internet marketing. He has tested what does and doesn't work and is always up to date on the latest techniques that are working online. A true veteran in the industry.
Want help putting this into production?
If this article named a problem you already feel, tell me what is stuck. We will turn the playbook into systems your team can run.
