AI Is Quietly Getting Your Brand Wrong. 84% Never Check.
Notes from building BrandGEO, an AI visibility tool for our agency’s clients.
For almost 30 years, being discovered worked one way: someone typed words into a search box, an engine returned ten links, and an entire industry, mine included, fought over the order of those links.
That era is ending faster than almost anyone’s marketing budget admits. Buyers now ask an AI assistant and take the recommendation. McKinsey pegs it at 44% of US consumers already using AI search as their primary source for purchase decisions, and the same research found that only 16% of brands systematically measure what those assistants say about them. Forrester found B2B buyers adopting AI search three times faster than consumers. ChatGPT alone serves hundreds of millions of people a week.
For a decade, you optimized how you’re found. Now you need to optimize how you’re described.
An AI doesn’t return ten links. It returns a verdict: a confident paragraph naming a few brands and characterizing each one. There’s no position two. You’re either in the answer, described accurately and favorably, or you functionally don’t exist for that buyer. And unlike every channel marketing has ever managed, nobody’s dashboard shows what those verdicts say.
This essay is the story of why I built BrandGEO to answer that question and, more usefully for you, everything building it taught me about how AI models decide which brands to recommend. I’ll give away the whole method.
The question nobody could answer
The origin was not visionary. Tools that answer “what do AI assistants say about this brand?” exist. None of them answered it the way I needed for our clients
When I went looking for tooling, the category had a hole in the middle. Free “AI visibility graders” that spit out a score with no methodology: lead magnets, not instruments. Enterprise platforms with real capability, gated behind sales calls and pricing with a comma in it. And AI-tracking checkboxes bolted onto traditional SEO suites: two or three engines, mentioned-yes/no, a feature shipped to close a roadmap gap.
What seemed reasonable to want (separate simple tool, enterprise-depth measurement, self-serve, at a price a small company pays without procurement) I could not find. So I built the tool I’d gone shopping for.
Building it was the real education. To automate the measurement, I had to answer questions most brands never ask: what does “AI visibility” decompose into? What moves it?
Lesson one: your brand now has two reputations
Every model knows you two ways, and the two diverge.
The first is memory. Models are trained on a snapshot of the internet; whatever the internet said about you before the cutoff got baked into the weights. Fast, fluent, and frozen in time. If you pivoted after the cutoff, memory doesn’t know. If your best coverage came after it, memory hasn’t read it. All other online tools use this engine via the fast AI API queries.
The second is live retrieval. When a model searches the web, it composes an answer from what it fetches right now: your current site, current reviews, whatever ranks today. Fresher, but only as good as what the AI’s crawlers can find and parse.
A huge share of buyer conversations run against memory; casual category questions often don’t trigger a search at all. So “we updated the website” fixes at most half of your reputation. The other half thaws only when the next training run ingests a better version of your story into the model’s new version.
The gap between the two answers is the most diagnostic number you can measure, because it tells you whether your problem is history or your current web presence. That’s why every audit my product runs is dual-mode per engine, with the gap reported. But you can see your own gap in twenty minutes: same questions, search off, then search on, note every disagreement.

Lesson two: recommendation is downstream of describability
Ask a model “what is [any brand]?” and watch the texture of the answer. For some brands it’s instant and crisp: category, audience, differentiator. For others it hedges: “appears to be,” “may offer.”
When a model composes a “best tools for X” answer, it reaches for entities it can state things about confidently. A brand it would have to hedge on is a liability, so it quietly doesn’t make the cut.
Designing a scoring rubric forced me to make this more precise. “AI brand visibility monitoring” decomposed into six measurable dimensions, and I’d argue they’re the checklist for any brand, whatever tool you use or don’t:
- Recognition. Does the model know you exist: name, what you do, where, since when.
- Knowledge Depth. Does it understand the product: features, audience, use cases.
- Competitive Context. On category questions, are you in the set, and who are you framed next to?
- Sentiment & Authority. What tone does it use, and does it treat you as a citable source?
- Contextual Recall. When someone describes a problem you solve, without naming brands, do you come up? This is where the gap between “known” and “recommended” lives.
- AI Discoverability. Can machines parse your site at all; is your name unambiguous; are you crawlable by the AI crawlers?
Each dimension failing points to a different fix, which is the entire reason to measure in dimensions instead of vibes.
Low recognition is an authority problem. Low recall with decent recognition is a content problem: nothing associates you with the problem space. Wrong facts are a freshness problem living at whatever sources the models learned them from. Low discoverability is a technical problem you can often fix in days: structured data, readable HTML, an llms.txt file, a robots.txt that isn’t accidentally blocking GPTBot and ClaudeBot.

Lesson three: the electorate is not your website
The most consequential mental-model shift, especially if you come from SEO, is that when a model with web access answers a category question, your homepage barely votes. The model leans on third-party sources: review platforms, comparison articles, community threads, industry publications. Those sources are the electorate; the model mostly counts votes.
You can map the electorate yourself: ask the engines your category’s buyer questions and write down every source they cite. The list is usually shorter than you’d expect, and it converts into the most effective PR targeting I know of.
Every domain sorts into three buckets: cites you favorably (defend), cites competitors but not you (that’s your outreach list), doesn’t cover your segment (ignore). Getting onto the handful of comparison pages the models keep citing will do more for your AI visibility than fifty generic backlinks. This insight became a feature in my product (citation intelligence, with the “cites them, never you” outreach queue), but the manual version is disciplined note-taking, and it works.
Lesson four: only a trend means anything
LLM answers are non-deterministic; a single answer is an anecdote. Models get updated and retrained without notice; a competitor lands one well-placed article and the recommendations tilt. A one-off check tells you where you were on a Tuesday.
So the real instrument is a loop, not a snapshot: a structured question set covering how buyers actually prompt (discovery, comparison, recommendation, sentiment, feature, use-case, plus the exact phrases you know your buyers use, verbatim), across multiple engines, weekly, with mention rates, sentiment, and share of voice against competitors logged over time. Measure → fix → track.
Price the manual version of that loop and it comes to a meaningful chunk of a marketer’s week, forever. That’s why I built it as the monitoring half of the product: weekly runs across all AI engines, competitor share of voice, citation tracking, a report every Monday. Whether you rent that or build your own, only a time series is real in this channel.

Where BrandGEO fits
BrandGEO currently audits how five major AI engines (OpenAI, Anthropic, Gemini, xAI, DeepSeek) see a brand: both modes, six dimensions, a 0–100 score, findings per engine, and a prioritized fix plan. Monitoring runs weekly with competitors and citations. Plans are $79, $149, and $349 a month depending on scale, with white-label reports for agencies on the top tier. Plus a free two-engine check, a 7-day trial without a card, and a 30-day money-back guarantee. If the data isn’t worth it, you’ll know before you’ve risked anything.
Why now, and what to do
There are no ad units in an AI’s answer. In search and social you can always buy your way back in; here you can’t. Visibility is earned through clarity, consistency, and presence in the sources models trust, and it accumulates slowly, because today’s coverage is next model’s training data. That’s what makes waiting expensive.
There’s no trick underneath any of this. Nobody outside the labs knows how the weights get set, and anything sold as a way to game them is a guess. What’s left is ordinary work: find out what the models currently say, fix the specific thing each weak dimension points at, check again next week.
Start with the gap between memory and live retrieval on your own brand. It tells you whether your problem is your history or your website, and it’s the number I built BrandGEO to report first. From there the six dimensions tell you what to fix, and the weekly trend tells you whether it worked.
