AI Marketing

How to Track AI Mentions of Your Brand (Tools + Method)

A statistically defensible method for measuring where your brand shows up in ChatGPT, Gemini, Claude, and Perplexity, plus the tools that automate it.

August 17, 2026By Ben LaRue8 min read
GEOAI SearchAI VisibilityMeasurementTools

Tracking AI mentions of your brand means running a fixed set of buyer questions through ChatGPT, Gemini, Claude, and Perplexity many times each, logging who gets named, and calculating a mention rate. A single answer to a single prompt proves close to nothing: ask the same question twice and you rarely get the same list of brands back, so the entire practice depends on how many times you ask, not just what you ask.

Key Takeaways

  • A single AI answer is close to worthless as evidence: SparkToro found under a 1-in-100 chance that ChatGPT or Google's AI returns the same brand list twice for the same prompt, across 2,961 real responses.
  • For a defensible +/-10 percentage point margin of error, that works out to roughly 100 runs per prompt, per engine, per market, not a screenshot or two.
  • Dedicated tracking tools range from $29/month (4 engines, 15 prompts) to $800+/month, and none of the self-serve tiers cover every major engine at the entry price.
  • You can build a statistically honest baseline by hand with a spreadsheet and zero software budget. It costs time, not money, roughly 100 prompts to run per cycle for one properly measured question.
  • Brands that skip measurement don't find out they've lost the AI answer until a client asks why a competitor got named instead.

Definition: AI Mention Tracking

AI mention tracking is the practice of repeatedly running a fixed panel of buyer questions through generative engines like ChatGPT, Gemini, Claude, and Perplexity, then logging whether, how, and how consistently a brand is named. It differs from keyword rank tracking in one fundamental way: a Google rank is a stable position you can check once and trust, while an AI answer is a probability distribution you have to sample repeatedly before you can trust it at all. This is a companion piece to how to measure your brand's visibility in AI search and focuses specifically on the tools and the mechanics of running the measurement.

Why is tracking AI mentions harder than tracking a Google ranking?

Rand Fishkin's team at SparkToro put a number on the instability marketers keep discovering the hard way. Working with Gumshoe.ai, they recruited 600 volunteers to run 12 prompts through ChatGPT, Claude, and Google's AI Overviews and AI Mode, a combined 2,961 times, then compared the brand lists each response returned (SparkToro, January 2026). The odds that two runs of the identical prompt named the identical set of brands: under 1 in 100 for ChatGPT and Google's AI. The odds they named those brands in the identical order: roughly 1 in 1,000. In a later session unpacking the data, Fishkin illustrated just how far that stretches on some engines: Google AI Mode needed on the order of 124 runs before the same two-brand list was even likely to repeat once (SparkToro Office Hours, May 2026).

That is the whole problem with a spot check. Ask ChatGPT once whether it recommends you, get an answer, and you have learned almost nothing about whether it recommends you, only what one draw from a wide probability distribution happened to return that day.

So how many times do you actually need to ask?

Treat "did the brand appear" as a coin-flip-style estimate and the math is standard survey statistics. For a 95% confidence interval at the hardest case (p = 0.5), the number of independent runs needed to hit a given margin of error looks like this:

| Margin of error | Runs needed per prompt, per engine, per market | |---|---| | +/-15 percentage points | 43 | | +/-10 percentage points | 97 | | +/-7.5 percentage points | 171 | | +/-5 percentage points | 385 |

Bar chart showing the number of independent AI prompt runs needed for a given statistical margin of error: 43 runs for plus or minus 15 points, 97 for plus or minus 10 points, 171 for plus or minus 7.5 points, 385 for plus or minus 5 points

Ninety-seven runs for a +/-10-point margin rounds cleanly to the "about 100" standard that shows up independently across the industry. Evertune, one of the enterprise tracking platforms, samples each prompt exactly 100 times per model as its stated methodology (Evertune, 2026). Search Engine Journal's practical baseline for teams without that budget recommends 20 to 30 prompts run at least twice each, monthly (Search Engine Journal, August 2026), which is honest about what it is: a directional audit that can catch a glaring absence, not a number precise enough to defend in a board meeting.

Talk to Etched

Which tools actually track this, and what do they cost?

A real category of software now automates the repeated prompting, and pricing scales almost entirely on two levers: how many prompts you track and how many engines you cover, not on any single "best" feature.

| Tool | Entry price | What the entry tier covers | Best for | |---|---|---|---| | Otterly | $29/mo | 15 prompts, 4 base engines (ChatGPT, Google AI Overviews, Perplexity, Copilot); Gemini, AI Mode, and Claude cost extra | Smallest budget, single brand | | Peec AI | $95/mo | 50 prompts, pick 3 engines, daily tracking, unlimited seats | Agencies wanting flexible engine choice | | Semrush AI Visibility Toolkit | $99/mo | ChatGPT, Google AI Overviews/AI Mode, Gemini, Perplexity | Teams already inside Semrush | | Profound | $99/mo, billed annually | ChatGPT only at Starter; 3 engines at the $399/mo Growth tier, also billed annually | Buyers who expect to upgrade fast | | Ahrefs Brand Radar | $199/mo per platform, on top of a base Ahrefs plan | One AI platform's pre-collected response index plus custom prompts | Teams already inside Ahrefs | | Trafiq | $299/mo | ChatGPT, Gemini, Perplexity, Claude (no Google AI Overviews or Copilot); mention frequency, sentiment, and share of voice bundled with existing SEO and paid-media reporting | Teams wanting AI visibility alongside SEO and paid media in one dashboard | | Evertune | $800/mo | 11 engines, 100 runs per prompt built into the methodology | Enterprise, paying for rigor over price |

Bar chart of entry-tier monthly price for six self-serve AI-visibility tools: Otterly $29, Peec AI $95, Semrush AI Visibility Toolkit $99, Profound $99, Ahrefs Brand Radar $199, Trafiq $299

Prices above are public list prices as of August 2026 and move often in this category; verify against the vendor's own pricing page before buying. Two entries need a footnote the chart can't carry: Profound's $99 and $399 are annualized rates billed yearly, not month-to-month, and Ahrefs Brand Radar's $199 is an add-on that requires an active base Ahrefs subscription underneath it, so it is not really a standalone $199 entry point. Read the fine print before comparing sticker prices generally, too. Otterly's headline $29 covers four engines at only 15 prompts; Profound's Starter tier covers exactly one engine, ChatGPT, until you upgrade to Growth for three. "Cheapest" and "most engines covered" are not the same axis, and no self-serve tier in this list covers every major engine at its entry price. For how these tool categories trade off structurally, independent of any single brand name, see how to calculate AI share of voice.

Key takeaway: Tool pricing scales on prompts tracked and engines covered, not on a single quality tier. Compare what's actually included at the price you'd pay, not the lowest number printed on the pricing page.

What should you log if you're not buying a tool yet?

The manual version works. It just costs time instead of a subscription, and it needs the same discipline a paid tool builds in by default: a fixed prompt list, clean sessions, and consistent logging. Build a panel of 15 to 30 real buyer questions, "best [category] for [use case]," "[competitor] alternatives," "who should I hire for X," covered in more detail in how to measure your brand's visibility in AI search, then run each one through ChatGPT, Gemini, Claude, and Perplexity from a logged-out or fresh session so personalization doesn't skew the answer.

For every run, log at minimum:

  • Prompt (exact wording) and engine
  • Date and run number within that prompt/engine pair
  • Whether your brand was mentioned at all (yes/no)
  • Whether it was actually recommended, not just named in passing
  • Which competitors appeared, and in what order
  • Whether any of your own pages were cited as a source
  • A short sentiment read: positive, neutral, or negative

Report every number as a fraction, not a percentage floating on its own: "12 mentions out of 100 valid runs," not "12% visible." The denominator is what makes the number auditable six months later when someone asks how it was calculated.

When does the manual method stop being worth it?

Run the arithmetic on your own program and the ceiling becomes obvious fast. Twenty-five prompts across four engines at 100 runs each is 10,000 queries in a single measurement cycle, and a trend line needs that repeated monthly. That's a research operation, not a task you bolt onto someone's existing job. A tool earns its subscription the moment the manual version would consume more hours than it costs in software.

What a tool buys you is scheduling, not a different number: the same prompt panel, run on a cadence, logged automatically, with sentiment and citation tagging done for you instead of by hand. It will not tell you why a competitor is winning the answer instead of you. The entity and content gaps that usually explain it are worth reading before you shop for software, because a subscription that tells you the score every month without changing it is just a more expensive way to watch yourself lose.

Key takeaway: The manual method is free but doesn't scale past a handful of prompts measured occasionally. A tool is worth paying for once the measurement itself becomes a monthly job, not because it knows something a spreadsheet can't.

Frequently Asked Questions

How many times do I need to run the same AI prompt before the results mean anything?

For a defensible +/-10 percentage point margin of error, plan on roughly 100 independent runs per prompt, per engine, per market. SparkToro found under a 1-in-100 chance that ChatGPT or Google's AI returns the same brand list twice for the same prompt, so a single answer is a sample of one, not a measurement.

What's the difference between a rank tracker and an AI-visibility monitor?

A rank tracker checks a stable position that barely changes between checks. An AI-visibility monitor repeatedly samples a probability distribution, because the same prompt asked twice on the same engine rarely returns the same answer. That is why AI mention tracking requires many runs per prompt where rank tracking only needs one.

Which AI-visibility tool should I start with?

It depends on budget and which engines matter most to your buyers. Otterly ($29/mo) is the cheapest entry point but covers only 4 engines at 15 prompts. Peec AI and Semrush's AI Visibility Toolkit sit around $95 to $99/mo with broader engine choice. None of the self-serve tiers cover every major engine, so check what's included before comparing sticker prices.

Can I track AI mentions without paying for a tool?

Yes. A fixed prompt panel, logged-out browser sessions, and a spreadsheet will produce a statistically honest baseline. It costs time instead of a subscription: roughly 100 runs per prompt per engine for one properly measured question. It stops being practical once you need to track more than a handful of prompts on a recurring monthly cadence.

How often should I re-run my AI visibility measurement?

Monthly, at minimum, using the same prompt wording, engines, and session conditions each cycle. Models update and competitors publish new content between cycles, so a single measurement is a snapshot; only a repeated, consistent cadence turns it into a trend you can act on.

How Etched tracks AI mentions for clients

We run this exact discipline, a fixed buyer-question panel, sampled at a defensible run count, across ChatGPT, Gemini, Claude, and Perplexity, then pair the numbers with the fix: closing the entity, schema, and content gaps the measurement points to.

Want a baseline without building the spreadsheet yourself? Run the free AI visibility check for an instant snapshot, or book the $5,500 GEO Audit for the full prompt-panel measurement, done properly.

Run the free AI visibility check