Methodology

How Mentionaly measures AI visibility

AI answers are non-deterministic — the same question can return different brands on different runs. Our entire methodology is built around making the measurement trustworthy anyway.

1. Prompt generation

When you enter a domain, we analyze the site to identify the business, its category, offerings, and competitors. From this we generate the buying questions real customers ask AI assistants in that category — comparison questions ("best X for Y"), alternative questions ("alternatives to Competitor"), use-case questions, and branded questions about your company directly. On paid plans, every prompt is reviewable and editable, and you can add your own.

2. Engines and access

We query each AI engine programmatically with web access enabled, so answers reflect the retrieval-augmented behavior real users experience. Engine coverage currently includes ChatGPT, Gemini, Claude, Perplexity and DeepSeek. Consumer chat products and their underlying APIs can differ in behavior (memory, personalization, model routing); we measure the closest reproducible equivalent and apply the same method consistently across every run, so trends over time are apples-to-apples.

3. Sampling

Every prompt runs multiple times per engine — 5 samples by default. A brand's result on a prompt is never "mentioned or not" from a single answer; it's a frequency across samples (e.g. mentioned in 4 of 5). This is the difference between measurement and anecdote, and it's why Mentionaly scores stay stable week to week unless something real changed.

4. Scoring

Your AI Visibility Score is a 0–100 composite of three sub-scores, computed across all prompts, samples, and engines:

score = w1 * mention_rate + w2 * position_score + w3 * citation_presence

mention_rate — % of samples in which your brand appears
position_score — average rank when mentioned, weighted toward early mentions
citation_presence — % of answers citing your website as a source

Competitor share of voice falls out of the same data: every brand detected in the sampled answers is scored on the identical basis, so comparisons are fair by construction.

5. Sources and citation gaps

For every answer, we extract the sources the engine cited. Aggregated across your category's prompts, this reveals which domains actually shape AI answers in your space — and which of them don't mention you. Gaps are ranked by citation frequency, producing a fix list ordered by expected impact.

6. Change tracking

Runs repeat on a weekly cadence (daily available as an add-on). Because each run uses the same prompts, sampling depth, and engines, score movement reflects genuine changes in how AI represents your brand — not sampling noise.

Honest limitations

  • Non-determinism is reduced, not eliminated. Sampling narrows the confidence interval; small week-to-week wiggles within it are expected.
  • Personalization is out of scope. Engines may tailor answers to individual users' history and location; we measure the neutral baseline.
  • Engines evolve. Model updates can shift answers for everyone at once. We flag runs that follow known model releases so you can read trends in context.

We publish this page because measurement tools deserve scrutiny. If you have questions about the methodology, we're happy to answer them.

See the methodology in action

Run a free report on your own domain and inspect every sampled answer.

Run a free report