The AEO Loop · Learn

How to measure your AEO performance

You can measure how often AI recommends your business, but only once you stop guessing and start reading what the engines actually say. This is a practical method for scoring your AI visibility, setting a baseline, and tracking whether it improves.

Anmol Talwar, founder of The AEO Loop

The blind spot

Most businesses have never looked

Most owners can tell you their Google ranking for a core term, because they have checked it a hundred times. Almost none can tell you what ChatGPT says when a prospective client asks for the best option in their category. AI answers happen inside a private conversation between a person and an engine, so unlike a page of search results you never see them unless you go looking. That gap matters because buyers are already asking these questions, and the engine is already naming someone. If it is not naming you, nothing tells you. Measuring AEO starts with the uncomfortable step of opening the engines yourself and reading what they say about your category.

Three outcomes

Recommended, mentioned, or excluded

Every answer puts you in one of three states, and scoring each one is how a vague sense of visibility becomes something you can count. Recommended means the engine names you as an answer to the question, near the front, as one of the options a buyer should consider. Mentioned means you appear somewhere in the response but not as a headline choice, maybe in a longer list or a passing reference. Excluded means the engine answered the question in full and never named you, while naming competitors. This scale is coarse on purpose, because it maps to what a buyer actually experiences. Being fifth on a list of ten is closer to invisible than to chosen, and reading each answer against these three states gives you a number instead of an impression.

The method

Run real buyer questions against live engines

The measurement is only as good as the questions, so use the ones a real buyer would type, not the ones that flatter you. A plastic surgery practice should test something like the best rhinoplasty surgeon in its city, not whether its own name is good. Run each question against the engines your buyers actually use, which today means ChatGPT, Gemini, Perplexity, and Claude, plus Google's AI Overviews. Keep the wording, the location, and the specialty consistent so the results stay comparable over time. Record the outcome for each question on each engine as Recommended, Mentioned, or Excluded, and note which competitors got named instead. Read the full answers, not just the names, because the language an engine uses about your category tells you as much as who it chooses to drop into it.

Sample, not snapshot

Why one answer proves nothing

Ask the same engine the same question twice and you can get two different answers, with different names and different phrasing. That variation is built in, not a glitch to wait out. These models sample their output and pull from a live set of pages that shifts through the day, so no two runs are guaranteed to match. A single screenshot showing you recommended is as unreliable as one showing you excluded. The honest way to measure is to sample, running each question several times across engines and reading the pattern rather than any one result. What you are really measuring is a probability, how often you get named across many answers, and that is the number that moves when the work is working. No one can promise a specific answer on a specific day, and any measurement that pretends otherwise is selling you a snapshot.

The baseline

What a baseline actually looks like

A baseline is the first honest read, taken before you change anything. In practice it is a table with your buyer questions down one side and the engines across the top, each cell marked Recommended, Mentioned, or Excluded based on several samples rather than one. From that grid you can pull a single share figure, the percentage of answers where you were named at all, and a recommended share for the answers where you led. Most businesses that have never done AEO work start lower than they expect, often excluded from the majority of answers in their own category, which is confronting but useful to know. The baseline also captures who is winning your questions right now, since those competitors are the ones the engines currently trust. Everything you do afterward gets measured against that starting line.

Month over month

Track the trend, not the day

One reading is a baseline, and the measurement is the repeat. Re-run the same questions on the same engines at a set interval, with monthly being a sensible default, keeping the wording identical so you are comparing like for like. Watch the share of answers where you are named and the share where you lead, and expect both to move unevenly, because engines update, competitors publish, and retrieval shifts underneath you. A single month that dips is usually noise, while a direction that holds across three or four months is signal. Log what you changed between readings, the pages, the citations, the profile fixes, so you can connect the work to the movement instead of guessing at it. Over time this turns AEO into a loop you can run on purpose, seeing where you stand, changing something, and measuring whether it landed.

Definitions

Key terms.

Recommended

An outcome where an AI engine names your business as one of the answers to a buyer's question, near the front of its response, among the options worth considering.

Mentioned

An outcome where your business appears somewhere in the answer but not as a leading choice, often in a longer list or as a passing reference.

Excluded

An outcome where the engine answers the question in full and never names your business, while typically naming competitors instead.

Sampling

Running the same question several times across engines and reading the pattern of outcomes, rather than trusting one answer, because AI responses vary from run to run.

Questions

Common questions.

How do I check what AI says about my business?

Open the engines your buyers use, ChatGPT, Gemini, Perplexity, Claude, and Google's AI Overviews, and ask the questions a buyer would ask, such as the best provider for a service in your city. Read the full answer and note whether you are named and whether competitors are named instead. Run each question a few times, since answers vary between runs, and record the pattern rather than a single result.

How often should I measure AEO performance?

Monthly is a sensible default for most businesses. It is frequent enough to catch real movement and spaced enough that engine updates and your own changes have time to show up. Keep the questions and wording identical each time so the readings stay comparable, and treat a single dip as noise until a direction holds across several months.

Why does AI give a different answer every time I ask?

Because these answers are non-deterministic. The model samples its output and pulls from a live set of web pages that changes through the day, so the same question can return different names and phrasing on repeat runs. This is why you sample across many answers instead of trusting one, and why no one can honestly promise a specific answer on a specific day.

What does a good AEO baseline look like?

A baseline is a grid of your real buyer questions against each engine, with every cell marked Recommended, Mentioned, or Excluded based on several samples. From it you read one share figure for how often you are named at all, and a recommended share for how often you lead. Many businesses new to AEO start excluded from most answers in their category, which is normal and gives you a clear line to measure against.

Can I measure AEO myself, or do I need a tool?

You can do it by hand by opening each engine, asking your buyer questions, and logging Recommended, Mentioned, or Excluded in a spreadsheet. A scanner speeds up the sampling and keeps the questions consistent, which matters most when you repeat the read every month. Either way the discipline is the same, real buyer questions, several samples each, and the same wording over time.

See it for your business

Find out what AI says about you.

Run the free scanner — real queries against four live engines — and see whether you are recommended, mentioned, or excluded.