Skip to content
Rank Flywheel
Client

How to Run an AEO Audit on a Client Page

Rank Flywheel Published July 31, 2026

Run a defensible AEO audit on one client page. Separate search and AI-answer evidence, record mentions and citations, and prioritize testable changes.

What an AEO audit on a single page actually covers

An answer engine optimization (AEO) audit, as used here, is a dated assessment of how one specific URL performs as a candidate source for AI-generated answers to a defined set of questions. It is a diagnostic of conditions observed at a point in time, not a forecast of how assistants will behave next week.

The unit of analysis matters more than the label. Every audit of this type fixes five things before any evidence is collected:

  • One URL, chosen because it is the page that should be answering the questions
  • One market and location context
  • One measurement period, recorded as a date
  • One defined set of traditional-search queries
  • One corresponding set of natural-language AI questions

Several things are explicitly out of scope, and saying so early prevents an awkward conversation later:

  • Site-wide crawling and technical health across the whole domain
  • Claims about coverage across every assistant or answer surface
  • Continuous monitoring or alerting
  • Any projection of future citation behavior or ranking movement

Traditional search evidence and AI-answer evidence are separate measurement surfaces. They use different inputs, produce different evidence and must be collected and reported separately. Blending them into a single figure reduces the diagnostic value of both.

Step one: scope the page and agree on the inputs before you gather evidence

A common way an AEO audit fails is testing inputs the client never cared about, then presenting findings the client cannot act on.

Start by confirming, in writing, what the page is supposed to do commercially and which buyer or research needs it should address. A service page, a comparison page and a definitional article carry different jobs, and the inputs should reflect that job.

Define two related but separate sets:

  • Traditional-search queries, written as someone would enter them in a search engine
  • Natural-language AI questions, written as someone would ask an assistant

Map each search query to the corresponding AI question when they address the same intent, but do not force them to use identical wording. A commercial search query such as “commercial HVAC maintenance contract” and a question such as “How do I choose a commercial HVAC maintenance provider?” can represent the same need while remaining valid inputs for different surfaces.

Record every query and question verbatim, along with the assumed market and location. Keep both sets small enough that every input receives a complete evidence record rather than a partial pass.

Question fit is itself a finding. A page can be well written, technically clean and still be a poor match for the questions the client wants to win. When that happens, the audit output is a scoping recommendation, not a list of on-page tweaks.

Close the scoping step by setting the expectation in writing: findings apply only to the exact queries, questions, market, assistant and conditions tested.

Step two: collect traditional search evidence for the URL

Traditional search evidence is the conventional baseline that makes the AI-answer findings interpretable. Collect it first, and keep it in its own part of the record.

For each agreed search query:

  • Capture whether and where the URL is found in traditional search results
  • Record not-found results explicitly, because a not-found result is evidence, not a blank cell
  • Log the observation date and the market context beside every result
  • Note which other pages on the client domain compete for the same query or underlying intent, since overlapping pages change which URL is even a candidate source

Cannibalization is worth flagging in plain terms. If two client pages both partially answer a question, the assistant and the search engine may each favor a different one, and your recommendations need to name the page that should own the answer.

Do not treat search position as a proxy for AI-answer inclusion. A page can hold a strong search position for a query and still be absent from a sampled AI answer to the corresponding question. Search visibility and AI-answer visibility are separate measurements, so neither result can stand in for the other.

Step three: collect AI-answer evidence as a separate record

The AI surface needs its own evidence-capture procedure, with its own fields and its own caveats.

For each agreed question, record four distinct things:

  • Mention: whether the brand or page was referred to in the answer
  • Citation: whether the specific URL under audit was cited
  • Sources: which source URLs the answer displayed or cited
  • Unavailable: whether evidence could not be retrieved for that question

Alongside every observation, log the assistant tested, the exact question wording, the market context and the timestamp. Those fields allow another reviewer to repeat the procedure and assess the classification, even though the resulting answer may differ.

A Standard report uses one AI sample per question. One valid sample is real evidence and enough to open a diagnosis, but the AI-answer finding remains preliminary rather than settled. Label it that way in the deliverable.

Failed or unavailable evidence is reported as unavailable. It is never counted as absence, never averaged away and never replaced with an estimate. That single discipline protects the credibility of everything else in the record.

Keep the mechanics of the check itself brief in the deliverable. If the client needs the underlying method, link to how Rank Flywheel measures AI visibility with defensible evidence rather than repeating it in full.

Step four: review the page for question fit and extractability

With evidence in hand, move to on-page diagnosis. The goal is to describe conditions that could plausibly affect whether an answer can be lifted from the page, without asserting that any one condition caused a missing citation.

Question fit and answer placement:

  • Does the page answer each tested question directly, and early, or is the answer buried inside narrative sections
  • Is the answer stated in a self-contained passage, or does it depend on several paragraphs of context
  • Does the page answer the question at the level of specificity the question implies

Structure and extractability:

  • Heading structure that maps to real questions rather than marketing labels
  • Answer-shaped passages, definitions, bulleted sets and tables that isolate a discrete answer
  • Consistent terminology, so the same concept is not named three different ways on one page

Entity clarity:

  • Is it obvious who the business is, what it sells, where it operates and who it serves
  • Are the claims on the page supported by something a source could rely on

Access and rendering:

  • Can the URL be crawled and rendered
  • Is the substantive answer content present in accessible page output, or is it loaded only after an interaction or script execution

Write each structural observation as a hypothesis to test, phrased that way in the deliverable. The honest formulation is that the answer is not currently extractable in a discrete form, and that changing this is a testable action, not that this is the reason the page was not cited.

Step five: examine which sources the sampled answers displayed instead

When the client page was not cited, first determine whether the sampled answer displayed any source URLs. Cataloging the sources that were displayed can turn the observation into diagnostic signal. When no source was displayed, record that explicitly rather than inferring what the answer used.

List the source URLs that appeared in your samples and classify them by type:

  • Vendor and provider pages
  • Directories and listing sites
  • Editorial and press coverage
  • Community forums and discussion threads
  • Product or technical documentation

Then record what each source answered that the client page did not, and at what level of detail. This is usually where the actionable finding lives. A directory that answers a narrow local question, or a documentation page that states a specific limitation plainly, tells you what shape of answer the client page is missing.

Watch for a specific pattern: cited sources that mention the client but are not owned by the client. That can change the recommended action from on-page work to source and coverage work, because the sampled answer is displaying third-party descriptions of the business.

Two restraints apply. And do not rank or score rival pages from a single sample, because one sample cannot support an ordering.

Step six: separate findings from interpretation in your audit record

An audit that mixes raw evidence with opinion is harder to review and easier to dispute. Structure the record in three layers, kept visually distinct in the deliverable.

  • Observations: the questions asked, the assistant tested, the market, the timestamps, the mention, citation, source and unavailable states, and the traditional search results
  • Interpretation: what you believe the observations indicate, stated as reasoning rather than fact
  • Recommended actions: what you propose doing, tied to the specific observation that motivated it

Attach a confidence statement to each interpretation and say what it rests on. Interpretations supported by several questions and consistent sources deserve more confidence than an interpretation resting on one sample of one question. Say which is which.

Contradictory evidence is normal and should be shown rather than smoothed over. A page found in traditional search for a query but absent from the sampled AI answer to the corresponding question is a genuine finding. Present both results side by side and treat the divergence as the condition to investigate.

Step seven: convert findings into prioritized next actions

The commercial value of the audit is the action list. It should be specific, grouped and ordered by something defensible.

Group recommended actions into four families:

  • Question-fit changes: retarget the page, split it, or move a question to a page better suited to own it
  • Content and structure changes: add a direct answer passage, restructure headings, isolate definitions, add a comparison table
  • Entity and evidence changes: clarify what the business does, where it operates and what supports each claim
  • Off-page source changes: address directories, documentation and third-party coverage that answered the question instead

Order the list by implementation effort and by how directly each action addresses an observed gap. Do not order by predicted lift, because predicted lift is not something this audit measured.

Recommend a re-check after implementation using the same search queries, AI questions, market and test conditions, with the new date recorded. A dated before-and-after comparison is the honest way to evaluate whether the observations changed and establishes a clear baseline for future re-checks.

State plainly in the deliverable that no action can be presented as guaranteeing a mention, a citation or a ranking change.

An illustrative prioritized list, using hypothetical findings rather than real client data:

  • Add a direct, self-contained answer to the pricing-model question in the opening section, since the sampled answers cited a third-party listing for that topic. Low effort, addresses an observed gap directly.
  • Make the service-area detail available in accessible page output if it is currently loaded only after interaction. Low effort, addresses an observed access condition.
  • Split the combined services page so the comparison question has a page that owns it. Higher effort, addresses an observed question-fit finding.
  • Update the two directory listings that were cited in place of the client page. Moderate effort, off-page source action.

That example is illustrative only. Your list should name the specific observation behind each item.

Limitations you should disclose in every AEO audit

Disclosure is a deliverable component, not a disclaimer buried in a footer. Including it consistently reduces expectation risk and makes the audit easier to defend.

  • Findings describe one URL under specific conditions on a specific date. They do not describe all AI answers everywhere.
  • One AI sample per question makes a finding preliminary.
  • Testing more questions broadens coverage across intents. Taking more samples of the same question helps characterize variation for that question. Neither removes uncertainty.
  • Answers vary by assistant, exact question wording, market, location, time and source availability.
  • Small convenience samples should not be treated as statistically representative or used to claim significance or a margin of error.
  • Unavailable evidence is labeled as unavailable. It is not evidence of absence.

A short disclosure block you can adapt for client deliverables:

This report records how one URL appeared in traditional search results and in sampled AI-generated answers for a defined set of questions, in a stated market, on the date shown. AI answers vary by assistant, question wording, market, location and time, so results may differ on another date or in another context. Findings based on a single sample per question are preliminary. Where evidence could not be retrieved, it is marked unavailable rather than treated as absence. Nothing in this report predicts or guarantees future mentions, citations or rankings.

Packaging the audit as a repeatable service

Once the workflow is stable, define exactly what the client receives and how the audit fits into the broader reporting relationship.

Define a fixed deliverable so scope does not drift between client projects:

  • The agreed search-query set and AI-question set, recorded separately and verbatim
  • The evidence record, with dates, market context and mention, citation, source and unavailable states
  • Your interpretation, with a stated confidence level per finding
  • A prioritized action list tied to specific observations
  • The disclosure block

Set scope boundaries for each audit and write them into the agreement. The variables worth fixing are the number of queries and questions, the number of URLs, the market context and whether a re-check is included. Vague scope on this kind of work turns into unpaid sampling, because there is always another question the client wants tested.

Fitting it into an existing reporting rhythm works best when the AI-answer evidence stays in its own section rather than being merged into traditional SEO reporting. The two surfaces answer different questions, and clients understand the distinction faster when the report structure reflects it.

The packaging tradeoff is straightforward. A one-off diagnostic answers what the current conditions are and what to change next, which suits onboarding and a stalled page. A periodic re-check answers whether conditions changed after implementation, which suits retained work, but it costs more delivery time and it still cannot establish a trend from sparse samples. Choose based on what the client actually needs to decide, and say which question each option can and cannot answer.

Choose the report that matches the audit scope

The Standard Visibility Report is a starting diagnostic. It infers up to three search queries and corresponding AI questions from the page, then records how the URL appears in traditional search and in one sampled AI answer per question. This shows what Rank Flywheel identifies from the page as it currently stands; it does not test a consultant-defined input set.

If they do not, treat that mismatch as a scoping finding rather than as evidence about the exact queries and questions the client wanted assessed.

When the audit must test exact search queries and AI questions agreed with the client, use the Deep Visibility Report. Deep supports selected target queries and confirmed AI test questions, three separate AI samples per question, deeper search evidence, structural analysis of accessible ranking pages and an included rerun.

The first Standard Visibility Report is free and can be generated and viewed without signup. Run a free Standard Visibility Report on one real client URL.

A practical AEO audit record

Use one row for each mapped search-query and AI-question pair. If a search query or AI question has no direct counterpart, keep the unused field blank rather than forcing a false match.

FieldWhat to record
PageExact URL and agreed commercial purpose
Search queryExact traditional-search input
AI questionExact natural-language question
Test contextMarket, location, assistant and timestamp
Search resultObserved position or not observed, plus the result depth checked
AI sample statusValid or unavailable
MentionYes or no, with a short supporting excerpt
CitationExact cited URL or none displayed
Other displayed sourcesExact source URLs shown in the answer, or none displayed
InterpretationWorking hypothesis and confidence
Recommended actionSpecific change tied to the observation

How is an AEO audit different from a standard SEO audit?

A standard SEO audit typically examines a whole site for technical health, content coverage and search performance. The AEO audit described here fixes a much narrower unit of analysis: one URL, one market context, one measurement period and defined sets of search queries and AI questions, assessed as a candidate source for AI-generated answers. Traditional search evidence is still collected, but it is recorded as a separate surface rather than blended into the AI-answer findings.

How many questions should I test for one client page?

Test as many as you can support with real recorded evidence, and no more. A small set of questions written in natural phrasing, each with a full evidence record, is more defensible than a long list with partial results. If the client wants broader coverage, treat that as a scope decision written into the agreement rather than an informal extension.

Is one AI sample per question enough to draw a conclusion?

One valid sample is real evidence and enough to open a diagnosis and identify actions worth testing. It is not enough to settle a question, so findings based on a single sample per question should be labeled preliminary. Taking more samples of the same question helps characterize variation for that question. Testing more questions broadens coverage across intents. Neither produces certainty, because answers vary by assistant, wording, market, location and time.

What do I do when the page ranks in search but is not cited in the sampled AI answers?

Report both results side by side and treat the divergence as the finding. Search position is not a proxy for AI-answer inclusion, so the correct next step is diagnostic: review whether the page answers the tested question directly and early, whether the answer can be lifted as a discrete passage, whether the substantive content is accessible without interaction, and which sources the sampled answer displayed or cited instead.

Can an AEO audit tell the client they will get cited after the changes?

No. The audit records observed conditions on a specific date and proposes actions tied to those observations. It cannot guarantee a mention, a citation or a ranking change. The defensible way to evaluate a change is a dated re-check using the same search queries, AI questions, market and test conditions after implementation, presented as a before-and-after comparison.

How should I handle evidence that could not be retrieved?

Label it unavailable in the record and in the client deliverable. Unavailable evidence is not evidence of absence, and it should never be counted as a not-found result or replaced with an estimate. Keeping that distinction visible is what allows the rest of the audit to be trusted.

#AEO audit for a client page

See the flywheel in action

This very article was published through Rank Flywheel.

Join the Waitlist