Skip to content
Back to Blog
Digital Strategy

AI Visibility Test vs GEO Audit: What Each One Measures

11 Eylül 2026
Next GEO Agency
AI Visibility Test vs GEO Audit: What Each One Measures

A free AI visibility test shows how ready your site is for AI crawlers and answer engines; it does not show whether your brand is actually named in ChatGPT, Gemini, Perplexity or Google AI Overviews answers. These are two separate questions, and what answers the second is a GEO audit: one run with your own sector's questions, dated, and with the raw answers kept on record.

In practice the distinction shows up like this. You enter your site address into a test tool and a few seconds later a colourful scorecard arrives: bot permissions green, schema partly amber, "citability" middling. The same week an agency offers you a "free GEO analysis". Both go by the same word, "visibility". Yet one is checking whether the door is open; the other ought to be checking whether you are being talked about inside.

This article gives you two things: how to read a test tool's result correctly, and how to weigh the quality of an audit proposal or report. Metric definitions and a monthly tracking routine are not covered here; we set those out separately in our article on the five metrics for measuring AI visibility. The subject here is the one-off audit and the anatomy of its report.

Why being ready and being named are not the same thing

The question "is our site visible in AI?" actually asks about three separate layers at once. An answer given without knowing which layer was examined gets misread, even when it is correct.

Readiness (the site side). Can crawlers get into your site, does the page content arrive from the server fully rendered, is there structured data, are the pages written in a layout that answers questions. These can be measured automatically and give the same result for the same site every time.

Visibility (the answer side). When a given set of questions is put to the engines, does your brand appear in the answer, is it cited as a source, is what is said about you correct. This layer can only be measured by asking the questions and reading the answers, and the result can shift a little on each run.

Entity and external signals. Are your business name, address and service area written the same way on your site, on Google Business Profile and in third-party listings; how are you mentioned in reviews and on other sites. Automated test tools usually do not look at this layer at all.

A site that has turned every line green on the readiness layer may not appear on the answer layer at all. The reverse happens too: a brand with messy technical infrastructure but widely mentioned in its sector can feature regularly in answers. Readiness is a precondition for being named, not the thing itself.

What free test tools actually check

When you enter an address into an AI visibility test tool and get a score, what is being checked is mostly the same family of signals on the readiness layer. As you read the result, separate what each item proves from what it does not.

  • AI bots allowed in robots.txt. Whether crawlers such as GPTBot, ClaudeBot and PerplexityBot are blocked. What it proves is the declaration in the file. What it does not prove is that the bot can actually get in; a security layer in between can contradict the file.
  • Presence of an llms.txt file. Whether the file exists and is readable. Most tools do not check whether its content is accurate and current.
  • Schema types. Which schema.org types are present on the page. The presence of a type is measured; whether the information inside it matches what is genuinely visible on the page usually is not.
  • Fully rendered HTML from the server. Whether the content arrives even without scripts running in the browser. This is the most valuable check on the list; a page that arrives empty makes everything else meaningless.
  • Citability signals. Structural cues such as question-shaped subheadings, short answer paragraphs, and date and author information. This is where tools differ from one another most, because there is no commonly agreed definition of the word "citable".

All five checks are useful, and all five are about your site. None of them reads an engine's answer.

The tool gave a score out of 100: how far should you trust that number?

A scorecard is persuasive, because a single number makes deciding easy. The problem is that how the score is calculated is often not disclosed. It does not say how many points each check carries, how much a missing item pulls the total down, or how a "partly present" status is counted.

This has two consequences. First, scores from two different tools cannot be compared. Scoring high on one and low on the other does not mean your site is contradictory; it shows that the two tools use different weightings. Second, even a rising score on the same tool is not always meaningful: you can pull the score up quickly by fixing an item that carries a heavy weight but has an uncertain effect on the answer layer.

The practical way to read it is this: set the total score aside and open the item list. For each red line, ask a single question: if this item is fixed, why would an engine be more likely to find me for a question my customer asks? Items with a concrete answer go onto your work list. Items whose answer is "the score goes up" can wait.

Four things a readiness test cannot see

The limits of a free visibility test come not from a technical shortcoming but from its design. The tool looks at your site; the four things below sit outside your site.

The answer itself. What an engine writes in response to your customer's question. Did your name come up, in which sentence, with what description, and which sources were shown next to it. The only way to know this is to ask the question and read the answer.

How well the question set fits your sector. Even work that looks at the answer side produces false confidence if it is done with the wrong questions. A generic question such as "En iyi ambalaj firmaları" (best packaging companies) and the question a purchasing manager actually types, "gıdaya uygun sertifikalı karton ambalajı küçük partiyle kim üretir" (who makes certified food-safe cardboard packaging in small batches), do not bring up the same brands. The question set should be built by someone who knows the sector and the customer's language.

Entity consistency. One business name on your site, a different abbreviation on your Business Profile, the address of a closed branch in an old directory listing. Engines recognise you by stitching these pieces together; when the pieces contradict each other, it becomes easier for them to produce incomplete or wrong answers about you. Because the tool only crawls your own site, it cannot see this contradiction.

Change over time. A single test is a snapshot of one moment. To see what has changed on the answer layer, the same questions have to be askable again later in the same way. A scorecard does not leave behind the record that would make that repetition possible.

Why the same question gets a different answer every time

Asking a question today, asking it again tomorrow and getting a different answer does not show that the measurement is broken. Generative engines can give different answers to the same question at different times, on different accounts and in different languages. How often this happens, and how much each variable contributes, is not known from the outside, because the engines do not publish how they select sources and that logic can change without notice.

What this means for an audit is clear: a single run is not evidence on its own. The sentence "we asked ChatGPT and your name did not come up" does not count as a finding unless it is written down on which day, with which account and in which language the question was asked, along with the full text of the answer. By the same token, "your name came up" does not count as success on its own either.

Every answer line in an audit report should carry these four pieces of information:

  1. The date the question was asked
  2. The engine and, if any, the mode used
  3. The language of the question and the state of the account (signed in or not, clean history or not)
  4. The raw text of the answer, unsummarised

These four pieces of information turn the report from a claim into a repeatable measurement. How to set up monthly tracking is a separate subject; what matters for the audit is that the starting point is recorded in a way that can be reproduced later.

The eight sections of a good GEO audit report

The quality of an audit report is measured not by its page count but by whether each finding comes with its evidence. The eight sections below are the minimum structure you should expect to find in a one-off GEO audit report.

SectionWhat it containsWhy it is neededHow it is evidenced in the report
Scope and methodWhich engines, which language, which date range, account conditionsSets the limits of what the results representA method note on the first page, a date on every answer line
Question setQuestions built for your sector, grouped by intentWrong questions produce false confidenceThe full list, with version name and date
Answer findingsAnswers where the brand appears, is cited as a source, or is described wronglyWhere visibility is actually measuredA pointer to the raw answer text for every finding
Cited sourcesURLs shown as sources in the answersShows which sites the engine relies onA URL list, with the question each came from
Site readinessBot access, fully rendered HTML from the server, schema, llms.txtThe precondition for the answer layerServer response codes, excerpts from the page source, validator output
Entity consistencyBusiness name, address, service area and team information compared across sourcesConflicting information feeds wrong answersA source-by-source side-by-side table
Content gapsQuestions with no page to answer themThe list of work to be writtenA mapping between each question and the relevant page
Prioritised findingsA work list ranked by impact, effort and ownerTurns the report into a planWhich finding each item rests on

On top of these, the raw data itself should be handed over: the full answer texts, in an editable file. The body of the report is interpretation; the appendix is what that interpretation rests on.

The last row of the table is the weakest part of most reports. If dozens of findings are listed and nothing says which comes first, what you have is an inventory, not an audit.

How to weigh a "free GEO analysis" offer

At the proposal stage you cannot yet see the report's quality, but a few questions will tell you what kind of report is coming. These questions apply to any offer, free or paid.

Can I see a sample report? A report with another business's details removed is enough. If the sample has no raw answer texts, yours will not have them either.

Can I see the question set before the audit and add to it? The questions should be your customers' sentences. A report that arrives without you ever seeing the set is a report where you do not know whose questions it was run with.

Which engines will be asked, and under what conditions? If the answer is "all of them", ask for the detail. Without the engine name, language and account condition written down, there is no repeatable measurement.

Will the raw answers stay with me? When the audit ends, the question set and the answer texts should be yours. It is the only way you can repeat the same measurement six months later, yourself or with someone else.

How are the findings tied to priorities? "Schema missing" is an observation. "This service page has no information that answers this question, which is why you do not appear in this question group" is a finding tied to a priority.

One sign to watch for: if most of the report consists of screenshots from a test tool and there is not a single line about the answer layer, what you are being offered is a readiness test. There is nothing wrong with that; it just needs to be called by its proper name. Choosing a tool for regular tracking is a separate decision; our comparison of AI visibility tools covers that side.

The minimum checks you can do yourself without waiting for an audit

Even before commissioning an audit, you can check three things yourself. The aim is not a full measurement; it is to spot an obvious blocker without handing work to anyone.

  1. Can bots actually get in? Permission in robots.txt is not enough; you need to look at what the server returns to a request arriving with a bot's identity. We set out the steps in our guide to verifying whether AI crawlers can reach your site.
  2. Is there an llms.txt, and is it current? If the page list in the file does not match your site as it stands today, the file's existence gains you nothing. What to write in it is covered in our llms.txt guide.
  3. Is your name the same everywhere? Open your site's contact page, your Google Business Profile and two or three directory listings that mention you, side by side. If the business name, address, phone number or service description differs, make a note. This check is usually absent from automated tools, and it is where the most common problem turns up.

If these three are clean, it is time to look at the answer layer.

The report is in your hands: how findings turn into a work plan

A good report ends with a long list of findings. When turning that list into a work plan, the order should be built not on the size of each finding but on its dependencies.

Access first. If bots cannot get to the page, or the page arrives empty, every piece of work done on content and schema stays invisible. Findings of this kind go to the top of the list, however many there are.

Then accuracy. Answers that give wrong information about you, and listings that contradict one another. Before trying to increase visibility, you need to be sure you are described correctly where you are already visible.

Then content gaps. Questions with no page to answer them. Each one goes onto the work list as a new page or as a section to be added to an existing page.

External signals last. Being present on the third-party sites cited as sources in answers, and work on reviews and mentions. This is the longest-running item; it starts early, but its effect shows up last.

Turning this order into a strategy document with target questions, a content calendar and structured data is a separate piece of work; we set out its steps in sequence in our guide to building a GEO strategy.

How an audit starts at Next GEO Agency

At Next GEO Agency an audit can be taken in two ways: as the first stage of a monthly engagement, or as a one-off audit without a monthly package. The scope and fee of both routes are written on our GEO pricing page. In the monthly engagement the first two weeks go to discovery and baseline measurement: the real questions in your sector are put one by one to ChatGPT, Gemini, Perplexity and Google AI Overviews and the answers are recorded with their dates; the state of access and structured data on the site side and the consistency of entity information are gathered in the same report. This baseline measurement report becomes the basis for comparison in the months that follow.

No guarantee of rankings or mentions is given; generative search systems are black boxes and their source selection logic changes without notice. What can be offered is a repeatable record showing what was measured and what changed. The full set of work items is written on our GEO agency service page. If you would like to start from your own position, for a free business analysis all you need to send is your site and a few of the questions your customers ask most often.

Frequently Asked Questions

What is a free AI visibility test good for?

It shows how ready your site is for AI crawlers and engines: whether bots are allowed, the llms.txt file, schema types, and whether the content arrives from the server fully rendered. It is useful for quickly finding obvious technical blockers. But it does not measure whether your brand is named in answers; that question can only be answered by putting real questions to the engines.

Does a high score on a test tool mean I will show up in ChatGPT?

No. A high score shows that your site's readiness signals are in place; readiness is a precondition for being named, but not the thing itself. Whether an engine recommends you also depends on the other sources in your sector, the external records about you and the question itself. Because most tools do not disclose how the score is calculated, scores from different tools cannot be compared with each other either.

What should a GEO audit report contain?

At least eight sections: a scope and method note, a question set built for your sector, answer findings, the sources cited in the answers, site readiness, entity consistency, content gaps and findings ranked by priority. On top of these, the full raw answer texts should be handed over, because the report's interpretations rest on those texts.

Is a one-off audit enough?

A one-off audit shows the starting point and the priority tasks; that is a valuable output. But because answers change over time, seeing progress requires the same question set to be asked again later under the same conditions. That is why an audit leaving the question set and the raw answers with you matters as much as the report itself.

I get different answers to the same question on different days; is the measurement wrong?

No, this is in the nature of generative engines; the same question can produce a different answer at a different time, on a different account or in a different language. Its effect on measurement is that a single run does not count as evidence on its own. Recording each answer with its date, engine, language, account state and raw text is what makes it possible to tell later whether a difference is a real change or ordinary fluctuation.