A team gets a better ChatGPT answer than last month and starts celebrating. Then somebody asks for the original prompt, the date it was run, the location setting and the scoring rule. Nobody has them. What looked like progress is now a screenshot with no control group.
For a Dubai B2B firm, that is an expensive way to run generative engine optimisation. Before changing service pages, adding structured data or commissioning a content programme, establish an AI visibility baseline that can distinguish a genuine shift from normal platform variation.
Decide what the baseline needs to prove
An AI visibility baseline is a dated record of how selected AI platforms respond to a fixed set of relevant prompts before a material website, content or authority change. It records whether the business appears, whether it is described accurately, whether it is cited or recommended, and which competitors or sources occupy the same answer space.
The aim is not to make model outputs pretend to be stable. They are not. The aim is to make your own testing stable enough that a later review is commercially useful.
Start by agreeing the decisions the measurement will support. A baseline should help answer questions such as:
- Are we visible for realistic category and problem-led discovery prompts?
- Does the platform identify our business, services and location correctly?
- Are we present as a citation, a recommendation, both, or neither?
- Which competitors repeatedly appear in comparable answers?
- Did a change improve answer quality, or did the test conditions change?
If the exercise cannot answer at least some of those questions, it is prompt collecting rather than measurement. For a wider definition of the categories worth tracking, use the AI share-of-voice measurement checklist. The baseline is the record that lets those categories be compared later.
Build a prompt portfolio before editing the website
Lock a prompt portfolio that reflects how a buyer might actually search. Do this before rebuilding your website, before paying for ads around a new offer, and before an agency starts reporting on supposed AI visibility gains.
Keep branded and discovery prompts apart
Branded prompts test recognition and entity accuracy. Discovery prompts test whether a platform can surface the business when the user does not already know its name. Combining them into one score produces a flattering but unhelpful result.
A useful portfolio usually has three groups:
- Branded prompts: business name, service name, founders where relevant, office location and common name variants.
- Category prompts: queries such as B2B AI visibility agency in Dubai or GEO consultancy for SaaS firms.
- Problem prompts: queries framed around a buyer need, such as how to improve AI search citations for a consultancy.
Keep the wording exact. A prompt asking for the best provider is materially different from one asking for a shortlist, an explanation or a local recommendation. Record the full prompt rather than a shortened label that later leaves room for interpretation.
Capture the conditions, not only the answer
Prompt tracking fails when teams save the attractive paragraph but omit the circumstances that produced it. Record the platform and model version where visible, access mode, location, date, language, prompt text and whether web search or browsing was enabled.
A practical detail matters here: check whether a test was run in a fresh conversation. A long chat thread can supply context that a prospective buyer would never provide. It may be useful for research, but it is not a comparable discovery test.
| Baseline field | What to record | Comparison rule |
|---|---|---|
| Prompt identity | Full wording, intent group and market | Reuse exactly unless creating a separately labelled test |
| Test conditions | Platform, date, location, language and browsing state | Match settings as closely as the platform allows |
| Business appearance | Absent, mentioned, listed or recommended | Apply the same label definitions each time |
| Entity accuracy | Business name, category, service, geography and factual claims | Log inaccuracies separately from non-appearance |
| Citations and sources | Linked sources, first-party pages and third-party references | Record source type, not merely citation count |
| Answer quality | Useful, partial, misleading or unusable for the stated prompt | Score against the same short rubric |
Score the answer without making the rubric theatrical
Use a small set of labels that different reviewers can apply consistently. For example, a recommendation means the business is positively presented as a suitable option for the stated need. A mention means it appears without a clear endorsement. A citation means a source link or source reference is attached to a relevant claim. These labels can overlap.
Also assess entity accuracy. A business that appears under the wrong category, wrong city or wrong service description has a different problem from a business that does not appear at all. This is where machine readability, page language and structured data can become relevant, but do not assume a schema change will force a platform to revise its answer.
One operational check is worth adding before interpreting a poor result: confirm that the service page intended to support the claim is crawlable, present in the XML sitemap and not mostly hidden behind client-side JavaScript. A page may render nicely in a browser while leaving little useful extractable content for other systems.
Set comparison rules before the first change
Write down what counts as a meaningful movement. A single favourable answer should not override a run of comparable tests. Nor should a one-off omission trigger a hurried rewrite of a service page.
For example, a Dubai consultancy might change its AI visibility service page after one poor ChatGPT response. If the next test uses different wording, a different browsing mode and a different date, there is no basis for attributing any change to the page. The platform may have varied, or the prompt may simply have asked a narrower question.
Use a dated review cadence instead. Run the core portfolio at a fixed interval, such as monthly for an active programme or quarterly for a stable site. Run an additional review after a substantial release, but label it as an interim check rather than treating it as a full trend point.
It also helps to separate baseline measurement from website diagnosis. A poor AI answer may point to an entity, content, technical or authority issue, but it does not identify the cause on its own. The distinction is covered in this AI visibility audit versus SEO audit comparison.
A one-page baseline record is enough to start
You do not need a grand reporting system. Start with one shared sheet or workbook containing the prompt portfolio, test conditions, captured answer, source references, labels, reviewer notes and the next review date. Save source screenshots or exports where permitted, but make the structured record the primary evidence.
Measure the starting line before arguing about the finish. It is less glamorous than publishing another page, but it stops normal answer variation being sold internally as success and gives later technical or content work a fair test.
Questions teams ask before setting a baseline
How many prompts belong in an AI visibility baseline?
Start with a manageable portfolio that covers branded, category and problem-led discovery. Ten to twenty well-defined prompts is often more useful than a large, loosely assembled list. Add prompts only when they represent a real buyer question, a key market or a meaningful service distinction. Consistency matters more than volume.
How often should an AI visibility baseline be rerun?
Use a fixed cadence that matches the pace of change. Monthly checks suit active GEO, content or authority programmes. Quarterly checks can suit stable websites with limited changes. Keep major releases and ad hoc tests separately labelled, since they should not be confused with a scheduled comparison point.
Should branded prompts be separated from unbranded prompts?
Yes. Branded prompts test whether a platform recognises and describes a known entity correctly. Unbranded prompts test discovery among alternatives when the business name is absent. They answer different commercial questions and should have separate labels, summaries and interpretation.
