Do not start by asking ChatGPT what your company does. That test may confirm that the model recognises a name you have already handed it. It does not show whether the business appears when a potential buyer asks for help without knowing you exist.
This distinction matters for professional-services firms in Dubai and the wider UAE. A sensible branded summary can create false confidence while competitors are being surfaced for the category, comparison and recommendation prompts that actually begin a buying process.
Do not touch the verdict after one branded prompt
A prompt such as What does [company name] do? tests factual brand recognition. It can reveal whether ChatGPT has confused your location, service range or sector. That is worth checking, but it is a representation check, not a discovery test.
It also gives the model too much assistance. The company name acts as the answer key. A buyer who asks which firms specialise in a particular service for a particular need has not supplied that clue.
Checking ChatGPT brand visibility properly means separating branded factual prompts from unbranded discovery prompts, recording what the model says and cites, comparing relevant competitors, and repeating the same test later. A branded answer tests whether a model can describe a known business. An unbranded buying prompt tests whether it may surface that business before the buyer knows its name. Neither result guarantees a future recommendation, but the second is closer to genuine market discovery.
Check discovery visibility before fixing anything
Build a small prompt set before changing copy, publishing more schema or commissioning a fashionable AI file. Ten to fifteen well-chosen prompts will tell you more than repeatedly asking a broad question until you receive an agreeable answer.
1. Split the prompt set into three groups
| Prompt group | What it tests | Example shape |
|---|---|---|
| Branded factual | Recognition and accuracy | What does [company] do in Dubai? |
| Unbranded category | Basic discovery | Which firms provide [service] in Dubai? |
| Buying and comparison | Commercial relevance | Which [service] providers suit [buyer need], and how do they differ? |
Use wording a real buyer could plausibly enter. Add geography only where it genuinely affects selection. Add a buyer constraint where it changes the shortlist, such as sector, company size, project type or a specific requirement.
A Dubai professional-services firm might ask ChatGPT what its own company does and receive a reasonable response. Its more useful test is a prompt such as: which Dubai firms specialise in [service] for [buyer type] with [relevant requirement]? If the firm is absent from that answer, the branded summary was never evidence of discovery visibility.
2. Keep the first run clean
Run prompts in a fresh session. Do not spend several messages discussing your firm and then ask for recommendations. Conversation context can alter the answer, and it makes the result harder to compare with a future run.
Record the exact prompt, date, platform, location settings where relevant, whether web search was used, brands mentioned, order of appearance, description accuracy and any visible sources. Save the response rather than relying on memory.
This boring detail prevents a common reporting error: comparing a searched answer in one session with an unsourced answer in another, then treating the difference as a visibility trend.
3. Look at sources, not only names
A mention is not always a useful representation. If sources are visible, note whether they are your website, an independent publication, a directory, a competitor page or something unrelated. Then ask whether those sources actually support the description given.
For your own website, check the obvious operational basics before drawing large conclusions. A core service page hidden behind client-side JavaScript, missing from the XML sitemap, or described with a different service name from the homepage makes extraction and entity interpretation unnecessarily difficult. It may not explain every result, but it is a legitimate audit finding.
Use a record sheet, not a collection of screenshots
A simple spreadsheet is enough for a first manual pass. Use one row per prompt and preserve the wording exactly. Screenshots are useful supporting material, but they are poor data if nobody can tell which query produced them.
- Prompt: the exact wording, including location and buyer qualifier.
- Test conditions: date, platform, session state and whether search was enabled.
- Your brand: mentioned, not mentioned, or mentioned inaccurately.
- Competitors: each relevant competitor named in the answer.
- Source behaviour: sources shown, no sources shown, or sources that do not substantiate the statement.
- Notes: qualification language, missing services, wrong geography or confused company identity.
Do not turn a single run into a scorecard with suspicious precision. The practical baseline is a prompt-by-prompt record that shows where you appear, where you do not, and where the description needs correction.
Escalate when the pattern is consistent
Results can change between runs because the platform, retrieval behaviour, available sources, conversation context and wording can change. That does not make testing pointless. It means the test needs controls.
Repeat the same prompt set after a defined interval, using the same conditions as far as possible. Look for patterns across several runs. If your business is repeatedly recognised by name but absent from commercially relevant unbranded prompts, you have a discovery gap worth investigating.
Do not assume the remedy is one thing. The cause could sit in unclear service definitions, weak location signals, thin answer-ready pages, inconsistent organisation details, limited independent corroboration, or the platform simply preferring other sources for that query. Structured data and crawlability can help make valid information easier to interpret, but they cannot compel an independent platform to recommend a company.
For a wider baseline across ChatGPT, Gemini, Perplexity and Copilot, a FlareFalcon AI visibility audit expands the manual exercise into tracked prompts, competitor comparison and website-level checks for machine readability and entity clarity. The value is not a magic recommendation score. It is a more defensible view of what needs investigation.
Questions that tend to appear during testing
What prompts should I use to test ChatGPT visibility?
Use a mix of branded factual, unbranded category, comparison and recommendation prompts. Make them specific enough to resemble buyer research, but do not over-specify every feature until only one business could fit. Include relevant location, buyer type and service constraints, then keep the prompt set stable for later comparison.
Should I include my brand name?
Yes, but only in a separate factual recognition check. Branded prompts are useful for finding incorrect service descriptions, wrong locations and confused identities. They should not be the only test, because they cannot show whether ChatGPT introduces your company when the buyer has not named it.
Why can ChatGPT results change between runs?
Generated answers are not fixed search rankings. Prompt wording, session context, search behaviour, source availability and platform changes can all affect the response. Use fresh sessions, preserve test conditions and assess repeated patterns rather than celebrating or panicking over one answer.
