A site can pass every crawler-access check and still have no evidence that an AI crawler requested its important service URLs. We see this often with organisations in Dubai and the UAE that have allowed GPTBot globally, checked robots.txt, then assumed the technical work is complete. During the period under review, the logs may show no request at all for the pages they expect an AI system to understand.
That is not a contradiction. It is the difference between permission and observation. A robots.txt rule can allow a crawler through the gate. It cannot show that the crawler arrived, which URL it asked for, whether the server returned useful content, or what happened after the request.
The incident starts after the access check
Start with the AI crawler access checklist to establish whether the relevant bots are permitted to fetch the site. Then move to server, hosting or CDN logs. This is where theoretical access becomes an observable event.
A practical example: GPTBot is allowed sitewide, the sitemap contains the main service pages, and a fetch test returns HTML. Yet a review of CDN request logs for the investigation period finds no verified GPTBot requests for those service URLs. Changing the robots.txt rule again would achieve little. The next questions are whether those URLs are discoverable, internally linked, consistently canonicalised and available without rendering dependencies.
AI crawler logs are records from a web server or CDN that can show whether a verified crawler requested a URL, when it did so and how the server responded. They prove observed HTTP activity, not that a platform indexed, understood, cited or recommended the page.
Match the request to the crawler you are investigating
Do not search logs for any user-agent containing AI and call it done. User-agent strings can be spoofed. Start with the crawler whose role matches the question, then verify its identity using the provider’s published method.
For OpenAI, the current crawler documentation distinguishes GPTBot, OAI-SearchBot and ChatGPT-User. Their jobs are different. GPTBot relates to content that may be used for model training, OAI-SearchBot is associated with search, and ChatGPT-User is used when a user asks ChatGPT to visit a site. A request from one is not evidence of activity by the others.
| Log signal | What to inspect | What it can support | What it cannot establish |
|---|---|---|---|
| User-agent | Exact string and request volume | A candidate crawler visit | That the request is genuine |
| Provider verification | Reverse DNS and forward DNS confirmation where documented | That the request originated from the claimed crawler infrastructure | Downstream indexing or use |
| Requested path | URL path, query string and host | Which page or asset was requested | That the page was fully interpreted |
| Status code | 200, redirect, 403, 404 or 5xx response | How the server responded at that time | That a 200 response contained useful page content |
| Sequence and timing | Requests before and after the page fetch | Basic crawl depth and asset dependency clues | A complete picture of browser rendering |
Verify before you count
A plausible GPTBot user-agent by itself is weak evidence. OpenAI’s documentation provides verification guidance based on reverse DNS lookup followed by a forward lookup check. In plain terms, resolve the requesting IP address to a hostname, confirm it sits within the provider’s documented domain, then resolve that hostname back to the same IP address.
This is worth doing before sending a crawler report around the business. Unverified user-agents are common enough that a spreadsheet of string matches can create confidence without proof. If your CDN retains the client IP, request path, timestamp, user-agent and response status, you have the ingredients for a defensible review.
Keep raw request fields intact
Export enough detail to reconstruct the event: timestamp in a known timezone, hostname, full path, status code, user-agent, client IP, referrer where available and edge or origin response information. A 200 from the CDN cache and a 200 generated by the origin are not always operationally identical, particularly when middleware, bot controls or regional rules intervene.
Read the URL trail, not just the bot total
A verified crawler may request the homepage, robots.txt, sitemap.xml and a few assets while never reaching the pages that carry the commercial explanation. A total request count hides that distinction.
- Filter verified requests by crawler and hostname.
- Group paths into homepage, sitemap, service pages, articles, contact pages, media and technical endpoints.
- Review response codes and redirect destinations for every priority URL.
- Compare requested URLs with the pages you want to support AI visibility, entity clarity and answer-ready content.
- Inspect the request order for sitemap retrieval, navigation paths and repeated failures.
Pay attention to boring faults. A service page may return a 301 to a trailing-slash version that is then blocked by a CDN rule. Or its main explanation may only appear after a JavaScript call that returns a 403 to a bot protection layer. The initial document request can look healthy while the useful content never becomes available in the form the crawler receives.
When absence in logs is meaningful
No request in logs is not proof that an AI provider will never visit the site. It is evidence limited to the log source, retention window, hostname and period examined. Check whether another CDN, subdomain, legacy origin or managed host receives the traffic before making a firm statement.
Still, a clean review with no verified requests to key pages changes the order of work. Before rewriting content, changing schema or buying authority coverage, inspect crawl discovery. Confirm sitemap inclusion, canonical consistency, internal links from crawlable pages, stable status codes and whether important copy is present in the initial response or depends on client-side rendering.
For wider improvements once the crawl path is understood, use the website optimisation guide for AI search as a practical reference for content, technical delivery and evidence quality. Logs answer a narrow but useful question. They should direct the investigation, not carry the whole GEO case.
Logs do not prove understanding or recommendation
A crawl request proves that a verified agent asked your infrastructure for a resource and received a particular HTTP response. It does not prove that the content was rendered successfully, retained, indexed, retrieved for a later answer, cited or used to recommend the organisation.
Those stages require separate evidence. Rendering is tested through delivery and content inspection. Discovery and description are tested through controlled prompts over time. Citations and referrals are their own observations. Treating one server request as a completed AI visibility journey is simply a neater version of the original robots.txt mistake.
Questions about AI crawler logs
Which AI crawler user-agents should we check?
Check the agents relevant to the platform and outcome you are investigating. For OpenAI, that may include GPTBot, OAI-SearchBot and ChatGPT-User, but do not treat them as interchangeable. Also review published documentation from other providers you care about, such as Google, Microsoft and Perplexity, because names, stated functions and verification methods can change.
How do you verify that a crawler request is genuine?
Use the provider’s current verification instructions rather than trusting the user-agent field. For OpenAI requests, this includes reverse DNS lookup of the client IP and a forward DNS lookup that resolves back to the same address, with the hostname matching the documented provider domain. Retain the lookup result with the log extract.
What does a crawler request actually prove?
It proves a verified crawler requested a particular URL at a recorded time and received the logged response. It can help diagnose discovery, redirects, errors and delivery dependencies. It does not prove complete rendering, indexing, model training, answer retrieval, citation or recommendation. Each is a separate claim with a separate evidence threshold.
