Blog | FELD M

AI crawler and AI SEO: The two faces of AI traffic | FELD M

Written by Dr. Ramona Casasola-Greiner, Dr. Matthias Böck | Oct 1, 2026, 9:19:55 AM

AI traffic series - Part 1 of 5

 For the kick-off of our five-part series about AI traffic, we'll be looking at the two fundamental phenomena behind AI traffic: AI bots and AI crawlers. These automated bots visit websites independently, comb through the web, and scrape content, all with the help of artificial intelligence. 

Series overview

Previous article:

Next articles:


Coming soon:

  • Part 3: Filtering AI bot traffic: How to clean up your data in GA4, Adobe & Piano

  • Part 4: Pay-per-crawl & co.: How companies are responding to AI crawlers

  • Part 5: AI agents and web analytics: New rules for attribution

In contrast, AI SEO refers to the use of AI to improve either the search engine optimization or the content itself, so that ChatGPT and co. can find and display your products, website, or content to their users.

Since people increasingly research directly on AI platforms such as Gemini and ChatGPT, AI SEO is becoming more and more important. Both aspects – AI crawlers and AI SEO – influence your web analytics, but in different ways.

 

AI crawlers: When bots sweep the web clean

AI crawlers are specialized bots that crawl websites to collect data for AI applications.

Unlike classic search engine crawlers (like Googlebot), which primarily index the web for web search, AI crawlers pursue three fundamental goals: They collect new training data for large language models (LLMs), support indexing for AI search, or search and take action directly on the page in response to a user prompt.

Well-known examples include:

  • GPTBot by OpenAI or ClaudeBot by Anthropic, which are used to scrape vast amounts of web content for training language models.
  • Other bots like ChatGPT user agent or PerplexityBot visit websites specifically to retrieve live information; for example, when a user asks a generative AI a question, the bot fetches the answer straight from a website.
  • Google-Extended or Applebot in AI mode serve to gather content for the two tech giants' in-house AI systems (Google Gemini, Apple Foundation Models, etc.)

On a technical level, AI crawlers can differ significantly in their behavior. Many of them do not execute JavaScript, meaning that while they do load HTML and often CSS/JavaScript resources, they don’t interpret client-side scripts.

An analysis from Vercel, for instance, showed that none of the major AI crawlers (GPTBot, ClaudeBot, PerplexityBot, etc.) currently actively render JavaScript content. However, some bots, such as Applebot, have a fully fledged headless browser engine and can render websites almost exactly like a real browser.

Beyond that, their focus also varies: a comparison of 570 million GPTBot requests vs. 370 million Claude requests in a single month revealed that GPTBot requested roughly 58% HTML, while ClaudeBot favored images around 35% of the time — apparently, different AI bots crawl different types of content.

One thing almost all the bots do have in common, though: they're extremely resource-intensive. A Vercel report found that GPTBot and Claude together account for just under 1.3 billion page views per month - that makes up roughly 28% of Googlebot's volume. Statistically, that means for every four times Google crawls a page, there is more than one visit by an AI bot.

That’s why, in addition to the robots.txt, there's currently an ongoing discussion about an llm.txt file. While the robots.txt gives traditional web crawlers instructions on which parts of a website they can crawl and index, llm.txt is intended to provide instructions and content specifically for AI crawlers (you can find an example here: docs.anthropic.com/llms-full.txt).

 

AI crawlers and their key features

Crawler name
Purpose/ type
Key features
GPTBot (OpenAI) Training an LLM (ChatGPT) Very high volume (100 million requests per month); crawls all links; doesn’t execute JavaScript; ca. 34% of the requests end in 404 (bot “hallucinates” non-existent URLs)
ClaudeBot (Anthropic) Training an LLM (Claude) Similarly high volume as GPTBot; no JS; noticeably high percentage of images (ca. 35%) in the requests; similarly high 404 rate
ChatGPT-User (OpenAI) Live data retrieval for ChatGPT’s browser mode Crawls a page at the request of a ChatGPT user; respects robots.txt, but only crawls specific pages.
OAI-SearchBot (OpenAI) Indexing for ChatGPT plugins/search Builds a search index to make websites searchable for ChatGPT; similar frequency to traditional crawlers, but limited to certain partner sites.
PerplexityBot Index and live search for Perplexity AI Crawler operated by the AI search provider Perplexity.ai; aims to reduce the dependency on Google, generates moderate traffic volume (e.g., 24 million requests per month).
Google-Extended AI training (Google Gemini, Vertex AI) Opt-in-crawler: website owners can control whether their content can be used for Google’s AI training; uses Googlebot infrastructure (can execute JS).
Applebot (AI-Modus) AI training (Apple Foundation Model) Extends Applebot (originally for Siri/Spotlight search) to cover AI use; respects “Applebot Extended” (a robots rule to prohibit AI training); like Googlebot, it renders complete pages. (JS, Ajax).
Amazonbot Web crawling for Alexa AI and others Partly for training, partly for live responses (e.g., Alexa); reports on unusually high crawl rates; reputable bots respect robots.txt, but they are not always clearly identifiable.
Common Crawl (CCBot) General web crawling for public datasets Crawls the web for the freely available “Common Crawl” dataset, which is often used for AI training; very broad crawl, respects robots.txt; doesn’t execute JS.

Source: Compiled by FELD M from various sources including Vercel, Wired, and Search Engine Land.

 

To summarize: AI crawlers operate in large numbers and crawl aggressively. They scour every corner of your website, often more thoroughly than human users or Google, and they don’t care about consent, sessions, or performance.

In doing so, they cause massive distortions in your analytics data. Or, like Dennis Schubert, who manages the infrastructure of the Diaspora Social Network, aptly put it:

AI crawlers „don't just crawl a page once and then move on. Oh, no, they come back every 6 hours because lol why not.“

–Dennis Schubert

 

AI SEO: Search engine optimization in the age of AI

On the other side of the coin is AI SEO. The term covers two things:

Firstly: the use of AI tools to carry out classic SEO tasks, such as automating keyword research, creating content with support from AI, or carrying out intelligent on-page optimization.

Secondly: adapting your own content to the new AI-driven search environment. This is increasingly crucial as search engines like Google and Bing have begun to include generative AI in their search results, e.g., with Google’s AI Overview or Bing’s chat mode.

AI SEO makes many tasks easier for the SEO team: AI can generate content ideas, speed up competitive analysis, and suggest titles, meta tags, or content structure thanks to machine learning. The goal stays the same: Increase visibility in traditional search results - but methods are becoming more efficient. However, for those of us in the analytics space, this topic is of lower importance.

For us, the second aspect of AI SEO is much more interesting: Generative Engine Optimization (GEO), which aims to ensure that your content is present in AI-generated responses.

With Google's traditional ten blue links on the search engine results page now a thing of the past, more and more users now get the answers they're looking for directly from AI. As a result, content strategists have to ensure that their content is chosen for and reproduced correctly by the AI's answers - ideally with source references.

That means: Semantically rich content that addresses the user intent rather than just the keywords is more important than ever. Structured data (Schema.org markups) and precise answers included in the text can help ensure AI treats you as a trustworthy source.

That aside, one side effect of this development is that websites can deliver excellent content and yet still suffer traffic losses because users instead get their answers directly in the AI overview and never click through to a website.

Initial investigations show, for example, that the CTR dropped by an average of 15.5% when a news article was cited in Google’s generative search results (AIO - AI overview). And this is based on relatively early test runs. If generative search continues to gain relevance, these effects could grow exponentially.

For SEO, this means that success is no longer measured by clicks alone, but also by the content's presence within AI responses. As a result, optimization work now goes beyond traditional ranking factors, focusing additionally on context, authority, and the completeness of the information.

 

Summary: Caught between the bot onslaught and the desire for visibility

AI crawlers drastically increase the non-human traffic on your website and muddy your analytics data. AI SEO requires you to adapt your content and strategies to the new AI-driven search landscape in order to remain visible – even if it still means that fewer users click through to your website.

In the next part of this blog series, we'll focus on the concrete effects of AI crawlers on your web analytics data and share what you can do to stop them skewing your data.

Don't want to wait to read the next article or need help with a specific problem you're facing with AI crawlers? If you want to understand how AI crawlers and generative AI affect your digital visibility and data, reach out to our analytics experts. We'll help you separate the opportunities from the risks and get ahead of the new rules of the game, tailored to your platform and your product.

 

 

 

Go to the next article in the series:

Part 2: AI traffic and data quality: How AI bots distort your data