Blog | FELD M

AI traffic and data quality: How AI bots distort your data | FELD M

Written by Dr. Ramona Casasola-Greiner, Dr. Matthias Böck | Oct 1, 2026, 9:25:53 AM

AI traffic series · Part 2 of 5 

 The sharp increase in bot traffic across the internet poses a considerable risk to the quality of your web analytics data. In the following blog post, we'll look at the metrics AI crawlers distort and why it’s a problem worth taking seriously. 

Series overview 

 Previous article: 

Coming soon:

  • Part 3: Filtering AI bot traffic: How to clean up your data in GA4, Adobe & Piano

  • Part 4: Pay-per-crawl & co.: How companies are responding to AI crawlers

  • Part 5: AI agents and web analytics: New rules for attribution

 

Distorted traffic volume and page views 

The most obvious effect is artificially inflated visitor numbers. When millions of additional requests suddenly arrive from AI bots, your total traffic volume might explode without a single additional real customer ever visiting your site.

A recent example comes from hosting provider Vercel, which reported that OpenAI's GPTBot alone generated 569 million requests in a single month, with Anthropic's Claude accounting for a further 370 million. That comes to almost a billion requests in one month combined, on a single hosting network. Those two bots thus reached roughly 28% of monthly Googlebot traffic.

Cloudflare has even reported that AI crawlers combined now account for more than 50 billion requests per day on its network — around 1% of all the web traffic Cloudflare handles. And we're only at the very beginning.

This means that, without properly filtering your analytics setup, a storm of bot traffic will dominate your visitor numbers. You risk believing that your website has seen an enormous traffic surge, when in reality it might have just been GPTBot and its partners in crime busily collecting content. Marketing campaigns could mistakenly be considered extremely successful, even though the "new visitors" were not people but machines.

In short: If AI bots are not stripped out of your analytics data, your data might no longer be meaningful or reliable.

AI crawlers also generate countless 404 errors and pointless requests. Because these bots often follow anything that vaguely looks like a URL (and sometimes even try out "hallucinated" paths), they can burden your servers with vast numbers of failed requests. Vercel found that over 34% of requests from the ChatGPT and Claude bots led to 404 errors. These failed accesses don’t always appear in analytics tools, because many of them are not captured by tracking setups. They can, however, be visible in log files and can fill up pages of crawl statistics in Google Search Console, for example.

For data quality, this means that even if you’re not seeing huge changes in your overall visitor numbers, AI bots can still lead you down the garden path to false conclusions at a page level. For example, your analytics might show frequent visits or high direct traffic to specific pages, when in reality nobody has ever seen that content.

 

 

Unreliable engagement metrics (bounce rate, session duration) 

AI traffic also heavily distorts traditional behavior analytics. Metrics such as bounce rate, session duration, or pages per session are key to measuring user interaction, but AI crawlers distort these metrics by behaving in a completely different way from human visitors:

Bounce rate: A bounce is normally a visit in which the user visits a single page and then leaves again. AI crawlers often request dozens of pages in a row. But because they generally don’t execute JavaScript tracking code, your analytics tool might not properly register these hits.

And if your analytics is measured exclusively via client-side JavaScript, such requests often don’t appear at all in tools like Google Analytics.

In other setups, such as server-side tracking, log file analysis, or inadequate bot filtering, bot requests can, by contrast, show up as isolated page views or unusual sessions. This can distort metrics such as bounce or engagement rates and create the false impression that real users are only briefly or superficially engaging with your content.

Session duration: average session duration is distorted similarly. Some bots race through several pages in milliseconds. The analytics tool can’t calculate any meaningful time on site and so often records the visit duration as 0 seconds, which drags your average down.

Other bots keep a connection open for a long time or keep reloading without any real interaction taking place. In some cases they stretch a "session" out to an extreme length — one that, from an analytics perspective, has no end, because no clean closing event arrives.

Metrics such as average time on page become unusable if a significant proportion of sessions show no genuine human navigation patterns at all.

Pages per session: While a human user might click through 2–5 pages per visit, a bot can speed run 50 pages in seconds or keep requesting the same pages over and over. Something like that can send your pages-per-session metric swinging up or down, taking you along for the ride.

Let’s take an example. If a bot requests every page once and disconnects immediately, your analytics will show huge numbers of one-page sessions and therefore bounces. But if it crawls the whole site sequentially, "monster sessions" arise with dozens of page views that a human would never reach. Taken together, the variance goes up, and the average of this metric loses its meaning.

All of this makes it harder to identify problems affecting real users:

  • If your bounce rate rises because of bots, you might wrongly assume that your landing page content is weak.
  • If average session duration falls, alarm bells may start ringing about user experience, when in truth technical bot activity is behind it.
  • Data quality suffers considerably, because without countermeasures you can no longer cleanly separate signal (human behavior) from noise (bot patterns).

 

Distorted conversion and performance metrics 

Crucially, your conversion rate data can also be devalued by AI traffic. The conversion rate is typically calculated from the number of conversions (e.g., purchases, form completions) divided by the number of sessions or users.

If the denominator (sessions) is then inflated by bot-generated visits, the conversion rate appears to fall. Your marketing ROI looks worse as a result, even though the actual human conversion figures remain unchanged.

For example, 100 real visitors with 5 orders would give a conversion rate of 5%. Add 900 bot sessions, which of course make no purchases, and the rate drops to 0.5%. Without filtering out bot traffic, you might incorrectly assume your campaign or website is performing miserably.

Ad analytics provider DoubleVerify has observed in the advertising context that so-called general invalid traffic (GIVT) — that is, recognizable bot impressions that don’t count as genuine ad views — rose by 86% in the second half of 2024. This increase is largely attributed to AI crawlers.

According to the same source, in 2024, 16% of all ad impressions identified as "known bot" traffic were attributable to bots associated with AI scraping (GPTBot, ClaudeBot, Applebot, etc.).

Transfer that to your web analytics, and it becomes clear: without cleaning up and filtering your data, 10–15% of your "visitors" might not be real visitors at all. That pushes every rate down — conversion, click-through, engagement.

Sometimes bots even trigger false conversions, for instance by technically clicking through a shopping cart process or firing tracking events, for example by loading a "thank you” page.

Modern AI agents can read out form fields or trigger certain API calls that your system counts as a conversion. Such cases are rarer than a general traffic overload, but they do occur — and then 0% conversion from bots suddenly becomes phantom conversions.

An analytics team might then see inflated success metrics, for example, an unusually high number of signups that never turn into real customers, because they came from a bot. Generally, these conversions wouldn’t be counted by sales systems since there’s no real customer behind them, but the analytics funnel is distorted just the same.

Finally, AI-based systems can also siphon off traffic without your analytics noticing. For example: a user asks ChatGPT a question, and the answer it provides is based on text from your website that ChatGPT reproduces from its training. The user has therefore made use of your content without ever visiting your site. These "invisible conversions” — in the form of knowledge transfer or brand impact — are not captured on any dashboard. Your website delivered value to the user and to the AI provider, but there’s nothing to see in your analytics: no page view, no visitor, and no session.

This is not bot traffic in the classic sense, but it is an indirect effect of AI on your usage statistics: fewer human visits despite unchanged demand for information.

In your web analytics tool, it looks as though your visibility is declining.

In reality, though, your content may be being picked up by AI answers; it’s just that the user no longer needs to click through to your website to get their answer. This phenomenon will make measuring content performance harder in the future, because you’re only measuring part of the actual reach.

Interim conclusion: AI crawlers and generative AI have negative effects on the quality of your data analytics on several levels, resulting in inflated visitor numbers, distorted engagement metrics, and unreliable conversion rates.

Decisions like budget allocation, website optimization measures, and campaign evaluations nowadays run the risk of being wrong if you base those decisions on raw data.

In our next blog post, we’ll therefore look at how you can detect, filter, and separately analyze AI traffic in order to make your analytics data meaningful again.

 

Summary: Securing data quality in the age of AI 

Whether from unscrupulous crawlers or from changed user behavior via chatbots, AI-generated traffic is here to stay. For us data and business analysts, that means we face a twofold task: taking protective measures to keep our web analytics data consistent and trustworthy, and at the same time recognizing the signs of the times and establishing new success metrics.

On the protection side, a combination of tool features (built-in bot filters), custom rules (user-agent/IP filters, segmentation) and, where applicable, technical measures (rate limiting, bot blocking via infrastructure) kit you out well to manage these new challenges. And as ever, clean, quality data remains essential to making well-founded decisions. Don’t allow AI bots to ruin your KPIs behind the scenes. Use the options available to make them visible and filter them out.

At the same time, new opportunities are opening up: by deliberately monitoring AI traffic, you can find out which content is particularly interesting to AI models — perhaps an indication of which topics are perceived as expert content. And through strategies such as pay-per-crawl or partnerships with AI providers, there could in future even be a return for the content provider. It’s still a way off, but the course is being set now.

In short: Securing data quality in the AI age is challenging, but achievable. In our next blog post, we’ll walk you through the “how”.

 

Are your web analytics not accurate anymore? Let’s work together to determine what percentage of your KPIs is attributable to AI traffic. We can show you how to ensure data quality and once again analyze your data reliably.

 

Coming soon:

Part 3: Filtering AI bot traffic: How to clean up your data in GA4, Adobe & Piano