Don’t close the door on useful AI bots — who’s visiting your website?

Details
In the summer of 2026, something happened online that wasn’t expected to occur until later: web traffic generated by AI bots and other bots has repeatedly overtook human traffic. This means websites are now more often visited by machines than by people.
But is this bot traffic to the website useful, or does it just waste bandwidth? And how can you distinguish bots that actually bring customers, like Google’s search bot?
AI bots are lining up at your website’s door
When we try to fetch, with an AI tool, the NBC News piece that served as the source for this article, the site blocks us. Search engines get in; the AI tool doesn’t. This small example says a lot: online, there’s an ongoing negotiation about which machines are allowed in and which are not.
This negotiation matters, because there are now more machines than human visitors. Cloudflare’s data shows bots accounted for about 57% of the requests to sites on its network in June, and humans 43%. The pace of bot traffic growth even surprised Cloudflare’s CEO, who had estimated the turning point would come as late as 2027.
You shouldn’t block all bots; overly strict blocking can also cause your page to be dropped from Google.

Not all bots are the same
Some bot traffic is pure waste. According to Cloudflare, over half of requests from even reputable bots target pages that haven’t changed since the last visit. That needlessly strains servers.
Still, some bots bring people to the site who buy or are genuinely interested in the service. According to data collected by Adobe from U.S. online retailers, visitors arriving via AI assistants buy more often than others. This is because when a person asks an AI for a recommendation, the site mentioned in the answer is likely to get a visitor who is already well along in their purchase decision. Read more about how AI is changing search behavior (in Finnish).
“A good bot strategy doesn’t shut doors. It decides who they’re opened to and on what terms.”
Three bot types, three different decisions
Cloudflare classifies AI traffic into three categories based on what a bot does on the site, and the site owner can decide how to handle each.
| Bot type | What it does | Recommended approach |
|---|---|---|
| Search | Indexes content so it can later answer questions and cite the source | Let it in |
| Agent | Acts on a person’s behalf in real time: searches, compares, buys | Treat it like a customer |
| Training | Collects content for training AI models | Limit with care |
Search bots: let them in. They index content so AI can later answer users’ questions and cite sources. They also bring visitors to your site, so it’s generally best to allow them.
Agent bots: treat them like customers. Agents act in real time on behalf of a person: they search, compare and increasingly buy. A real person is behind them. Read how agent-driven buying is changing the e-commerce purchase journey.
Training bots: limit them thoughtfully. They collect content to train AI models. Blocking them can be justified because the benefit to your site is indirect. On the other hand, appearing in training data can influence how models understand your brand. The decision belongs to the business, not just the firewall.
Check where the block applies — Google may be left out, too
Some of the largest bots fall into several categories at once: the same crawler fetches content both for search and for AI training. That’s why a rule that restricts Training bots can also hit search engines like Googlebot, Bingbot and Applebot. A well-intentioned protection setting can therefore reduce your site’s visibility in traditional search results.
On Cloudflare sites, rules are activated in two ways. Existing customers have been able, since July 2026, to choose whether they allow Search, Agent and Training bots. For new customers, default settings took effect in September 2026: on pages that display ads, Training and Agent bots are automatically blocked and Search is allowed.
Cloudflare treats multi-purpose crawlers by the strictest rule. If you choose “Block Training”, the block will also apply to crawlers that fetch content for search engines.
The same logic applies to other services, such as Akamai, Fastly or your own firewall. If you change bot-blocking settings, always make sure which bots the block actually hits.
From strategy to practice
Once rules for different bot types have been set, they can be described in the site’s robots.txt file. However, remember that robots.txt is a guideline, not an actual gatekeeper: not all bots follow it.
Real enforcement happens elsewhere. More granular access control is implemented using WAF rules and rate limiting. Authenticity is verified via lists of verified bots or by checking logs to ensure the request comes from the bot’s official IP address.
Freshness signals cut unnecessary load. When the site tells machines what has changed (for example with the IndexNow protocol and up-to-date metadata), bots fetch the right pages instead of repeatedly requesting unchanged ones. A well-functioning cache handles the rest.
Measure separately. People, search engines, AI search, training bots and agents are different user groups. The most interesting metric is often how many times a bot requests pages compared to how many visitors it ultimately directs to the site.
Want to know who’s really visiting your site? Get in touch and we’ll find out together.