Creuto is now an OpenAI Select Partner Read More
The agentic web, in shipped mechanisms: the AI crawler user agents in your logs, crawl-to-referral ratios, what llms.txt really does, and what to change.

More than half of internet traffic is no longer human. On the agentic web, your second audience fetches a page, lifts the three facts it came for, and leaves without a click, an ad impression or a session. Cloudflare says daily requests from AI agents on its network grew by more than 1,700% in a year. This is what to change, and what not to bother with.
Almost everything written about this is vision. What follows is the shipped part: the crawlers hitting you today, the numbers vendors publish about what they give back, which emerging standards have real adoption, and a short ranked list of changes worth making.
Cloudflare's 30 September 2026 post is the clearest statement of the commercial problem. Its network went from an average of 63 million HTTP requests a second at the end of 2024 to almost 115 million, with peaks above 150 million. In spring 2025, 22% of crawler requests it saw were for AI training by the crawlers' own stated purpose; by June 2026 that was 52%.
The consequence is not abstract. Cloudflare reports that heavily crawled categories — retail, computer software, IT services and financial services — "have seen human traffic decline as much as 40% in less than one year". The thirty-year arrangement where being found and getting paid were the same event has come apart.
Worth noting what site owners actually chose when given the controls. After Cloudflare split its single AI switch into separate search, agent and training controls in July 2026, fewer than 1% of sites blocked search crawlers while 17% blocked training. Nobody is trying to hide. They want to be found without being used for free.
Start here, because it is checkable tonight. These are the user agents the vendors document, as of 1 October 2026.
| User agent | Operator | What it is for |
|---|---|---|
| GPTBot | OpenAI | Crawls content for training generative models. Blocking it signals your content should not train foundation models. |
| OAI-SearchBot | OpenAI | Surfaces sites in ChatGPT search results. Opting out removes you from ChatGPT search answers. |
| ChatGPT-User | OpenAI | User-initiated fetches only. OpenAI states it "is not used for crawling the web in an automatic fashion". |
| ClaudeBot | Anthropic | Collects web content for model training. |
| Claude-User | Anthropic | Fetches pages when a person asks Claude something. Blocking it reduces visibility in Claude's web search. |
| Claude-SearchBot | Anthropic | Indexes to improve search result quality. |
| PerplexityBot, Meta-ExternalAgent, Amazonbot, Bytespider, Applebot | Various | Observed by Cloudflare across its network in its crawler analysis. |
The practical point is the split inside a single vendor. OpenAI documents three agents and Anthropic documents three, and in both cases training, search indexing and user-initiated retrieval are separate names. A blanket Disallow on the training bot does not remove you from the assistant's answers. A blanket block on everything does.
Run the check before you argue about strategy. Grep a week of access logs for those strings. Most teams we work with are surprised twice: by how much of it there is, and by how little of it is the bot they were worried about.
Cloudflare defines the crawl-to-refer ratio as "how many pages a platform crawls compared with how often it drives users to a website". In its July 2025 figures, the ratios were Anthropic 38,065.7, OpenAI 1,091.4, Perplexity 194.8 and Google 5.4 — Google being the only one anywhere near the old bargain.
Two honest caveats. These are one vendor's measurements of its own network, not an industry audit. And they move fast: Cloudflare recorded Anthropic's ratio falling 86.7% between January and July 2025 while Perplexity's rose 256.7% over the same window. Treat the direction as the signal and any single figure as dated.
What does not move is the structural point. A crawler that reads your page and returns a summary consumes your bandwidth and origin capacity and returns no referral, no ad impression and no subscription. If your business model assumes a visit follows a fetch, that assumption now fails for a growing share of your traffic. We covered the blunt version of this response in our post on the Cloudflare AI crawler block and whether it hides you from AI search.
This is where most agentic web advice goes wrong, so take the measurement over the enthusiasm. Ahrefs studied 137,210 domains in May 2026. Twenty-eight per cent published a valid llms.txt. Ninety-seven per cent of those files received zero requests in the study period. Of the requests that did arrive, 96% came from bots, and the largest categories were SEO audit tools (21.7%), unidentified bots (14.9%) and general web crawlers (13.1%). AI retrieval bots accounted for 1.1% of total requests.
Google is explicit. Its AI features documentation states there are "no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary", and that "you don't need to create new machine readable files, AI text files, or markup to appear in these features". That is Google saying, in its own words, that it does not use llms.txt.
The strongest case for publishing one anyway: it costs an afternoon, it does no harm, and standards sometimes get adopted after they look dead. That case is fair. It is also an argument for doing it last, not first, and for not telling a client it will improve AI visibility — on the published evidence, it will not.
The strongest objection to this entire topic comes from Google itself: "There's also no special schema.org structured data that you need to add." If the largest answer engine says normal SEO is the whole job, why do anything?
Because that sentence answers a narrower question than it appears to. It says you need no new schema type invented for AI. It does not say the facts on your page are equally extractable in every form. The same page that renders a price in a styled span and the same page that also emits a Product offer in JSON-LD present identical pixels to a person and very different structures to a parser. Google's own guidance in the same document asks you to ensure "your structured data matches the visible text on the page" — which only means something if the structured data is being read.
Structured data is also the lowest-risk work on this list, because it was already justified for search. That is the test to apply to everything here: would you still do it if the agentic web turned out to be overstated? If the honest answer is no, it goes at the bottom. If you are emitting JSON-LD in a React app, the one thing to get right first is escaping — we wrote up the JSON-LD script injection escape most sites skip.
Today, overwhelmingly HTML, because that is what exists. That is also why the extraction is lossy: an agent asked for your pricing parses a page built to persuade, and discards the hero video, the nav, the testimonial and the call to action. Every layout decision made for a human is noise to it.
The proposals aiming at this are real but early. WebMCP lets a browser agent call your web app's tools instead of clicking through its interface. Web Bot Auth takes a different angle: Cloudflare says operators including OpenAI, Google and AWS cryptographically sign their agents' requests, and that it sees more than 500 billion verified bot requests a week — which makes identity, not the user-agent string, the thing you gate on. On payments, Cloudflare's Pay Per Use and its x402-based Monetization Gateway are the commercial layer, and there is more than one competing design; we compared two of them in Agentic Commerce Protocol and UCP are not the same job.
None of these is a reason to build an agent API this quarter unless you already have an API product. Build one when a named customer's agent needs it, not in anticipation.
| Change | Effort | Payoff |
|---|---|---|
| Grep your logs for the documented AI user agents and count them by path | An hour | Highest. Every decision below depends on knowing your own numbers. |
| Split robots.txt by purpose: training, search, user-initiated | An hour | High. The single on/off switch is the mistake 17% of Cloudflare sites made differently from the 1% blocking search. |
| Put the facts an agent came for in text, not only in images, video or script-rendered DOM | Days | High. Extraction fails silently; nobody files a bug. |
| Correct JSON-LD matching the visible text, on the page types that carry facts | Days | Medium-high, and already justified by search alone. |
| Server-render the pages you want quoted | Days to weeks | Medium. Varies by crawler; measure before assuming. |
| Bot identity verification rather than user-agent string matching | Weeks, or a vendor | Medium. Rises with how much you intend to charge or gate. |
| Publish llms.txt | An afternoon | Low on current evidence. Do it last. |
| Build an agent-facing API or MCP server | Weeks | Low until a named customer asks. Then high. |
The first four are things we would have recommended for search in 2023 and would still recommend if every agent vanished tomorrow. That is deliberate. The honest version of this advice is that most of the work is unchanged — it is the justification that got stronger, not the task list.
Where it genuinely is new is accounting. Your analytics counts sessions, and a growing share of your audience never starts one. Until you can answer "how many of our fetches came from agents, and on which pages", you are making web application decisions against a number that no longer describes your traffic. Start the log count this week; the rest can wait for what it tells you.
The agentic web is the web as it is now used by software acting for people, not only by people. Cloudflare reported in September 2026 that more than half of internet traffic is non-human and that daily AI agent requests on its network grew over 1,700% in a year.
Today almost entirely HTML, because that is what sites publish. An agent parses a page built to persuade a human and keeps only the facts it needs, discarding layout, imagery and calls to action. Agent-callable interfaces such as WebMCP exist but have limited adoption so far.
On current evidence, not as a priority. Ahrefs found in May 2026 that 97% of llms.txt files across 137,210 domains received zero requests, and Google states you do not need AI text files to appear in its AI features. It costs little, so do it last.
Count AI crawler hits in your logs first, then split robots.txt by purpose rather than using one switch, put the facts agents want in plain text rather than images or script-rendered DOM, and emit JSON-LD that matches the visible page. Those four cover most of the value.
Cloudflare defines it as how many pages a platform crawls compared with how often it sends users to a site. In its July 2025 figures the ratios were Anthropic 38,065.7, OpenAI 1,091.4, Perplexity 194.8 and Google 5.4. Treat specific figures as dated; they move quickly.
Rarely. Training, search indexing and user-initiated retrieval use separate user agents at both OpenAI and Anthropic, so a blanket block removes you from assistant answers as well as training sets. Fewer than 1% of Cloudflare sites block search crawlers, while 17% block training.
Google says no special schema.org type is needed for its AI features, and that your structured data should match the visible text. Correct JSON-LD still makes facts unambiguous to any parser and was already justified by ordinary search, which makes it low-risk work either way.
Ready to take the first step towards unlocking opportunities, realizing goals, and embracing innovation? We're here and eager to connect.
11th Floor, O-Hub, Chandaka Industrial Estate, Infocity, Bhubaneswar, Odisha 751024
Level 4, 11 York Street Sydney Startup Hub Sydney, NSW – 2000
30 N. Đinh Nghệ, Phước Mỹ Sơn Trà, Đà Nẵng / Da Nang City – 550000
Level 25, AIDP Business Tower, Dubai Marina, United Arab Emirates
50 Beauchamp Street, Wellington, WGN 5028, New Zealand