You want to know which bots crawl your site: which AI crawlers, how often, and which pages they get a 404 on. A tool offers to show you, and the setup screen asks for a Cloudflare API token with Workers Scripts Edit. With that token it deploys its own Worker and sets up a route that runs it on every request to your site.

Ahrefs Bot Analytics works this way, and so do Known Agents (formerly Dark Visitors), Profound and Scrunch. The Worker is short and does what it says. The route in front of it is the part that changes how your site runs, and none of these setup screens mention the cost.

This article covers what that route does on Cloudflare’s Free plan, what Cloudflare already gives you without it, and how to log bots yourself in a Worker you already have. Every limit below was checked against Cloudflare’s documentation on 30 September 2026, and every traffic number comes from this site, 301.sh, which runs on the Free plan.

What the Worker does

Ahrefs asks for three permissions: Account Workers Scripts Edit, Zone Workers Routes Edit and Zone Read. Its setup screen lists three steps: “Find the zone matching” your site, “Deploy the Ahrefs analytics worker”, “Set up a route to run on all requests”. It adds: “Your token is used once and not stored.” The manual setup in its documentation uses the route pattern *yourwebsite.com/*.

The Worker passes each request through unchanged with fetch(event.request). After the response goes out, it sends a report inside event.waitUntil, so the report does not delay the page. The report carries the full URL including the query string, the visitor’s IP from cf-connecting-ip, the user agent, the referer, the method, the status and the content type. It also carries three headers from Web Bot Auth: signature, signature-input and signature-agent. Those are how a signed AI agent proves who it is, and they are the most interesting part of the payload.

Known Agents and Profound publish Workers of the same shape. Profound’s skips /checkout, /cart, /admin and /api before it reports (its docs). The Ahrefs one reports everything.

What the route costs

fetch(request)

Workers with a custom domain

Static assets alone

An origin server

waitUntil

Any request:
page, image, CSS

Route Worker
*example.com/*
counts toward the quota

What serves
the site?

Your own Worker
runs after the route

Assets
free only when
no Worker ran

Origin

Report to
the tool

The request path with a route Worker in front

Every request becomes a Worker invocation. The Workers Free plan allows 100,000 requests a day, reset at midnight UTC (limits). A route on *example.com/* matches the page and every image, stylesheet, script and font the page loads. A page with twenty files spends twenty one requests of that quota on one visit. The report the Worker sends is a subrequest, and Cloudflare does not bill subrequests; the invocation itself counts.

Static files stop being free. If your site is Workers static assets, Cloudflare says “Requests to static assets are free and unlimited”, and bills only the requests that invoke a Worker (billing). A route Worker in front makes every one of them invoke a Worker.

Here is the difference on 301.sh. In the seven days to 30 September the zone served about 19,200 requests, and our own Worker ran 2,421 times, because it runs only on page URLs and lets the static files go out without it. A route on /* would have run a Worker on all 19,200, eight times as many. This site stays far below 100,000 a day either way. A site with 10,000 page views a day and twenty files a page would not.

The route runs before your own Worker. Cloudflare’s documentation on custom domains: “Any Workers running on routes before your Custom Domain can optionally call the Worker registered on your Custom Domain by issuing fetch(request)” (custom domains). So on a site that already runs its own Worker, the tool’s Worker is the first code that sees each request. The documentation does not say how the pair is counted against the quota.

Know what happens at 100,000. Past the daily limit a route either fails open, which “Bypasses the Worker. Requests behave as if no Worker is configured”, or fails closed, which “Returns a Cloudflare 1027 error page”. The mode is a toggle on the route. The documentation does not say which one a new route gets, and the route request in the Ahrefs manual setup does not set it. Open Workers & Pages, find the route the tool created, and check it. A route that fails closed turns a traffic spike into an error page for your whole site.

A more specific route wins. “The most specific route pattern wins” (routes). If you already have a Worker on www.example.com/* or example.com/api/*, those requests go to your Worker, and the analytics tool never sees them. Nothing tells you they are missing.

Every visitor’s data leaves. The route cannot tell a bot from a person before it runs, so the report goes out for people too: their IP, the page they opened, the query string with whatever your site puts in it. That makes the tool a processor of your visitors’ personal data, which your privacy policy may need to say.

The token is broader than the job. The Ahrefs documentation suggests a token on All accounts and All zones, and mentions narrowing it only as an alternative. Workers Scripts Edit on an account lets the holder upload any Worker to it. Create the token for the one zone the tool needs, and delete it once the Worker is running. “Not stored” is then something you do not have to take on trust.

What lifting the quota costs

The Workers Paid plan is $5 a month per account with 10 million requests included, then $0.30 per additional million (pricing). There is no daily limit on Paid, so the fail mode stops mattering. For most sites that settles the quota question for $5. It does not change the other items above: static files are still billed as invocations, and the data still leaves.

What Cloudflare already shows you

Before adding a Worker, look at what Cloudflare already records for the zone.

Tool Plan What it shows
GraphQL Analytics API all plans; on Free 31 days back, up to 30 days per query every bot by user agent and Cloudflare’s verified category, with path and response status
AI Crawl Control all plans; referral data on paid plans AI crawlers by operator, the most crawled paths, response status codes
Security Analytics all plans; the last seven days on Free sampled individual requests with their properties and scores
Bot Analytics Business and Enterprise bot scores and verified bot traffic
Logpush Enterprise every request, sent to your storage

AI Crawl Control is the dashboard answer for AI crawlers. The API answers for every bot, and it is the row we checked on our own zone.

Ask the API

On 301.sh, a Free zone, the settings node of the GraphQL API lists userAgent, verifiedBotCategory, clientRequestPath and edgeResponseStatus among the fields of httpRequestsAdaptiveGroups. That is enough to ask which verified bots got a 404 this week:

query ($zone: String!, $since: Time!, $until: Time!) {
  viewer {
    zones(filter: { zoneTag: $zone }) {
      httpRequestsAdaptiveGroups(
        limit: 50
        filter: {
          datetime_geq: $since
          datetime_lt: $until
          verifiedBotCategory_neq: ""
          edgeResponseStatus: 404
        }
        orderBy: [count_DESC]
      ) {
        count
        dimensions { verifiedBotCategory userAgent clientRequestPath }
      }
    }
  }
}

POST it to https://api.cloudflare.com/client/v4/graphql with a token that has Zone Analytics Read, the zone ID and a time range in the variables. Drop the edgeResponseStatus line to see all verified bot traffic. For the bots Cloudflare has not verified, filter on verifiedBotCategory: "" and userAgent_like: "%bot%" instead.

On our zone, over the seven days to 30 September, about 1,400 of 19,200 requests came from verified bots in nine categories. Search engine crawlers led, then AI crawlers, SEO tools and aggregators. The 404 query found two things worth fixing or knowing. DuckDuckBot asked for /favicon.ico thirteen times; the site had only /favicon.svg, and has both since this query. AI crawlers asked for addresses such as /2018-gopherconuk-slides that this site has never had.

The unverified query found something else. A client with the user agent Claude-SearchBot/1.0 asked 36 times for /api/uploads/..%2F..%2F.env, a path that tries to climb out of an upload folder to a file of secrets. Cloudflare did not verify it, and no search crawler needs that path. A tool that counts by user agent alone files those 36 requests under Anthropic.

The category values are the dashboard names with spaces, such as "AI Crawler" and "Search Engine Crawler". A filter on "AICrawler" returns nothing and no error.

What the API does not give you is history past 31 days, or the Web Bot Auth headers. For those you need your own log.

Log bots in the Worker you already have

If your site already runs a Worker for its pages, add the log there. The request already invokes that Worker, so the log costs no extra invocation, and you decide what leaves your account. 301.sh is built this way: its Worker runs first on page URLs only, set with run_worker_first in wrangler.toml, and images and styles are served as free static assets without it.

Workers Analytics Engine is on the Free plan: 100,000 data points written and 10,000 read queries a day (pricing). The Workers Paid plan at $5 a month raises that to 10 million data points a month. Bind a dataset:

# wrangler.toml
[[analytics_engine_datasets]]
binding = "BOTS"
dataset = "bot_hits"

Then record bots on the way out:

const BOT = /bot|crawl|spider|slurp|GPTBot|ClaudeBot|PerplexityBot|Bytespider|CCBot|meta-externalagent/i;

export default {
  async fetch(request, env, ctx) {
    const response = await handle(request, env, ctx); // your existing logic

    const ua = request.headers.get("user-agent") ?? "";
    const verified = request.cf?.verifiedBotCategory ?? "";
    if (verified || BOT.test(ua)) {
      env.BOTS.writeDataPoint({
        indexes: [verified || "unverified"],
        blobs: [
          ua.slice(0, 256),
          new URL(request.url).pathname,
          String(response.status),
          request.headers.get("signature-agent") ?? "",
        ],
        doubles: [1],
      });
    }
    return response;
  },
};

Three details matter.

Only bots are written. People’s requests never reach the dataset, so there is nothing personal in it to disclose, and the 100,000 daily data points go to the traffic you want to study.

verifiedBotCategory is Cloudflare’s check, the user agent is the bot’s claim. Anyone can send GPTBot in a header. Cloudflare fills request.cf.verifiedBotCategory only for bots it has verified. Write both and compare.

writeDataPoint is not awaited. It returns at once and the runtime stores the point in the background, the same as in counting clicks on a redirect, which also covers the read side: one SQL query, and SUM(_sample_interval) rather than COUNT().

Which bots get a 404:

SELECT index1 AS category, blob1 AS agent, blob2 AS path,
       SUM(_sample_interval) AS hits
FROM bot_hits
WHERE blob3 = '404' AND timestamp >= NOW() - INTERVAL '7' DAY
GROUP BY category, agent, path
ORDER BY hits DESC
LIMIT 50

If your site has no Worker today, adding one only for this brings back the quota question. You can still narrow it: a route on the page paths rather than /* leaves images and styles alone.

Where this breaks

The regex decides what a bot is. A crawler that sends a browser user agent is not in your dataset unless Cloudflare verified it. Every tool that reads the user agent has the same limit, the paid ones included.

Analytics Engine keeps data for three months. For a year of history, export it on a schedule.

The API numbers are estimates. Cloudflare samples high volume and “returns an estimate derived from the sampled value” (sampling). Our week came back with an average sample interval of about 2.5. Good for which bots and which paths; not for exact counts of a rare event.

Thirty one days is all the API keeps on Free. Save a weekly result if you want a trend.

Web Bot Auth headers are rare today. Recording signature-agent costs nothing, and the column will fill as more agents sign their requests. Do not read an empty column as proof that no agents visit.

When 301.st is worth it

For one site you do not need us. Run the query above once a week, and if you want a longer history, add the log to your Worker.

301.st manages redirects and traffic rules across a portfolio of domains. On the domains where you give it rules, it runs a Worker in your Cloudflare account, so this is the same trade described above, taken for routing the traffic rather than only watching it. That Worker sorts each request into seven bot categories, from search engines and AI crawlers to link previews and uptime monitors, using the user agent and Cloudflare’s verified bot flag. A rule checks those categories before any other rule on the domain: AI Guard blocks or diverts AI training crawlers and leaves search engines alone, and Bot Shield stops every bot on a domain that should never be crawled.

A screen that shows those bots per domain, with paths and response codes, is what we are building next.