You want to know which bots crawl your site: which AI crawlers, how often, and which pages they get a 404 on. A tool offers to show you, and the setup screen asks for a Cloudflare API token with Workers Scripts Edit. With that token it deploys its own Worker and sets up a route that runs it on every request to your site.
Ahrefs Bot Analytics works this way, and so do Known Agents (formerly Dark Visitors), Profound and Scrunch. The Worker is short and does what it says. The route in front of it is the part that changes how your site runs, and none of these setup screens mention the cost.
This article covers what that route does on Cloudflare’s Free plan, what Cloudflare already gives you without it, and how to log bots yourself in a Worker you already have. Every limit below was checked against Cloudflare’s documentation on 30 September 2026, and every traffic number comes from this site, 301.sh, which runs on the Free plan.
What the Worker does
Ahrefs asks for three permissions: Account Workers Scripts Edit, Zone Workers Routes Edit
and Zone Read. Its setup screen lists three steps: “Find the zone matching” your site,
“Deploy the Ahrefs analytics worker”, “Set up a route to run on all requests”. It adds:
“Your token is used once and not stored.” The manual setup in
its documentation uses
the route pattern *yourwebsite.com/*.
The Worker passes each request through unchanged with fetch(event.request). After the
response goes out, it sends a report inside event.waitUntil, so the report does not delay
the page. The report carries the full URL including the query string, the visitor’s IP from
cf-connecting-ip, the user agent, the referer, the method, the status and the content
type. It also carries three headers from Web Bot Auth: signature, signature-input and
signature-agent. Those are how a
signed AI agent proves who it is, and they are the most
interesting part of the payload.
Known Agents and Profound publish Workers of the same shape. Profound’s skips /checkout,
/cart, /admin and /api before it reports
(its docs). The Ahrefs one
reports everything.
What the route costs
Every request becomes a Worker invocation. The Workers Free plan allows 100,000 requests
a day, reset at midnight UTC
(limits). A route on
*example.com/* matches the page and every image, stylesheet, script and font the page
loads. A page with twenty files spends twenty one requests of that quota on one visit. The
report the Worker sends is a subrequest, and Cloudflare does not bill subrequests; the
invocation itself counts.
Static files stop being free. If your site is Workers static assets, Cloudflare says “Requests to static assets are free and unlimited”, and bills only the requests that invoke a Worker (billing). A route Worker in front makes every one of them invoke a Worker.
Here is the difference on 301.sh. In the seven days to 30 September the zone served about
19,200 requests, and our own Worker ran 2,421 times, because it runs only on page URLs and
lets the static files go out without it. A route on /* would have run a Worker on all
19,200, eight times as many. This site stays far below 100,000 a day either way. A site with
10,000 page views a day and twenty files a page would not.
The route runs before your own Worker. Cloudflare’s documentation on custom domains:
“Any Workers running on routes before your Custom Domain can optionally call the Worker
registered on your Custom Domain by issuing fetch(request)”
(custom domains).
So on a site that already runs its own Worker, the tool’s Worker is the first code that
sees each request. The documentation does not say how the pair is counted against the quota.
Know what happens at 100,000. Past the daily limit a route either fails open, which
“Bypasses the Worker. Requests behave as if no Worker is configured”, or fails closed, which
“Returns a Cloudflare 1027 error page”. The mode is a toggle on the route. The
documentation does not say which one a new route gets, and the route request in the Ahrefs
manual setup does not set it.
Open Workers & Pages, find the route the tool created, and check it. A route that fails
closed turns a traffic spike into an error page for your whole site.
A more specific route wins. “The most specific route pattern wins”
(routes). If you
already have a Worker on www.example.com/* or example.com/api/*, those requests go to
your Worker, and the analytics tool never sees them. Nothing tells you they are missing.
Every visitor’s data leaves. The route cannot tell a bot from a person before it runs, so the report goes out for people too: their IP, the page they opened, the query string with whatever your site puts in it. That makes the tool a processor of your visitors’ personal data, which your privacy policy may need to say.
The token is broader than the job. The Ahrefs documentation suggests a token on All accounts and All zones, and mentions narrowing it only as an alternative. Workers Scripts Edit on an account lets the holder upload any Worker to it. Create the token for the one zone the tool needs, and delete it once the Worker is running. “Not stored” is then something you do not have to take on trust.
What lifting the quota costs
The Workers Paid plan is $5 a month per account with 10 million requests included, then $0.30 per additional million (pricing). There is no daily limit on Paid, so the fail mode stops mattering. For most sites that settles the quota question for $5. It does not change the other items above: static files are still billed as invocations, and the data still leaves.
What Cloudflare already shows you
Before adding a Worker, look at what Cloudflare already records for the zone.
| Tool | Plan | What it shows |
|---|---|---|
| GraphQL Analytics API | all plans; on Free 31 days back, up to 30 days per query | every bot by user agent and Cloudflare’s verified category, with path and response status |
| AI Crawl Control | all plans; referral data on paid plans | AI crawlers by operator, the most crawled paths, response status codes |
| Security Analytics | all plans; the last seven days on Free | sampled individual requests with their properties and scores |
| Bot Analytics | Business and Enterprise | bot scores and verified bot traffic |
| Logpush | Enterprise | every request, sent to your storage |
AI Crawl Control is the dashboard answer for AI crawlers. The API answers for every bot, and it is the row we checked on our own zone.
Ask the API
On 301.sh, a Free zone, the settings node of the GraphQL API lists userAgent,
verifiedBotCategory, clientRequestPath and edgeResponseStatus among the fields of
httpRequestsAdaptiveGroups. That is enough to ask which verified bots got a 404 this week:
query ($zone: String!, $since: Time!, $until: Time!) {
viewer {
zones(filter: { zoneTag: $zone }) {
httpRequestsAdaptiveGroups(
limit: 50
filter: {
datetime_geq: $since
datetime_lt: $until
verifiedBotCategory_neq: ""
edgeResponseStatus: 404
}
orderBy: [count_DESC]
) {
count
dimensions { verifiedBotCategory userAgent clientRequestPath }
}
}
}
}
POST it to https://api.cloudflare.com/client/v4/graphql with a token that has Zone
Analytics Read, the zone ID and a time range in the variables. Drop the edgeResponseStatus
line to see all verified bot traffic. For the bots Cloudflare has not verified, filter on
verifiedBotCategory: "" and userAgent_like: "%bot%" instead.
On our zone, over the seven days to 30 September, about 1,400 of 19,200 requests came from
verified bots in nine categories. Search engine crawlers led, then AI crawlers, SEO tools and
aggregators. The 404 query found two things worth fixing or knowing. DuckDuckBot asked for
/favicon.ico thirteen times; the site had only /favicon.svg, and has both since this
query. AI crawlers asked for
addresses such as /2018-gopherconuk-slides that this site has never had.
The unverified query found something else. A client with the user agent
Claude-SearchBot/1.0 asked 36 times for /api/uploads/..%2F..%2F.env, a path that tries
to climb out of an upload folder to a file of secrets. Cloudflare did not verify it, and no
search crawler needs that path. A tool that counts by user agent alone files those 36
requests under Anthropic.
The category values are the dashboard names with spaces, such as "AI Crawler" and
"Search Engine Crawler". A filter on "AICrawler" returns nothing and no error.
What the API does not give you is history past 31 days, or the Web Bot Auth headers. For those you need your own log.
Log bots in the Worker you already have
If your site already runs a Worker for its pages, add the log there. The request already
invokes that Worker, so the log costs no extra invocation, and you decide what leaves your
account. 301.sh is built this way: its Worker runs first on page URLs only, set with
run_worker_first in wrangler.toml, and images and styles are served as free static
assets without it.
Workers Analytics Engine is on the Free plan: 100,000 data points written and 10,000 read queries a day (pricing). The Workers Paid plan at $5 a month raises that to 10 million data points a month. Bind a dataset:
# wrangler.toml
[[analytics_engine_datasets]]
binding = "BOTS"
dataset = "bot_hits"
Then record bots on the way out:
const BOT = /bot|crawl|spider|slurp|GPTBot|ClaudeBot|PerplexityBot|Bytespider|CCBot|meta-externalagent/i;
export default {
async fetch(request, env, ctx) {
const response = await handle(request, env, ctx); // your existing logic
const ua = request.headers.get("user-agent") ?? "";
const verified = request.cf?.verifiedBotCategory ?? "";
if (verified || BOT.test(ua)) {
env.BOTS.writeDataPoint({
indexes: [verified || "unverified"],
blobs: [
ua.slice(0, 256),
new URL(request.url).pathname,
String(response.status),
request.headers.get("signature-agent") ?? "",
],
doubles: [1],
});
}
return response;
},
};
Three details matter.
Only bots are written. People’s requests never reach the dataset, so there is nothing personal in it to disclose, and the 100,000 daily data points go to the traffic you want to study.
verifiedBotCategory is Cloudflare’s check, the user agent is the bot’s claim. Anyone
can send GPTBot in a header. Cloudflare fills request.cf.verifiedBotCategory only for
bots it has verified. Write both and compare.
writeDataPoint is not awaited. It returns at once and the runtime stores the point in
the background, the same as in
counting clicks on a redirect, which also covers the read
side: one SQL query, and SUM(_sample_interval) rather than COUNT().
Which bots get a 404:
SELECT index1 AS category, blob1 AS agent, blob2 AS path,
SUM(_sample_interval) AS hits
FROM bot_hits
WHERE blob3 = '404' AND timestamp >= NOW() - INTERVAL '7' DAY
GROUP BY category, agent, path
ORDER BY hits DESC
LIMIT 50
If your site has no Worker today, adding one only for this brings back the quota question.
You can still narrow it: a route on the page paths rather than /* leaves images and styles
alone.
Where this breaks
The regex decides what a bot is. A crawler that sends a browser user agent is not in your dataset unless Cloudflare verified it. Every tool that reads the user agent has the same limit, the paid ones included.
Analytics Engine keeps data for three months. For a year of history, export it on a schedule.
The API numbers are estimates. Cloudflare samples high volume and “returns an estimate derived from the sampled value” (sampling). Our week came back with an average sample interval of about 2.5. Good for which bots and which paths; not for exact counts of a rare event.
Thirty one days is all the API keeps on Free. Save a weekly result if you want a trend.
Web Bot Auth headers are rare today. Recording signature-agent costs nothing, and the
column will fill as more agents sign their requests. Do not read an empty column as proof
that no agents visit.
When 301.st is worth it
For one site you do not need us. Run the query above once a week, and if you want a longer history, add the log to your Worker.
301.st manages redirects and traffic rules across a portfolio of domains. On the domains where you give it rules, it runs a Worker in your Cloudflare account, so this is the same trade described above, taken for routing the traffic rather than only watching it. That Worker sorts each request into seven bot categories, from search engines and AI crawlers to link previews and uptime monitors, using the user agent and Cloudflare’s verified bot flag. A rule checks those categories before any other rule on the domain: AI Guard blocks or diverts AI training crawlers and leaves search engines alone, and Bot Shield stops every bot on a domain that should never be crawled.
A screen that shows those bots per domain, with paths and response codes, is what we are building next.