Original research · Updated Aug 10, 2026

The AI Crawler Observatory

Retrieval crawl rose +594.9%. Attributed AI visits stayed at 0.

Nice Pick now tracks crawl, search visibility, citations, and human referrals as separate signals. The latest edge snapshot counted 2,300 retrieval-class requests. The current citation result is unavailable because the last probe is stale.

Current verdict

Fetching improved. Distribution has not followed.

Comparable seven-day edge snapshots show retrieval-class requests rising from 331 in July to 2,300 now. Training crawl rose much more slowly, and the training-to-retrieval ratio narrowed from 769:1 to 134:1. Cloudflare Analytics Engine recorded 0attributed AI arrivals in the latest seven days. The most recent citation batch is 38days old, so it cannot answer whether citation performance changed.

Latest scorecard

Each number keeps its own source and reporting window. None stands in for another.

Zone-wide retrieval fetches
2,300

Aug 3, 2026 to Aug 10, 2026

Change from July baseline
+594.9%

Same seven-day method

Attributed AI arrivals
0

Latest seven days

Current citation rate
Unknown

Probe is 38 days old

Two comparable edge snapshots

Both snapshots use Cloudflare zone-wide adaptive request groups, one-day query slices, the same user-agent classifier, and seven-day windows.

Comparable July and August Cloudflare edge request snapshots
MeasureJuly baselineLatestChange
All edge requests513,534846,171up 64.8%
Training-class requests254,595307,145up 20.6%
Retrieval-class requests3312,300up 594.9%
Search-engine crawler requests1,9263,576up 85.7%
Training requests per retrieval request769:1134:1Narrower
HTTP 5xx responses5,223150.002% rate

Retrieval actors in the latest snapshot

OAI-SearchBot supplied most retrieval-class requests. A request can make a page eligible for an answer. It cannot prove selection or citation.

OAI-SearchBot (ChatGPT search)1,856 requests · 0 5xx
ChatGPT-User (live user fetch)343 requests · 0 5xx
PerplexityBot101 requests · 0 5xx

The two retrieval counts answer different questions

Zone-wide edge audit

2,300

Requests across every route and response status seen by Cloudflare. This is the broad infrastructure view.

Content-route telemetry

33

Matched content middleware routes and explicit server logging. This narrower series fell -83.9% against the preceding seven days.

The scopes are deliberately published side by side. Subtracting one from the other would mix routing coverage, assets, redirects, and response classes into a fake conversion rate.

Crawl has not become measurable distribution

Search visibility is moving, citations are currently unmeasured, and attributed AI referrals remain scarce.

Current search, retrieval, citation, and AI referral signals
SignalCurrent evidenceWhat it proves
Google search visibility8,812 impressions, 2 clicks, average position 61.8Jul 12, 2026 to Aug 8, 2026Google surfaced Nice Pick. Impressions do not establish a visit or an answer-engine citation.
Persisted citationsCurrent performance unavailableLatest batch: 0/11, observed Jul 3, 2026The old batch establishes a baseline. Its age prevents a current citation claim.
Attributed AI referrals0 in 7 days; 3 in 30 daysA recognized answer product sent a browser visit. Attribution is heuristic and separate from crawler identity.

Methodology

The edge series queries Cloudflare HTTP Requests Adaptive Groups in one-day slices, then groups results by user-agent identity and HTTP response status. The public snapshot removes raw unknown user-agent strings and retains aggregate actor classes and recognized crawler names.

The content-route series comes from Cloudflare Analytics Engine and uses SUM(_sample_interval) to reconstruct sampled counts. Search visibility comes from exact, consecutive 28-day Google Search Console windows with a two-day reporting lag. Citation status reads persisted probe rows and marks batches older than 14 days stale. AI referrals are aggregate edge events from the hardened post-July 25 attribution series.

Crawler roles follow the operators' published descriptions. OpenAI documents OAI-SearchBot for search and GPTBot for model training. Anthropic and Perplexity publish separate crawler controls for their products. These descriptions guide classification; they do not reveal what happened after a request.

Primary references: OpenAI crawlers, Anthropic crawler controls, Perplexity crawlers, and Cloudflare Analytics datasets.

Limits

  • Cloudflare HTTP Requests Adaptive Groups are analytics estimates, not raw access logs.
  • Crawler class totals cover up to the top 300 user-agent groups in each daily query; the overall edge total uses the zone-wide aggregate.
  • User agents can be spoofed, and actor identity does not prove how fetched content was used.
  • Zone-wide retrieval and content-route retrieval have different scopes and must not be compared as one series.
  • Requests are not unique pages, people, referrals, recommendations, or citations.
  • Google Search Console data is delayed by about two days.
  • A stale citation probe means current citation performance is unavailable, not zero.
  • AI referral attribution is heuristic and uses the post-2026-07-25 hardened series.
  • This is one website. Its traffic mix should not be generalized to the web.

Use the data

The public JSON contains both comparable edge snapshots, source windows, actor totals, definitions, current search evidence, citation freshness, referral counts, and stated limits. The original July data file remains available for audit history.

Suggested citation

Nice Pick. “AI Crawler Observatory: Crawl, Citation, and Referral Data.” Updated Aug 10, 2026. https://nicepick.dev/research/ai-crawler-reality-report