The AI Crawler Observatory
Retrieval crawl rose +594.9%. Attributed AI visits stayed at 0.
Nice Pick now tracks crawl, search visibility, citations, and human referrals as separate signals. The latest edge snapshot counted 2,300 retrieval-class requests. The current citation result is unavailable because the last probe is stale.
Current verdict
Fetching improved. Distribution has not followed.
Comparable seven-day edge snapshots show retrieval-class requests rising from 331 in July to 2,300 now. Training crawl rose much more slowly, and the training-to-retrieval ratio narrowed from 769:1 to 134:1. Cloudflare Analytics Engine recorded 0attributed AI arrivals in the latest seven days. The most recent citation batch is 38days old, so it cannot answer whether citation performance changed.
Latest scorecard
Each number keeps its own source and reporting window. None stands in for another.
- Zone-wide retrieval fetches
- 2,300
- Change from July baseline
- +594.9%
- Attributed AI arrivals
- 0
- Current citation rate
- Unknown
Aug 3, 2026 to Aug 10, 2026
Same seven-day method
Latest seven days
Probe is 38 days old
Two comparable edge snapshots
Both snapshots use Cloudflare zone-wide adaptive request groups, one-day query slices, the same user-agent classifier, and seven-day windows.
| Measure | July baseline | Latest | Change |
|---|---|---|---|
| All edge requests | 513,534 | 846,171 | up 64.8% |
| Training-class requests | 254,595 | 307,145 | up 20.6% |
| Retrieval-class requests | 331 | 2,300 | up 594.9% |
| Search-engine crawler requests | 1,926 | 3,576 | up 85.7% |
| Training requests per retrieval request | 769:1 | 134:1 | Narrower |
| HTTP 5xx responses | 5,223 | 15 | 0.002% rate |
Retrieval actors in the latest snapshot
OAI-SearchBot supplied most retrieval-class requests. A request can make a page eligible for an answer. It cannot prove selection or citation.
The two retrieval counts answer different questions
Zone-wide edge audit
2,300
Requests across every route and response status seen by Cloudflare. This is the broad infrastructure view.
Content-route telemetry
33
Matched content middleware routes and explicit server logging. This narrower series fell -83.9% against the preceding seven days.
The scopes are deliberately published side by side. Subtracting one from the other would mix routing coverage, assets, redirects, and response classes into a fake conversion rate.
Crawl has not become measurable distribution
Search visibility is moving, citations are currently unmeasured, and attributed AI referrals remain scarce.
| Signal | Current evidence | What it proves |
|---|---|---|
| Google search visibility | 8,812 impressions, 2 clicks, average position 61.8Jul 12, 2026 to Aug 8, 2026 | Google surfaced Nice Pick. Impressions do not establish a visit or an answer-engine citation. |
| Persisted citations | Current performance unavailableLatest batch: 0/11, observed Jul 3, 2026 | The old batch establishes a baseline. Its age prevents a current citation claim. |
| Attributed AI referrals | 0 in 7 days; 3 in 30 days | A recognized answer product sent a browser visit. Attribution is heuristic and separate from crawler identity. |
Methodology
The edge series queries Cloudflare HTTP Requests Adaptive Groups in one-day slices, then groups results by user-agent identity and HTTP response status. The public snapshot removes raw unknown user-agent strings and retains aggregate actor classes and recognized crawler names.
The content-route series comes from Cloudflare Analytics Engine and uses SUM(_sample_interval) to reconstruct sampled counts. Search visibility comes from exact, consecutive 28-day Google Search Console windows with a two-day reporting lag. Citation status reads persisted probe rows and marks batches older than 14 days stale. AI referrals are aggregate edge events from the hardened post-July 25 attribution series.
Crawler roles follow the operators' published descriptions. OpenAI documents OAI-SearchBot for search and GPTBot for model training. Anthropic and Perplexity publish separate crawler controls for their products. These descriptions guide classification; they do not reveal what happened after a request.
Primary references: OpenAI crawlers, Anthropic crawler controls, Perplexity crawlers, and Cloudflare Analytics datasets.
Limits
- Cloudflare HTTP Requests Adaptive Groups are analytics estimates, not raw access logs.
- Crawler class totals cover up to the top 300 user-agent groups in each daily query; the overall edge total uses the zone-wide aggregate.
- User agents can be spoofed, and actor identity does not prove how fetched content was used.
- Zone-wide retrieval and content-route retrieval have different scopes and must not be compared as one series.
- Requests are not unique pages, people, referrals, recommendations, or citations.
- Google Search Console data is delayed by about two days.
- A stale citation probe means current citation performance is unavailable, not zero.
- AI referral attribution is heuristic and uses the post-2026-07-25 hardened series.
- This is one website. Its traffic mix should not be generalized to the web.
Use the data
The public JSON contains both comparable edge snapshots, source windows, actor totals, definitions, current search evidence, citation freshness, referral counts, and stated limits. The original July data file remains available for audit history.
Suggested citation
Nice Pick. “AI Crawler Observatory: Crawl, Citation, and Referral Data.” Updated Aug 10, 2026. https://nicepick.dev/research/ai-crawler-reality-report