The AI Crawler Observatory
Retrieval crawl fell 81.8% from the prior snapshot. Attributed AI visits stayed at 0.
Nice Pick now tracks crawl, search visibility, citations, and human referrals as separate signals. The latest edge snapshot counted 419 retrieval-class requests. The current citation result is unavailable because the last probe is stale.
Current verdict
The August retrieval spike did not hold. Measurable distribution remains absent.
The latest seven-day edge snapshot counted 419 retrieval-class requests, down from 2,300 in the prior published snapshot but still above the July baseline of 331. The training-to-retrieval ratio widened from 134:1 to 437:1. Cloudflare Analytics Engine recorded 0 attributed AI arrivals in the latest seven days. The most recent citation batch is 58 days old, so it cannot answer whether citation performance changed.
Latest scorecard
Each number keeps its own source and reporting window. None stands in for another.
- Zone-wide retrieval fetches
- 419
- Change from prior snapshot
- -81.8%
- Attributed AI arrivals
- 0
- Current citation rate
- Unknown
Aug 23, 2026 to Aug 30, 2026
July baseline: +26.6%
Latest seven days
Probe is 58 days old
Comparable edge snapshot history
Every snapshot uses Cloudflare zone-wide adaptive request groups, one-day query slices, the same user-agent classifier, and a seven-day window.
| Window | Retrieval | Training | Search crawlers | 5xx | Training / retrieval |
|---|---|---|---|---|---|
| Jul 8, 2026 to Jul 15, 2026 | 331 | 254,595 | 1,926 | 5,223 | 769:1 |
| Aug 3, 2026 to Aug 10, 2026 | 2,300 | 307,145 | 3,576 | 15 | 134:1 |
| Aug 23, 2026 to Aug 30, 2026 | 419 | 183,280 | 2,380 | 1 | 437:1 |
Retrieval actors in the latest snapshot
OAI-SearchBot supplied most retrieval-class requests. A request can make a page eligible for an answer. It cannot prove selection or citation.
The two retrieval counts answer different questions
Zone-wide edge audit
419
Requests across every route and response status seen by Cloudflare. This is the broad infrastructure view.
Content-route telemetry
91
Matched content middleware routes and explicit server logging. This narrower series rose 111.6%against the preceding seven days.
The scopes are deliberately published side by side. Subtracting one from the other would mix routing coverage, assets, redirects, and response classes into a fake conversion rate.
Crawl has not become measurable distribution
Search impressions fell 24.9%between consecutive 28-day windows. Citations remain unmeasured, and attributed AI referrals remain absent.
| Signal | Current evidence | What it proves |
|---|---|---|
| Google search visibility | 7,433 impressions, 1 clicks, average position 70.2Aug 1, 2026 to Aug 28, 2026 | Google surfaced Nice Pick. Impressions do not establish a visit or an answer-engine citation. |
| Persisted citations | Current performance unavailableLatest batch: 0/11, observed Jul 3, 2026 | The old batch establishes a baseline. Its age prevents a current citation claim. |
| Attributed AI referrals | 0 in 7 days; 0 in 30 days | A recognized answer product sent a browser visit. Attribution is heuristic and separate from crawler identity. |
Methodology
The edge series queries Cloudflare HTTP Requests Adaptive Groups in one-day slices, then groups results by user-agent identity and HTTP response status. The public snapshot removes raw unknown user-agent strings and retains aggregate actor classes and recognized crawler names.
The content-route series comes from Cloudflare Analytics Engine and uses SUM(_sample_interval) to reconstruct sampled counts. Search visibility comes from exact, consecutive 28-day Google Search Console windows with a two-day reporting lag. Citation status reads persisted probe rows and marks batches older than 14 days stale. AI referrals are aggregate edge events from the hardened post-July 25 attribution series.
Crawler roles follow the operators' published descriptions. OpenAI documents OAI-SearchBot for search and GPTBot for model training. Anthropic and Perplexity publish separate crawler controls for their products. These descriptions guide classification; they do not reveal what happened after a request.
Primary references: OpenAI crawlers, Anthropic crawler controls, Perplexity crawlers, and Cloudflare Analytics datasets.
Limits
- Cloudflare HTTP Requests Adaptive Groups are analytics estimates, not raw access logs.
- Crawler class totals cover up to the top 300 user-agent groups in each daily query; the overall edge total uses the zone-wide aggregate.
- User agents can be spoofed, and actor identity does not prove how fetched content was used.
- Zone-wide retrieval and content-route retrieval have different scopes and must not be compared as one series.
- Requests are not unique pages, people, referrals, recommendations, or citations.
- Google Search Console data is delayed by about two days.
- A stale citation probe means current citation performance is unavailable, not zero.
- AI referral attribution is heuristic and uses the post-2026-07-25 hardened series.
- This is one website. Its traffic mix should not be generalized to the web.
Use the data
The public JSON contains all comparable edge snapshots, source windows, actor totals, definitions, current search evidence, citation freshness, referral counts, and stated limits. The original July data file remains available for audit history.
Suggested citation
Nice Pick. “AI Crawler Observatory: Crawl, Citation, and Referral Data.” Updated Aug 30, 2026. https://nicepick.dev/research/ai-crawler-reality-report