The store is invisible to AI assistants — by a switch, not by fate.
Cloudflare returns 403 to AI crawlers at the network edge, and robots.txt says two contradictory things about them. ChatGPT, Claude, and Perplexity cannot retrieve or cite a single page — while the only AI surface still working (Google AI Overviews) works by accident of using regular Googlebot.
Everything else this audit found is ordinary, well-understood content work, most of it already scoped by the June SEO playbook. The ask for this review: approve Stage 0 — unblock AI retrieval crawlers this week. It's hours of configuration, and nothing else on this page matters until it lands.
Prove it works, and make sure the fix stays fixed.
Low, recurringMonths 2–3 · ongoing
GA4 AI-referral channel + a monthly citation panel logged against ~10 buyer-intent prompts.
What's in
GA4 "AI Assistant" channel group — native for ChatGPT/Gemini/Claude since May 2026; one manual channel for Perplexity. Configure first — every week without it loses baseline
Monthly citation panel — does ChatGPT recommend USA-CLEAN when a technician asks about a Tennant T7 squeegee blade? Fixed prompts, logged monthly
Quarterly crawler-access re-check — the 403 test above takes ten minutes; CDN defaults shift silently
Video transcripts + VideoObject — when the YouTube repurposing lands
Cost
Recurring attention, not build effort — a runbook, not a project
No prototype — this lives in GA4 and a monthly runbook.
The dividing line: retrieval vs training
One distinction resolves the whole robots.txt policy question — and it's the client conversation at the heart of Stage 0.
The dividing line
Blocking AI training is a policy choice; blocking AI retrieval is invisibility.
A facilities manager asks ChatGPT which squeegee blade fits a Tennant T500e. Today, USA-CLEAN structurally cannot be the answer — the retrieval crawler is 403'd, so a competitor gets the citation and the order. Unblocking retrieval doesn't require giving anything away to model training: the ai-train=no signal and the training-bot blocks can stay.
✓ Allow — retrieval & citation agents
OAI-SearchBot / ChatGPT-User — ChatGPT search index + live user visits
PerplexityBot — Perplexity's answer index
ClaudeBot / Claude-Web — Claude retrieval + user-triggered fetches
Google-Extended — if Gemini grounding is wanted (explicit call)
○ Client's call — training-only crawlers
CCBot, Bytespider, meta-externalagent — dataset collection; blocking is a defensible IP stance
GPTBot as trainer — OpenAI honors the retrieval/training split via separate bots
Content-Signal: ai-train=no — can stay exactly as is
Who can see the store
Per engine: visibility today vs after Stage 0. The amber column is what the one-week fix buys.
Engine
Today
Why
After Stage 0
ChatGPT (search + browsing)
Blocked
OAI-SearchBot / ChatGPT-User 403'd at the WAF; GPTBot also disallowed
Visible
Claude
Blocked
ClaudeBot disallowed in robots.txt and 403'd at the WAF
Visible
Perplexity
Blocked
Allowed by robots.txt on paper — 403'd at the WAF in practice
Visible
Gemini (grounding)
Excluded
Google-Extended disallowed by the Cloudflare-managed block
Client's call
Google AI Overviews
Visible
Uses the standard Googlebot index — unaffected by the AI-bot block
Visible
Visible can be retrieved & citedClient's call depends on the Google-Extended decisionBlocked cannot appear in answers
The recommendation for this review
Suggested path
Approve Stage 0 this week: unblock AI retrieval crawlers in Cloudflare, reconcile robots.txt into one deliberate policy, and ship llms.txt. It's hours of configuration with no content work, and it moves four of the five engines above from Blocked to Visible.
Honesty about the stakes: AI referrals are under 2% of sessions industry-wide and the conversion evidence is mixed (B2B leans positive) — this is cheap insurance on the fastest-growing referral channel, not a promised lift. A go-ahead here is logged with the next PD-NNN id in the decision log; Stages 1–2 then proceed as content and measurement work alongside the existing SEO playbook.