GEO / AEO Audit — AI Visibility at a Glance

What ChatGPT, Claude, Perplexity & Gemini can see of shop.usaclean.com — and the one-week fix
July 2026 · audit review
Full brief →
Executive summary (PDF) →

The store is invisible to AI assistants — by a switch, not by fate.

Cloudflare returns 403 to AI crawlers at the network edge, and robots.txt says two contradictory things about them. ChatGPT, Claude, and Perplexity cannot retrieve or cite a single page — while the only AI surface still working (Google AI Overviews) works by accident of using regular Googlebot.

Everything else this audit found is ordinary, well-understood content work, most of it already scoped by the June SEO playbook. The ask for this review: approve Stage 0 — unblock AI retrieval crawlers this week. It's hours of configuration, and nothing else on this page matters until it lands.

Stage 0
Turn AI visibility on
Hours · week 1
Recommended now
Stage 1
Make content quotable
Medium · weeks 2–6
Stage 2
Measure & compound
Low, ongoing · months 2–3

What the audit found

Four layers, graded on the live site (2026-07-08). Full evidence and method in the brief.

Layer Grade Headline evidence Fixed by
1 · Crawler access — can AI engines fetch pages at all?BlockedCloudflare 403s AI bots at the edge; robots.txt holds two contradictory policiesStage 0
2 · Structured data — can machines parse meaning?B−Product pages strong (Product + Offer + Breadcrumb); guides and About have zero schemaStage 1
3 · Extractable content — is there a quotable answer?CContent hub exists but isn't answer-first; "600 technicians / 40 years" not in crawlable HTMLStage 1
4 · Measurement — would anyone know if it worked?AbsentNo AI-referral tracking, no citation monitoring, no re-check cadenceStage 2
Good working today B− / C partial — real gaps Blocked not functioning at all

The receipts

Two captures from the live site, 2026-07-08.

Chrome (browser)200 OK
GPTBot (OpenAI)403 Forbidden
ClaudeBot (Anthropic)403 Forbidden
PerplexityBot403 Forbidden
Same URL, four user agents — Cloudflare blocks AI bots at the edge, before robots.txt applies.
# Cloudflare Managed content
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /

# …further down, the origin policy
User-agent: OAI-SearchBot
User-agent: PerplexityBot
Crawl-delay: 10
Disallow: /cart.php  # otherwise allowed
robots.txt, arguing with itself — one file, two policies. Allowed on paper, 403'd in practice.
0

Turn AI visibility on

Unblock retrieval at the edge — nothing else in this audit matters until this lands.
Hours · week 1 Recommended now Cloudflare + robots.txt

What's in

  • Unblock AI retrieval crawlers in Cloudflare — the 403 is an edge setting, not code; review the AI-bot toggle and managed-robots.txt injection
  • Reconcile robots.txt into one deliberate policy — retrieval bots allowed; training is the client's call (ai-train=no can stay)
  • Decide Google-Extended explicitly — labeled a training control, but it also gates Gemini grounding
  • Ship llms.txt — a same-afternoon checkbox, budgeted as exactly that

Deliberately out (→ Stage 1)

  • Any content or schema rewrites — pointless while the crawlers are still locked out
Read Stage 0 in the brief
1

Make content quotable

Once crawlers can get in, give them answer-shaped facts worth citing.
Medium · weeks 2–6 Content + Stencil theme

What's in

  • Entity-clarity copy pass — lead category/guide/About pages with the flattest extractable fact, then elaborate
  • Credibility stats into crawlable HTML — "600 technicians / 40 years" currently isn't machine-readable anywhere
  • Article + author + dates schema on guides — the hub pages AI engines would quote today carry zero markup
  • Q&A blocks + FAQPage on top 20 categories — 3–5 real buyer questions each
  • SEO carry-over — PDP meta descriptions (still raw part dumps) and the homepage title (still just "USA-CLEAN")

Deliberately out (→ Stage 2)

  • Video transcripts + VideoObject — rides on the playbook's YouTube-repurposing initiative when it happens
Full recommendation table
2

Measure & compound

Prove it works, and make sure the fix stays fixed.
Low, recurring Months 2–3 · ongoing

What's in

  • GA4 "AI Assistant" channel group — native for ChatGPT/Gemini/Claude since May 2026; one manual channel for Perplexity. Configure first — every week without it loses baseline
  • Monthly citation panel — does ChatGPT recommend USA-CLEAN when a technician asks about a Tennant T7 squeegee blade? Fixed prompts, logged monthly
  • Quarterly crawler-access re-check — the 403 test above takes ten minutes; CDN defaults shift silently
  • Video transcripts + VideoObject — when the YouTube repurposing lands

Cost

  • Recurring attention, not build effort — a runbook, not a project
No prototype — this lives in GA4 and a monthly runbook.

The dividing line: retrieval vs training

One distinction resolves the whole robots.txt policy question — and it's the client conversation at the heart of Stage 0.

The dividing line

Blocking AI training is a policy choice; blocking AI retrieval is invisibility.

A facilities manager asks ChatGPT which squeegee blade fits a Tennant T500e. Today, USA-CLEAN structurally cannot be the answer — the retrieval crawler is 403'd, so a competitor gets the citation and the order. Unblocking retrieval doesn't require giving anything away to model training: the ai-train=no signal and the training-bot blocks can stay.

✓ Allow — retrieval & citation agents

  • OAI-SearchBot / ChatGPT-User — ChatGPT search index + live user visits
  • PerplexityBot — Perplexity's answer index
  • ClaudeBot / Claude-Web — Claude retrieval + user-triggered fetches
  • Google-Extended — if Gemini grounding is wanted (explicit call)

○ Client's call — training-only crawlers

  • CCBot, Bytespider, meta-externalagent — dataset collection; blocking is a defensible IP stance
  • GPTBot as trainer — OpenAI honors the retrieval/training split via separate bots
  • Content-Signal: ai-train=no — can stay exactly as is

Who can see the store

Per engine: visibility today vs after Stage 0. The amber column is what the one-week fix buys.

Engine Today Why After Stage 0
ChatGPT (search + browsing)BlockedOAI-SearchBot / ChatGPT-User 403'd at the WAF; GPTBot also disallowedVisible
ClaudeBlockedClaudeBot disallowed in robots.txt and 403'd at the WAFVisible
PerplexityBlockedAllowed by robots.txt on paper — 403'd at the WAF in practiceVisible
Gemini (grounding)ExcludedGoogle-Extended disallowed by the Cloudflare-managed blockClient's call
Google AI OverviewsVisibleUses the standard Googlebot index — unaffected by the AI-bot blockVisible
Visible can be retrieved & cited Client's call depends on the Google-Extended decision Blocked cannot appear in answers

The recommendation for this review

Suggested path

Approve Stage 0 this week: unblock AI retrieval crawlers in Cloudflare, reconcile robots.txt into one deliberate policy, and ship llms.txt. It's hours of configuration with no content work, and it moves four of the five engines above from Blocked to Visible.

Honesty about the stakes: AI referrals are under 2% of sessions industry-wide and the conversion evidence is mixed (B2B leans positive) — this is cheap insurance on the fastest-growing referral channel, not a promised lift. A go-ahead here is logged with the next PD-NNN id in the decision log; Stages 1–2 then proceed as content and measurement work alongside the existing SEO playbook.

GEO / AEO audit — full brief docs/geo-aeo-audit.html Executive summary (PDF) docs/geo-aeo-executive-summary.pdf Markdown source docs/geo-aeo-audit.md Live receipt: shop.usaclean.com/robots.txt the double policy, verbatim All review decks reviews/