Observed Signal · Aug 8, 2026 · Technical Release · Source: DEV Community · Impact: 3/5 · Sentiment: Neutral
Firewalls determine AI crawler access — 18-site audit
An author built an open-source tool, geo-crawl-audit, to probe how major websites treat AI crawler user-agents and how much readable content exists in raw HTML before JavaScript runs. The probe (against 18 sites on August 7, 2026) found that many sites' firewall and bot-management rules — not robots.txt alone — determine which AI crawlers can fetch pages, that several prominent crawlers (e.g., GPTBot, ClaudeBot, PerplexityBot) do not execute JavaScript, and that some well-known sites either deliberately or inadvertently present almost-empty raw HTML to most AI crawlers. Five sites blocked the probe's baseline requests entirely, highlighting the difficulty of measuring crawler access from arbitrary networks. The author published the tool and a public scanner to help operators check AI readability of their domains.
The audit and tool highlight concrete, operational gates (WAF/bot management, TTFB, server-rendering, robots.txt) that determine AI visibility of web content — relevant to publishers, SEO/GEO practitioners, and any organization aiming to control or enable LLM access, but not an industry-shifting platform policy change.
Track The Guardian Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- The author published an open-source tool, geo-crawl-audit, that probes sites using the user-agents of major AI crawlers and measures raw-HTML word counts.
- The tool was run against 18 major sites on August 7, 2026, producing a per-crawler status matrix and timing data.
- Many AI crawlers (e.g., GPTBot, ClaudeBot, PerplexityBot) do not execute JavaScript; Googlebot and Applebot are among the few AI-adjacent crawlers that render JS.
- Observed HTTP status patterns often aligned with public business relationships: e.g., The Guardian served GPTBot and related UAs a 200 while disallowing others in robots.txt; The New York Times returned 403s to most crawlers but permitted bingbot and Amazonbot.
- Five sites (Quora, OpenAI, Perplexity, Bloomberg, GitHub) challenged or filtered the probe's baseline browser request, causing the tool to flag them as BASELINE_ANOMALY.
Connected Companies & Entities
18 Entities mapped“The Guardian — which has a content deal with OpenAI — serves my simulated GPTBot, OAI-SearchBot, and ChatGPT-User a clean 200....”
“The Guardian — which has a content deal with OpenAI — serves my simulated GPTBot, OAI-SearchBot, and ChatGPT-User a clean 200....”
“The New York Times — in litigation with OpenAI — 403s nearly everyone: GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Common Crawl, meta-e...”
“Reddit's robots.txt blocks every AI bot in my list — fourteen tokens, no exceptions....”
“Figma runs the inverse configuration: serves a 200 to every bot UA in the list, but robots-disallows GPTBot, OAI-SearchBot, ChatGPT-User, Cl...”
“Airbnb returns 403 to nine of the twelve AI user-agents I probe — every OpenAI, Perplexity, Microsoft, Amazon, Meta, and Common Crawl token ...”
“Airbnb returns 403 to nine of the twelve AI user-agents I probe ... while, curiously, all three Anthropic user-agents get a 200....”
“LinkedIn served my probe a 23-word bot-check interstitial — whatever a verified crawler negotiates, the raw HTML a stranger gets is effectiv...”
“Quora, OpenAI, Perplexity, Bloomberg, and GitHub challenged even my baseline browser request from my network....”
“Stripe (1,957 raw-HTML words, a 200 to every UA in my list) ... — full server-rendered HTML, fast, no bot differentials against my probe....”
“OAI-SearchBot, Claude-SearchBot, PerplexityBot, Amazonbot, and meta-externalagent all received a 200 from the same IP, seconds apart....”
“OAI-SearchBot, Claude-SearchBot, PerplexityBot, Amazonbot, and meta-externalagent all received a 200 from the same IP, seconds apart....”
“Quora, OpenAI, Perplexity, Bloomberg, and GitHub challenged even my baseline browser request from my network....”
“Stripe (1,957 raw-HTML words, a 200 to every UA in my list), Anthropic (fastest warm response in the test at 0.102s), Vercel, MDN, Shopify —...”
“Stripe (1,957 raw-HTML words, a 200 to every UA in my list), Anthropic (fastest warm response in the test at 0.102s), Vercel, MDN, Shopify —...”
“Quora, OpenAI, Perplexity, Bloomberg, and GitHub challenged even my baseline browser request from my network....”
“Quora, OpenAI, Perplexity, Bloomberg, and GitHub challenged even my baseline browser request from my network....”
Ontology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
Study: Nearly Half of Top German Sites Block AI Crawlers
A Hotwire study titled "The AI Coverage Gap" found that 48% of the 100 highest-reach German online media are not or only partially accessible to AI crawlers. The research evaluated access by ten AI crawlers (from providers such as OpenAI, Anthropic and Perplexity) to the domains of those top-100 publishers in June 2026. Major publishers (e.g., Bild, Der Spiegel, stern, Focus Online) largely block crawlers, while national newspapers like Frankfurter Allgemeine, Süddeutsche Zeitung and Die Zeit generally allow access. The study highlights strategic dilemmas for publishers and communications teams about visibility in LLM-generated answers, commercial implications for traffic and subscriptions, and technical mitigation options (robots.txt, CDN/server rules, paywalls). Hotwire recommends companies consider AI accessibility when selecting target media and measure AI visibility as part of communications KPIs.
Let AI Crawlers In: AEO Not SEO in 2026
The article argues that 'AEO' (Answer Engine Optimization) is a distinct discipline from classic SEO and that publishers who want to be cited by AI assistants should allow AI crawlers to index their pages. It explains that AI crawlers are vendor- and task-specific (training, search-indexing, live fetch) and lists explicit user-agent names for OpenAI, Anthropic, Perplexity and Google. The piece reminds readers that robots.txt (RFC 9309, 2022) is a voluntary request, not an enforcement mechanism, and that hard blocks require server/CDN/WAF rules. Recommended practical steps include per-bot robots.txt entries, server-rendered HTML, structured data (schema.org), adopting the emerging llms.txt convention, and ensuring CDN bot‑management aligns with robots.txt.
One-Call API to Check AI Search Visibility
A developer (Daniel Igel) published a small audit engine and web tool that scores whether a website can be crawled, parsed and cited by AI search assistants (e.g., ChatGPT Search, Perplexity, Google AI Overviews). The audit evaluates robots.txt access for AI crawlers (GPTBot, OAI-SearchBot, PerplexityBot, ClaudeBot), presence of llms.txt, JSON-LD structured data, and passage citability, producing a 0–100 score with per-category pass/warn/fail results and concrete recommendations. The API is keyless and callable via a POST to https://citeready-api.sprytools.com/v1/audit; a web UI at https://citeready.sprytools.com offers a free tier (three audits/day) without signup.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
