Observed Signal · Aug 11, 2026 · Guidance · Source: DEV Community · Impact: 3/5 · Sentiment: Neutral
Permissive robots.txt Is Not a License
The author audited ten sites his scraper had been reading and found that robots.txt permissiveness does not equate to permission to reuse content. Two sites passed his commercial-reuse licence check; eight did not. He distinguishes two separate questions: robots.txt controls whether a bot may fetch pages, while a licence (or explicit written grant) controls whether fetched content may be republished or reused commercially. Examples include Simon Willison (welcoming robots.txt but no licence), Troy Hunt (robots.txt permissive and footer licensed under CC BY 4.0), Julia Evans (robots.txt disallow and a visible "NO LLM PLZ" message but no formal licence), and a Pragmatic Engineer header using a machine-readable Content-Signal. The author argues the admission-time licence gate (check licence before ingesting) is the right control to avoid legal and trust risks.
Practical guidance on ingest/licensing controls affects legal/compliance risk for organizations that crawl web content and reuse it commercially (relevant to datasets, publishers, and any business using scraped content).
Track DEV Community Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- The author audited ten sites his scraper had been crawling since early July; two passed his commercial-reuse licence requirement and eight failed.
- Robots.txt answers "may your bot fetch this page?" while a licence answers "may you republish what your bot fetched?" — they are distinct permissions.
- Simon Willison's robots.txt explicitly allowed an AI user-agent but his site had no written licence in the footer.
- Troy Hunt's site included a Creative Commons Attribution 4.0 International License in the footer, which explicitly permits reuse with attribution.
- The Pragmatic Engineer uses a machine-readable header example: "Content-Signal: search=yes, ai-input=yes, ai-train=no" (retrieval allowed, training disallowed).
Connected Companies & Entities
5 Entities mapped“DEV Community...”
“Powered by Algolia...”
“Neon is the official database partner of DEV...”
“Built on Forem — the open source software that powers DEV...”
Ontology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
llms.txt vs robots.txt vs ai.txt Explained
This developer guide compares three site-level files—robots.txt, llms.txt, and ai.txt—used by crawlers and AI assistants to discover, index, and (in some cases) respect publisher intent. robots.txt (since 1994) remains the standard for crawl access and path-based Allow/Disallow rules. llms.txt is an emerging Markdown-based convention (adopted by Anthropic, Perplexity and some GPTBot variants) that documents site context for LLMs and AI-search engines rather than controlling access. ai.txt is a newer permission-focused proposal (AI-txt.com initiative) combining key-value directives and JSON blocks to grant or deny assistant usage, but it currently lacks major enforcement. The article includes Next.js App Router examples for dynamically generating robots.txt and llms.txt, a sample static ai.txt, and a recommended crawl decision flow. Practical advice: always publish robots.txt, add llms.txt for accurate AI citations, and include a simple ai.txt to signal intent.
Firewalls determine AI crawler access — 18-site audit
An author built an open-source tool, geo-crawl-audit, to probe how major websites treat AI crawler user-agents and how much readable content exists in raw HTML before JavaScript runs. The probe (against 18 sites on August 7, 2026) found that many sites' firewall and bot-management rules — not robots.txt alone — determine which AI crawlers can fetch pages, that several prominent crawlers (e.g., GPTBot, ClaudeBot, PerplexityBot) do not execute JavaScript, and that some well-known sites either deliberately or inadvertently present almost-empty raw HTML to most AI crawlers. Five sites blocked the probe's baseline requests entirely, highlighting the difficulty of measuring crawler access from arbitrary networks. The author published the tool and a public scanner to help operators check AI readability of their domains.
robots.txt vs llms.txt vs sitemap.xml: What Each Does
The article explains the distinct roles of three small web files: robots.txt (an access policy at /robots.txt that tells well-behaved crawlers which paths to fetch), sitemap.xml (a sitemap that lists URLs and metadata to help search engines discover pages), and llms.txt (a 2024-spec markdown brief at /llms.txt intended to curate high-signal pages for AI assistants). It clarifies common misconceptions — e.g., robots.txt does not deindex content or protect private data, sitemaps do not force indexing, and llms.txt does not bind models or replace linked pages — and describes how crawlers and assistants may read these files in different orders. The piece advises shipping robots.txt first, then sitemap.xml for discoverability, and llms.txt optionally for AI citation, while keeping all files in sync.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
