Observed Signal · May 15, 2026 · Technical Audit · Source: DEV Community · Impact: 3/5 · Sentiment: Neutral

Audit Finds Most Companies Lack llms.txt

Executive Signal Summary

An audit of 70 prominent sites (AI labs, docs platforms, dev infra, SaaS, WordPress ecosystem, and community sites) found only 31 returned a real llms.txt file; after deduping there were 29 unique samples. Many high-profile model and platform providers either return 404, serve an HTML single-page-app fallback at /llms.txt, or block access. Of the 31 existing files, 12 (41%) do not follow the simple llms.txt spec (H1 + blockquote). File sizes vary widely (648 bytes to ~280 KB), reflecting divergent interpretations of the file’s purpose. The auditor ran vanilla HTTP GET probes on 2026-05-16 and published the probe script and samples; the article was published 2026-05-15.

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

Technical audit of llms.txt adoption highlights gaps and a high‑leverage operational fix (HTML fallback) relevant to how LLM crawlers discover and cite web content; useful but not industry‑shifting by itself.

SIGNAL RADAR

Track Anthropic Signals & Market Shifts in Real-Time

Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.

Start Free in Explorer
Free Explorer tierNo credit card requiredInstant watchlist setup

Key Takeaways & Evidence Grounding

  • The author probed 70 candidate sites on 2026-05-16 using vanilla HTTP GET and a desktop browser User-Agent.
  • 31 of 70 (44%) returned a real /llms.txt; the other outcomes were 22 404s (31%), 9 HTML fallbacks (13%), and 8 403/530/timeouts (11%).
  • After deduping mirror domains there were 29 unique llms.txt samples; 17 (59%) followed the required H1 + blockquote pattern and 12 (41%) did not.
  • File sizes for real llms.txt samples ranged from 648 bytes to 279,743 bytes (median 9.4 KB; mean 34.7 KB).
  • Several high-profile model and docs providers (examples cited: OpenAI, Hugging Face, Google AI, Mintlify, GitBook, Readme, fast.ai) did not have a compliant /llms.txt at the probed root path.
Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: DEV Community•Published: May 15, 2026
Original Coverage Title: “I Audited 70 Companies' llms.txt Files. Most Don't Have One.”

Related Market Signals & Shifts

Recent verified developments and strategic activity across this market segment.

SEO, GEO & SEM PlatformMay 6, 2026

llms.txt: Optimize Sites for ChatGPT and LLMs (2026)

A DEV Community post (published 2026-05-06) by Dilm Informatique explains the llms.txt standard — a text file placed at the root of a domain that summarizes a website’s content for large language models. The article positions llms.txt as the equivalent of robots.txt for AI agents and says it helps services like ChatGPT, Perplexity and Google AI Overviews cite a site. It also lists complementary technical signals to improve AI discoverability: SSR prerendering, JSON-LD structured data (LocalBusiness, Service, FAQ), thematic sitemaps and optimized Core Web Vitals. The post cites a live example (Fix72’s implementation) and links to source code on GitHub.

Read assessment
SEO & AI OptimizationApr 11, 2026

Implementing llms.txt for AI Optimization

This technical guide explains how to implement the llms.txt specification to make websites AI-friendly. llms.txt is a single Markdown file placed at the web root that provides LLM-targeted summary information about an organization, similar in intent to robots.txt but aimed at large language models. The article documents Immagina Group’s real-world implementation as part of its AI Optimization (AIO) framework, provides file structure and code examples, and describes a broader knowledge-file ecosystem (llms-full.txt, ai-knowledge.json, entities.txt, citations.txt, brand.txt). Practical advice covers serving files as plain text, robots.txt rules to allow AI crawlers, registering in llms directories, and combining llms.txt with Schema.org markup. The author reports measurable impact: major LLMs cited Immagina Group within 30 days and a client (Omega Professional) saw +25% AI-sourced leads and +15% revenue within five months.

Read assessment
SEO, GEO & SEM PlatformMay 23, 2026

llms.txt vs robots.txt vs ai.txt Explained

This developer guide compares three site-level files—robots.txt, llms.txt, and ai.txt—used by crawlers and AI assistants to discover, index, and (in some cases) respect publisher intent. robots.txt (since 1994) remains the standard for crawl access and path-based Allow/Disallow rules. llms.txt is an emerging Markdown-based convention (adopted by Anthropic, Perplexity and some GPTBot variants) that documents site context for LLMs and AI-search engines rather than controlling access. ai.txt is a newer permission-focused proposal (AI-txt.com initiative) combining key-value directives and JSON blocks to grant or deny assistant usage, but it currently lacks major enforcement. The article includes Next.js App Router examples for dynamically generating robots.txt and llms.txt, a sample static ai.txt, and a recommended crawl decision flow. Practical advice: always publish robots.txt, add llms.txt for accurate AI citations, and include a simple ai.txt to signal intent.

Read assessment

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.