Skip to main content
AI Visibility Updated 2 August 2026 12 min read Originally published January 2026

AI Discovery Files: Complete Guide to All 10 Essential Files

10 structured files tell AI systems like ChatGPT and Claude exactly what your business does. Most websites have zero. Here's what each file does, how they work together, and why you need all 10.

MM
Mark McNeece Founder & Managing Director, 365i
Ten AI discovery file icons arranged in an interconnected ecosystem diagram showing how llms.txt, ai.txt, brand.txt, and seven other files work together

Ever wondered why ChatGPT recommends some businesses but ignores others? AI systems aren't figuring this out magically from your website. They're looking for specific signals: structured files that tell them exactly what you do, how to describe you, and what they're allowed to say about you.

And here's where it gets interesting. There's not just one file. There are ten different files that work together to create your complete AI Site Identity. Most businesses have zero. Some have one or two. Almost nobody has all ten.

I ran our free AI Site Identity checker on 200 UK business websites last month. Twelve had proper AI discovery files in place. That's 6%. And these weren't tiny startups. Established companies with marketing budgets, SEO consultants, the works. They just didn't know this existed.

The AI Discovery File Ecosystem

When someone asks ChatGPT "recommend a good WordPress host in the UK," it has to make a decision. Based on what? Your homepage? Your About page? That blog post you wrote three years ago and forgot about?

Websites are messy. They're full of contradictions, marketing speak, outdated information, and content written for humans who understand context. AI systems don't do context the way we do. They need clear, structured, factual signals.

That's what AI discovery files are: your business's identity card for AI systems. Instead of letting ChatGPT or Claude guess based on fragments, you're giving them a proper dossier.

Most people think creating one file (usually llms.txt, because that's the one everyone talks about) is job done. It's not. That's like a CV that only lists your name. Technically it exists, but it's not doing you any favours. We covered the llms.txt foundation in our complete guide to creating AI identity files, but it's just one piece of the puzzle.

The ten files work as an ecosystem:

  • llms.txt - Your core business identity for Large Language Models
  • llm.txt - Compatibility variant that redirects to llms.txt
  • llms.html - Human-readable HTML version with Schema.org structured data
  • ai.txt - Permissions and usage policies for AI systems
  • ai.json - Machine-parseable AI interaction guidance in JSON format
  • identity.json - Structured canonical data aligned with Schema.org
  • brand.txt - How you want to be named and described
  • faq-ai.txt - Verified questions and answers AI can reference
  • developer-ai.txt - Technical context about your platform and stack
  • robots-ai.txt - Specific crawler directives for AI bots

Each one serves a purpose the others don't cover. Think of them like different departments in a company: marketing, legal, technical, customer service. They all need to exist and they all need to be saying the same things. Together, they form what we've started calling the web's identity layer, following the same pattern as robots.txt and sitemaps before them.

What Each File Actually Does

llms.txt: The Foundation

This is your core business identity file. When an AI system wants to know "what does this company actually do," llms.txt is where it looks first. It contains your company name (exactly as you want it referenced), what you do (factual, not marketing waffle), where you operate, key services, and target audience.

The biggest mistake? Making it too marketing-focused. "We're passionate about excellence" means nothing to AI. "WordPress hosting provider, Kettering, UK, established 2002" tells it everything.

llm.txt: The Compatibility Variant

Some AI systems request llm.txt (singular) rather than llms.txt (plural). Both filenames exist in the wild, so the specification includes llm.txt as a compatibility variant. The recommended approach is a 301 redirect from llm.txt to llms.txt, so you maintain a single source of truth while catching both requests.

ai.txt: The Permissions File

This tells AI systems what they're allowed to do with your content. Can they quote you? Summarise you? What about commercial use? Without ai.txt, AI systems make their own decisions about how to use your content. Sometimes they get it right. Sometimes they quote you out of context or make claims you never made.

ai.json: Machine-Readable Interaction Guidance

Where ai.txt is human-readable, ai.json delivers the same permissions and interaction guidance in structured JSON with schema validation. AI systems that process JSON natively can parse your rules without interpreting free-form text. It's the difference between handing someone a paragraph and handing them a form with checkboxes: less room for misinterpretation.

brand.txt: Naming Consistency

Controls how AI writes your company name, product names, and key terminology. Sounds trivial until you've seen ChatGPT call you three different names in the same response. Inconsistent branding in AI responses looks unprofessional and confuses potential customers.

faq-ai.txt: Verified Q&A

Pre-verified questions and answers AI can reference when asked about your business. When someone asks ChatGPT "how long does WordPress hosting setup take," and you've provided a verified answer, AI is far more likely to reference yours than guess from random web content.

developer-ai.txt: Technical Context

Technical information about your platform, stack, and infrastructure. When developers research WordPress hosting options and ask AI technical questions, you want accurate answers about your infrastructure, not guesses. "Does 365i support PHP 8.3?" should get a confident, correct answer.

robots-ai.txt: Crawler Directives

Specific instructions for AI web crawlers about what they can access. This is separate from standard robots.txt because AI crawlers often need different rules than traditional search engine bots.

identity.json: Structured Data

Your business information in machine-readable JSON. The most structured of all ten files and the easiest for AI systems to parse. API endpoints, social profiles, contact details, service catalogue: all in one clean, parseable format.

llms.html: Human-Readable Version

A formatted HTML page containing the same information as your text-based files. It serves a dual purpose: AI systems that prefer HTML can parse it, and humans (including potential customers) can read it to understand what AI knows about your business. We publish ours publicly.

How They Work Together

These files don't exist in isolation. They reference each other, support each other, and they need to tell consistent stories. This is where DIY implementations fall apart.

I reviewed someone's files last month. Their llms.txt said "Company Name Ltd" and their brand.txt said "CompanyName Ltd" (no space). Tiny detail. But when ChatGPT sees conflicting signals, it doesn't know which to trust. It might use one version, or the other, or make up a third that's neither.

The ten files operate in three layers:

Core Identity: llms.txt, llm.txt, identity.json, and llms.html establish who you are. These four have to agree with each other down to the character. Same company name. Same address. Same service descriptions. (llm.txt simply redirects to llms.txt, so that's handled automatically.)

Control: ai.txt, ai.json, brand.txt, and robots-ai.txt manage how AI interacts with you. ai.json delivers the same rules as ai.txt in structured JSON. If ai.txt says "yes, AI can quote our content" but robots-ai.txt blocks all AI crawlers, that's a contradiction AI systems will notice.

Enhancement: faq-ai.txt and developer-ai.txt enrich the core identity with specific, useful information. They add depth to what the foundation files establish.

Consistency is the whole game, and it's the reason a generated set usually beats a hand-written one. Not because writing the files is hard, but because ten files written across ten different afternoons drift apart. Running a managed hosting platform since 2002 has taught us that the thing which breaks is almost never the clever part. It's the boring part nobody went back and re-checked.

Do AI Crawlers Actually Read AI Discovery Files?

This question should come before the ten-file checklist, and most articles about llms.txt skip it. So here is the case against, put as strongly as the people making it put it.

"The cons outweigh the pros right now. If you want to show up in AI search, there are more reliable ways to improve your visibility than this file."

Louise Linehan, Ahrefs, We Analyzed 137K Sites: 97% of llms.txt Files Never Get Read

Ahrefs looked at 137,000 sites and found that 97% of llms.txt files never get requested at all. That's a real finding on a serious sample, and we're not going to argue with it, because for the median website it is almost certainly correct.

But 97% getting zero requests describes a distribution, and a distribution has two ends. Most llms.txt files sit on abandoned sites, or contain broken syntax, or were generated once and linked from nowhere. Naturally nothing fetches them. The question worth answering isn't whether the average file gets read. It's whether an actively maintained, valid, complete set gets read.

We can answer that one, because we run it. Over the 30 days to 2 August 2026, the ten files on mcneece.com (one of our own properties, running the free plugin) were fetched 58 times by 10 distinct AI crawlers, and not one request was blocked. Seven operators appear in that log: OpenAI, Anthropic, Perplexity, Apple, Microsoft, Meta and ByteDance.

Three things in that data matter more than the headline number:

  • llm.txt was fetched zero times. The singular-filename variant, which exists in the specification only because both spellings turned up in the wild, has never once been requested on that site. We publish it and it does nothing. That's the honest half of the result.
  • brand.txt punches above its weight. In one 30-day window the naming file was read more often than llms.txt itself. The file everybody writes about is not reliably the file that gets read most, which is an argument for the full set rather than the famous one.
  • Crawlers arrive late. In an 11-day window shortly after install, only Meta and Bing appeared and the headline LLM crawlers were entirely absent. By around day 40 they had all turned up. A site that looks ignored in week one is not necessarily being ignored.

Now the limits, because a number without its caveats is just marketing. A fetch is not a citation: knowing GPTBot read your brand.txt tells you nothing about whether ChatGPT went on to recommend you. User agents are a claim rather than a credential, and Ahrefs found roughly a fifth of llms.txt requests come from audit tools rather than AI systems, so some share of any crawler dashboard is scanners. And this is one site. An n of 1 doesn't overturn a 137,000-site aggregate and isn't meant to. It answers a narrower question: does a valid, maintained, actively-served set get crawled? On the evidence in front of us, yes.

If you'd rather have that answer for your own site than ours, the free WordPress plugin logs every AI crawler that touches your files, and our AI Bot Checker tells you whether those crawlers can reach you in the first place, which is the prior question and the one more sites fail.

DIY vs Professional Setup

You've got two paths.

DIY: Start with our guide to creating llms.txt, then work through the other nine using this article as your roadmap. Budget 10-15 hours and be prepared for technical challenges. Use our free checker to validate as you go.

Professional setup: All ten files created, tested, and validated. Delivered in 24-48 hours. The biggest advantage isn't speed, it's consistency. We've seen too many DIY implementations where files contradict each other because different sections were written at different times.

Either way, you'll know exactly where you stand. The free checker takes 30 seconds and scores your current AI visibility out of 100. Most sites score under 15.

Free WordPress Plugin

Generate AI Discovery Files from your dashboard

Using WordPress? Install the plugin and create all 10 files in minutes. No coding, no configuration files to edit manually.

Get the Plugin →

Testing Your AI Discovery Files

Creating the files is half the job. The other half is making sure they actually work.

The free AI Site Identity Checker validates all your files at once. It checks for proper formatting, required fields, and consistency across files. But here's what it can't check: whether your content is actually accurate and useful. That's on you.

Common issues the checker finds: missing required fields in llms.txt, contradictory information between files (different company names, conflicting location data), broken JSON syntax in identity.json, and robots-ai.txt rules that accidentally block the AI crawlers you're trying to reach.

Beyond the checker, test manually. Ask ChatGPT about your business before and after implementing files. Check if the responses improve. Ask specific questions your faq-ai.txt should answer. If AI is still guessing when you've given it verified answers, something's misconfigured. We've published a dedicated guide to validating your AI discovery files that covers exactly what to check and how to fix common issues.

"You can track brand presence frequency with statistical rigor, if you run prompts enough times."

Rand Fishkin, Co-founder, SparkToro, Near Media Podcast

The words doing the work there are "enough times". Ask ChatGPT about your business once and you have learned close to nothing, because the same prompt can come back differently on different days. Ask the same five questions every month and write down what you get, and after a quarter you have something you can read a trend from. It's tedious, and it's the only honest way to measure the output side of this.

Measuring the input side is easier and most people never bother. Your server already knows which AI crawlers requested which files and when. On WordPress the free plugin surfaces that log in the dashboard; on any other stack it's in your access logs. Inputs and outputs answer different questions, and you want both.

For WordPress sites specifically, these files sit in your root directory alongside robots.txt. Standard text files any host can serve. There's now a free WordPress plugin on WordPress.org that generates all 10 files from your dashboard, which removes the manual file creation entirely. Sites on our global CDN already have the infrastructure optimised for AI crawlers: proper caching rules, appropriate headers, and server config that ensures AI systems can access files without triggering rate limits. We covered the broader picture of how AI discovery files help businesses get recommended on our sister site.

Get This Sorted

Ten files. One consistent story. That's all AI needs to understand your business properly.

The sites that get this right aren't doing anything clever. They're just giving AI systems what those systems are looking for, in the format they can actually read. The 94% of businesses without these files are leaving the conversation entirely. And this isn't just our opinion: investors just valued an AI visibility tracking company at $1 billion, which tells you where the industry thinks this is heading.

Check your current score with the free checker. It takes 30 seconds. Then decide whether to tackle it yourself or get it handled properly. Either way, the clock's ticking. Every week you wait is another week AI is recommending your competitors instead.

We also explored how GEO differs from traditional SEO and what website design changes are needed for AI search. Worth reading alongside this guide if you're mapping out your full AI visibility strategy.

Frequently Asked Questions

Do I really need all ten AI discovery files?

For full AI visibility, yes. Each file serves a specific purpose the others don't cover. llms.txt handles core business identity, ai.txt manages permissions, brand.txt controls naming, and ai.json provides machine-readable interaction rules. Having just one or two leaves gaps in how AI systems understand you. Start with llms.txt, ai.txt, and brand.txt as your foundation, then add the rest.

Which AI discovery file should I create first?

Start with llms.txt. It's the foundation everything else builds on. Then add ai.txt for permissions, followed by brand.txt for naming consistency. The remaining seven can follow in any order once those core three are in place.

What happens if my AI discovery files contradict each other?

AI systems get confused and may ignore your files entirely or cherry-pick incorrect information. If llms.txt says you're in London but identity.json says Manchester, AI doesn't know which to trust. Consistency across all ten files is more important than having all ten with conflicting data.

How long does it take to create all ten files myself?

Budget 10-15 hours if you understand the specifications and have your business information organised. That breaks down to 3-4 hours researching and planning, 6-8 hours creating files, and 2-3 hours testing and fixing consistency issues.

Will these files guarantee ChatGPT mentions my business?

No. AI systems make independent decisions about what to recommend based on multiple factors. What the files guarantee is that when AI does reference your business, it has accurate information to work from. You're controlling the narrative, not forcing the mention.

How often should I update my AI discovery files?

Update whenever core business information changes: new services, pricing updates, location changes. A quarterly review works well. Set a calendar reminder and check your files against current business reality. Small additions (new FAQs, updated tech specs) can happen any time.

How are AI discovery files different from SEO?

SEO helps people find your website in search results. AI discovery files help AI systems understand and describe your business when answering questions directly, which increasingly means no website click at all. You need both working together. Perfect SEO doesn't guarantee ChatGPT describes you correctly, and perfect AI files don't guarantee Google rankings.

How do I check if my AI discovery files are working?

Use the free AI Site Identity Checker to validate formatting, required fields, and cross-file consistency. It scores your visibility out of 100 and identifies specific issues. Then test manually by asking ChatGPT about your business before and after implementing files to see if responses improve.

Do AI crawlers actually read llms.txt?

Some do, on some sites. Ahrefs analysed 137,000 sites and found 97% of llms.txt files are never requested, which is accurate for the median website: most sit on abandoned or invalid pages that nothing links to. On an actively maintained, valid, complete set the picture differs. Our own mcneece.com logged 58 fetches from 10 distinct AI crawlers across seven operators in the 30 days to 2 August 2026, with none blocked. The honest reading is that publishing files is necessary but not sufficient, and that a fetch is not the same thing as a citation.

Which AI discovery file gets read the most?

Not always the one you'd expect. llms.txt usually leads, but in one 30-day window on our own site brand.txt (the naming and terminology file) was fetched more often than llms.txt, and one OpenAI crawler requested brand.txt more than any other file. At the other end, llm.txt (the singular-spelling compatibility variant) has never been requested once. That spread is the practical argument for publishing the full set rather than only the famous file.

Your AI Visibility Starts with the Right Hosting

365i's managed WordPress hosting is built for the infrastructure AI crawlers look for: proper caching, fast response times, and server configurations that don't accidentally block the systems trying to understand your business.

Explore WordPress Hosting

Sources