AI Crawler Setup
llms.txt and robots.txt setup for the AI bots reading your site.
We check which AI crawlers your robots.txt already blocks, help you decide what should be allowed, and publish a llms.txt file so an AI assistant has a clear list of what to read first.
- A written quote within one business day
- No long-term contract
- 30-day guarantee, pay only if satisfied
What we set up
Clear rules for every AI bot touching your site.
Bot audit
We check whether your current robots.txt already blocks known AI-crawler user agents, GPTBot, ClaudeBot, PerplexityBot, Google-Extended, and CCBot among them, and flag any accidental blocks.
Allow or block decisions
For each bot we write a plain decision: readable for AI answers, usable for model training, or both, since some crawlers let you set those separately.
A draft llms.txt file
We publish a plain markdown file at your site root listing the pages an AI assistant should read first, in the format currently proposed for that convention.
Log review where available
If server or CDN logs are accessible, we check which AI bots are already visiting and how often, so decisions rest on real traffic instead of guesses.
The scope, in plain terms.
What we need from you
DNS, host, or CMS access to edit robots.txt and add a file at your root, plus server or CDN logs if you want the traffic review.
What you get
An updated robots.txt, a llms.txt file at your root, and a one-page summary of which AI bots are allowed, blocked, or left undecided.
How we prove it works
We confirm the new robots.txt rules parse correctly and that llms.txt returns a normal response at your root. No AI company publishes a compliance confirmation for that file today, so we say plainly that we cannot prove any assistant reads it, only that it is published correctly.
What this does not cover
Rewriting the content those crawlers read, or negotiating a licensing or paid-crawl agreement with an AI company, which is a business decision outside a technical setup.
How the engagement runs.
01
Bot and log audit
We check the current robots.txt against known AI-bot user agents and pull any available crawl logs.
02
Decide the rules together
We walk through each bot with you and agree on allow, block, or leave open, based on what you actually want AI tools doing with your content.
03
Publish and confirm
The updated robots.txt and the new llms.txt file go live, and we confirm both return correctly before calling it done.
AI crawler setup FAQ.
No. It is a proposed, unofficial convention, and adoption across AI companies is inconsistent today. We treat it as a low-cost hedge worth having, not a guarantee that any assistant will use it.
It depends on what you want. Blocking a training crawler like Google-Extended keeps your content out of some model training sets, but it can also keep an AI answer engine from citing you at all. We walk through that tradeoff bot by bot instead of applying one blanket rule.
It should not, since AI-bot rules in robots.txt are separate from the rules that govern Googlebot. That said, a badly written robots.txt has accidentally blocked Googlebot itself before, so checking for that mistake is part of this job.
We cannot promise that. No AI company publishes exactly how it selects sources. What this setup controls is whether your content is readable and machine-parseable at all, which is the part actually within your control.
This is scoped to your site and current setup, so we send a written price after a quick look at your robots.txt and platform, free, within one business day.
Find out what your robots.txt is telling AI bots.
Send us your domain and we will check what is already blocked, recommend changes, and quote the llms.txt setup, free, within one business day.