Bot policy
This page is the contact URL in our User-Agent. If you operate a site we fetched, this is how to identify us and how to ask us to stop.
Who we are
Crawl Policy Index records site-declared AI crawl policy. We are not a search engine, we do not train models, and we do not follow links into page content.
What we fetch
At most three URLs per domain, once per daily run:
/robots.txt(primary)/llms.txt(often absent; a 404 is recorded as a fact)/sitemap.xml(top-level file only; we do not fetch child sitemaps or page URLs)
Raw bytes are stored to reproduce parses. We publish derived facts and short excerpts, not wholesale copies of your files.
How often, and from where
- At most one successful fetch per resource per domain per day
- One in-flight request per host
- User-Agent:
CrawlPolicyIndex/1.0 (+https://crawlpolicyindex.org/bot) - Current fetch server IPv4:
178.128.255.20
How to be excluded
Email exclude@crawlpolicyindex.org with the registrable domain. We honour exclusion within 24 hours and keep an auditable log. Historical blobs already stored are not used to continue fetching you; new panel versions omit the domain.
If that mailbox is not yet live, open an issue on github.com/Brinnen/crawl-policy-index with the domain to exclude.
We do not collect personal data. If any appears in a fetched file, it is purged, not analysed.