CrawlSEOBot, the crawler in crawlseo

What it fetches, how it reads robots.txt, and how to slow it down or block it.

User-Agent

crawlseo 0.2.1 and later send this User-Agent with every request:

CrawlSEOBot/1.0 (+https://crawlseo.cloud)

Earlier versions send this one:

CrawlSEOBot/1.0 (+https://crawlseo.dev; self-hosted SEO audit)

In robots.txt both are CrawlSEOBot.

Who runs it

CrawlSEOBot is the crawler of crawlseo, an open-source SEO app under the MIT license that people install on their own servers. Each install is run by its owner, so the requests come from their server, not from us. We don’t see who runs an install or which sites it crawls, and we can’t stop someone else’s install.

A crawl starts when someone using an install asks for one, from its dashboard or its MCP server. The code has no crawl schedule, and an install runs one crawl of a site at a time.

What it fetches

  • Your robots.txt, once per host in each crawl, before its first request to that host.
  • Your sitemap: the Sitemap lines in robots.txt, then /sitemap.xml and /sitemap_index.xml, until one of them lists pages. From a sitemap index it reads up to 5 sitemaps.
  • Pages on your site: the home page, up to 100 URLs from the sitemap, and the links it finds on the pages it fetches. www and the bare domain count as one site. Links to other sites are recorded, not fetched.
  • GET requests only. It doesn’t load images, scripts or stylesheets, run JavaScript or submit forms. A linked file that isn’t HTML, such as a PDF, is downloaded and not read.

How much and how fast

  • Up to 15 requests at a time. It fetches in batches of 15 and waits 100 ms after each batch. On a host with a Crawl-delay, one request at a time.
  • 200 pages per crawl by default. The owner of the install can raise that to 2,000, never more.
  • 30 minutes per crawl by default. The owner of the install can change that with CRAWL_MAX_MINUTES.
  • 12 seconds per request, download included, and each redirect is a new request. After that it gives up on the URL.
  • Up to 5 redirects per URL, and up to 10 MB per response.

Robots.txt

From crawlseo 0.2.2, it follows your robots.txt as described here. Installs on older versions behave differently: see the release notes for 0.2.2.

  • Each host has its own robots.txt. It reads robots.txt separately for each origin (scheme, host and port), the first time it is about to fetch a URL there, and keeps it for the rest of the crawl. A change you make applies from the next crawl.
  • example.com and www.example.com count as one site, but each uses its own robots.txt. Sitemap files, and the URLs listed in them, are checked against the robots.txt of their own host too.
  • Before each request it checks the URL’s path and query against your CrawlSEOBot group, or against the * group if there is no CrawlSEOBot group. If a Disallow rule matches, the URL is not requested. Allow rules, * and $ in paths are supported.
  • Before it follows a redirect, it checks the target against the robots.txt of the target’s host. If the target is disallowed, it stops there and records the page as blocked by robots.txt.

What happens depends on how your robots.txt answers:

  • 2xx: your rules apply. It reads only the first 500 KiB.
  • 4xx other than 429, including 404, 401 and 403: there are no rules, everything is allowed.
  • 429, 5xx, a network error or a timeout after 12 seconds: it skips your whole host for that crawl.
  • A redirect: it follows up to 5 redirects to reach robots.txt, and the rules it finds apply to your original host. After more than 5, it treats your host as having no robots.txt.

The person running the crawl sees which URLs and hosts were skipped, and why.

How to slow it down

Add a Crawl-delay, in seconds, to your robots.txt:

User-agent: CrawlSEOBot
Crawl-delay: 10
  • It uses the Crawl-delay in your CrawlSEOBot group, or in the * group if there is no CrawlSEOBot group.
  • Then it makes one request at a time on that host, and waits the delay between the end of one request and the start of the next. The robots.txt request counts as the first.
  • A delay over 60 seconds is never shortened: it skips your host for that crawl instead.
  • If the next request on your host would start after the crawl’s time limit (30 minutes by default), it stops crawling your host.

How to block it

To keep it off your whole site, add this to your robots.txt:

User-agent: CrawlSEOBot
Disallow: /

To keep it out of one folder:

User-agent: CrawlSEOBot
Disallow: /private/

A CrawlSEOBot group replaces your * group for this crawler, so repeat in it any * rules you want it to follow. Installs on versions before 0.2.2 read robots.txt differently, so the sure way is to refuse requests whose User-Agent contains CrawlSEOBot at your server or CDN, for example with a 403.

Report abuse

We can’t switch off an install we don’t run. We can change how the crawler behaves, and that reaches each install when its owner updates.

  • For a crawler bug, or behaviour this page doesn’t describe, open an issue on GitHub.
  • For anything you’d rather not post in public, use the contact page.

Include your domain, the dates and times, and the User-Agent you saw. The requests come from the install owner’s server, so the abuse contact for that IP address can reach them.