CheckTheCal

About our calendar checker

CheckTheCal watches iCalendar feeds on behalf of the people who subscribe to them and emails those people when an event moves, is added, or is cancelled. This page describes the software that does the fetching, for anyone who has found it in a server log.

The User-Agent it sends

CheckTheCal/1.0 (+https://checkthecal.com/bot; calendar change monitor)

Every request we make carries this string. We do not vary it, and we do not present ourselves as a browser.

What it fetches, and what it never does

It fetches one thing: a calendar URL that one of our users pasted into our site in order to watch it. Nothing else.

It does not crawl. It does not follow links out of a document, does not read your sitemap, does not look for other URLs on your site, and does not discover addresses on its own. If we are requesting a URL, a person typed it in. In this respect it behaves like Apple Calendar or Outlook fetching a feed somebody subscribed to, rather than like a search engine indexing a site.

It only ever issues GET requests, and it reads only the calendar file itself.

How often it asks

At most two requests at a time to any one hostname, with a minimum one-second gap between them, however many of our users happen to watch calendars on your site.

It sends If-None-Match and If-Modified-Since, so a 304 costs you nothing but a header. It honours Retry-After exactly as sent. On any failure it backs off exponentially, and it stops asking on a schedule that widens to hours rather than retrying in a loop.

Verifying a request is really us

A User-Agent string is trivially forged, so please do not treat one as proof. Requests genuinely from us originate from the addresses published at /bot/ips.json. Anything claiming to be us from another address is not.

robots.txt

We do not consult robots.txt, and we want to be direct about why rather than leave you to discover it. That file governs crawling — the automated discovery of URLs a site owner did not hand out. We do not discover anything; we fetch a single address at the explicit request of a person who already has it, which is the same thing their calendar app does when they subscribe.

If you would rather we did honour it for your site, say so and we will add the rule.

Allowing us

If our requests are being turned away by a firewall or bot filter, allowing the User-Agent above is usually enough. This is worth doing only for one reason: your own visitors have subscribed to your calendar through us and are currently not being told when it changes.

Blocking us

You are entitled to refuse us and we would rather you did that than rate-limit us silently. Return 403 to the User-Agent above and we will stop: our software recognises a refusal, widens its schedule to about one check a day, and tells the affected users that your site is not permitting this rather than implying their link is broken.

That is the whole mechanism, and it needs nothing from us — but if you would rather we stopped fetching your site regardless of what you return, email the address below with the hostname and we will take it off our side too.

Contact

A person reads support@checkthecal.com. If you are writing about traffic to your site, include the hostname and we will answer with exactly which of our users are watching what.