Skip to content
ThonRetire

Machine access

How ThonRetire treats AI crawlers

Every figure on this site is published under an open licence, with the authority that issued it and the month it was last checked sitting next to it. Software that wants to read that is welcome to — whether it reads the page to answer someone right now, or reads it to learn what this place is. The line we draw is not between those two. It is between software that could one day put this source in front of a person, and software that never will.

Retrieval agents that citeAllowed on all public pages
Search engine crawlersAllowed on all public pages
Training crawlers, namedAllowed on all public pages
Scrapers and data resellersRefused — 403
/api/ and /depth/Refused to every agent
Imagesnoimageai — not ours to license onward
Content-Usage headertrain-ai=y, search=y

Allowed: agents that answer a person and link back

These run only when someone has asked a question. They fetch the page, answer, and cite the source — which is the entire transaction we want. They get the same 17-country dataset and the same public pages a search engine gets, no more and no less. They do not get /api/, because that is a metered application endpoint rather than a document.

  • OAI-SearchBot
  • ChatGPT-User
  • Claude-SearchBot
  • Claude-User
  • PerplexityBot
  • Perplexity-User
  • DuckAssistBot
  • YouBot
  • MistralAI-User

If you operate a retrieval agent that cites its sources and it is not on this list, it is missing rather than excluded — the default rule allows it, and we would rather add the name explicitly than rely on a default.

Allowed: ordinary search crawlers

These are not named in robots.txtat all, because they are allowed by default and always have been. They are listed here because several AI answer layers run on top of them rather than on their own crawler — Google’s AI Overviews are built from what Googlebot fetched, not from Google-Extended — so a page about AI access that omitted them would be describing half the picture.

  • Googlebot
  • Bingbot
  • Applebot
  • DuckDuckBot

Allowed: crawlers that collect for training

Until 6 August 2026 these were refused, and the reasoning was that they take the pages and give nothing back. That is true about any single fetch and wrong about the thing that matters. An assistant reaches a source two ways: it fetches the page because someone asked, or it already knows the place exists and therefore thinks to go and look. The second route runs through model weights. A model that has never encountered this site will go and check the sites it has heard of instead, which means refusing the training crawler quietly guarantees losing the round where names are recalled.

The cost of the new position is also real, and worth stating plainly: we give away for free something a few publishers now charge for. We think the trade is right here because the dataset is already CC-BY-4.0 — there was never a fee to forgo — and because a reference work nobody has heard of has a reputation to build rather than a price to hold.

  • GPTBot
  • ClaudeBot
  • Claude-Web
  • anthropic-ai
  • CCBot
  • Google-Extended
  • Applebot-Extended
  • meta-externalagent
  • meta-externalfetcher
  • FacebookBot
  • AI2Bot
  • cohere-ai
  • cohere-training-data-crawler
  • Amazonbot

Permission covers the text and the figures. It does not cover the photographs, which are licensed to us for use on this site and are therefore not ours to pass on.

Refused: scrapers and data resellers

These are refused in robots.txt and again with a 403 at the server, so the answer does not depend on anyone choosing to honour a text file. The test for this list is not whether something is AI. It is whether a person could ever hear of this source because of it. A reseller repackages the content and sells it onward under its own name; a scripted scraper is not an assistant at all and has no reader to tell.

  • Bytespider
  • TikTokSpider
  • PanguBot
  • PetalBot
  • Diffbot
  • omgili
  • webzio
  • ImagesiftBot
  • Timpibot
  • VelenPublicWebCrawler
  • SemrushBot-OCOB
  • Scrapy
  • python-requests
  • python-httpx
  • aiohttp
  • Go-http-client

The last several entries are the default user agents of scripted scrapers. A site that publishes its whole dataset under an open licence gives nobody a reason to scrape it — the download button is right there, and unlike a scrape it is versioned, documented and counted.

What you may do with the content

The dataset is published under Creative Commons Attribution 4.0 International (CC-BY-4.0). Use it anywhere, including commercially. The only condition is that you credit ThonRetire and link back. That covers reuse by a person and reuse by a program equally — including inside an answer generated by an assistant, provided the attribution survives.

Training is the one thing a content licence does not settle either way, so we declare it separately and in machine-readable form, using the IETF AIPREF vocabulary: train-ai=y, search=y, served as a Content-Usage header on every page. Both halves are deliberate. train-ai=y is a permission, and stands where a reservation of rights under Article 4 of the EU DSM Directive would otherwise sit — we held that reservation until 6 August 2026 and have released it on purpose. search=y is an invitation to exactly the applications that send readers back to the source, and remains the half we care about most.

If you are reading this to decide what you may do: attribution is the whole ask. The licence requires it, the header does not override it, and a figure that arrives without the authority and check date attached is worth less to your reader than one that does.

Nguyễn Thành Dân (2026). ThonRetire Retirement Destination Dataset 2026, v1.3. Verified July 2026. https://thonretire.com/data — CC-BY-4.0

If you are building on this

Start with the answer hub, where each figure sits with its issuing authority and its verification month; then the CSV and JSON files, which are served with CORS enabled and a licence header; then the method page for how each column is built, and the changelog for what changed and when. There is also an llms.txt pointing at the same things in a form meant for a machine to read first.

A frozen edition is published each year with a sha256 over its own contents, so a citation written today still resolves to the figures it was written about after the live pages move on.

Dataset last verified July 2026.