Used to improve Amazon's products and services. Amazon says the data may be used to train Amazon AI models.
Source: developer.amazon.comOperator
Amazon
Job
Model training
Collects pages for training future models. Slower to pay off, and separate from search and citations.
robots.txt
Follows robots.txt
Amazon says its automated crawlers respect the Robots Exclusion Protocol, honouring user-agent and allow/disallow directives.
Source: developer.amazon.comMozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amazonbot/0.1) Chrome/W.X.Y.Z Safari/537.36
Add one of these groups to the robots.txt at the root of your domain.
A crawler obeys only the most specific group that names it (RFC 9309), so a named group does not inherit the rules under User-agent: *. Repeat any paths you keep private, as the example does with /admin/.
User-agent: Amazonbot Allow: / Disallow: /admin/
User-agent: Amazonbot Disallow: /
Read from the robots.txt we serve. We want AI engines to read and cite this site, so every AI crawler is allowed everywhere except the app, the API and the sign-in pages (/dashboard/, /api/, /login, /signup, /sign-in, /sign-up). The file also declares Content-Signal: ai-train=yes, search=yes, ai-input=yes.
User-agent: Amazonbot Allow: / Disallow: /dashboard/ Disallow: /api/ Disallow: /login Disallow: /signup Disallow: /sign-in Disallow: /sign-up
See the whole file. The same rules repeat in every group, because a crawler that matches a named group ignores the wildcard one.
Amazon's crawlers do not support the crawl-delay directive. They honour a noarchive robots meta tag as a signal not to use the page for model training.
Sources, checked 2026-09-24