Glossary
robots.txt
robots.txt is a file at the root of a site that tells crawlers which paths they may request. It controls crawling, not indexing, and it is a convention that well-behaved crawlers follow voluntarily.
Glossary
robots.txt is a file at the root of a site that tells crawlers which paths they may request. It controls crawling, not indexing, and it is a convention that well-behaved crawlers follow voluntarily.
Blocking a path stops compliant crawlers fetching it, but a blocked URL can still appear in results if other pages link to it: the engine knows the URL exists and simply cannot see its content. To keep a page out of an index, allow it to be crawled and use a noindex directive.

Deciding which of them may fetch your content is now a strategic choice: blocking them protects content from being used, and generally removes you from being cited in the answers those systems produce.
A stray disallow left after a migration can remove an entire section from crawling silently, and the damage is usually noticed weeks later in traffic rather than immediately.
No. It stops compliant crawlers fetching the page, but the URL can still be listed if other pages link to it. To keep a page out of an index, let it be crawled and serve a noindex directive.
It is a genuine trade. Blocking protects content from being used in training or answers, and generally removes the possibility of being cited by those systems. It should be a deliberate decision, not a default.
robots.txt is a file at the root of a site that tells crawlers which paths they may request. It controls crawling, not indexing, and it is a convention that well-behaved crawlers follow voluntarily.