Not every page on a site belongs in search results. A search engine will happily list anything it can reach, including pages that only make sense after a visitor has done something first, or pages that were never meant for the public at all.
Typical candidates are:
Before choosing a method it helps to separate two things search engines do:
robots.txt controls crawling. noindex controls indexing. They
are not interchangeable, and using the wrong one is the most common reason a page refuses
to leave the results.
For an HTML or PHP page, add a robots meta tag inside <head>:
<meta name='robots' content='noindex'>
The content value can combine rules, separated by commas:
To target only Google, use name='googlebot' instead of
name='robots'.
If your pages share one head file, the cleanest approach is a flag that each page sets before including it, so the tag is written in one place only:
// in the page
$No_Index = true;
require_once $Lib.'head.php';
// in head.php
if (isset ($No_Index))
{$Robots = 'noindex';}
else
{$Robots = 'all, index';}
...
<meta name='robots' content='<?php echo $Robots;?>'>
A PDF or a RAR file has no <head> to put a meta tag in. For these,
send the same rule as an HTTP response header instead. Search engines treat
X-Robots-Tag exactly like the meta tag.
Place this in the .htaccess of the folder holding the files. It needs
Apache's mod_headers module enabled:
<IfModule mod_headers.c>
<FilesMatch "\.(pdf|rar|zip|docx?)$">
Header set X-Robots-Tag "noindex, nofollow"
</FilesMatch>
</IfModule>
Leave out the FilesMatch lines to apply the header to everything in the
folder.
If a PHP script delivers the file, send the header before any output:
header ('X-Robots-Tag: noindex, nofollow');
A Disallow line in robots.txt asks crawlers not to fetch a
URL:
User-agent: *
Disallow: /test/
That stops the page being read, but not being listed. If other pages link to the URL, Google can still show it in results - usually with no description, because it was never allowed to read the page.
Worse, a blocked page can never be seen carrying a noindex tag. Google
states plainly that if a page is blocked by robots.txt, the crawler will never
see the noindex rule and the page can still appear in results. So:
noindex and leave the page
crawlable.robots.txt to save crawl effort on folders that have no business
being fetched - images, libraries, scratch areas - not to hide pages.noindex to a page that is already disallowed, remove the
Disallow line until the page has dropped out.noindex is for pages that stay online but stay out of search. If the page
is deleted, return 410 Gone (or 404) instead - search engines drop it on their next
visit. If it is private, put it behind a password - a crawler that cannot log in cannot
index it. Both are covered in the related articles below.
Nothing changes until the search engine visits the page again, which can take days or weeks for a page that rarely changes. To hurry it along:
noindex, 404/410 or a password
does the permanent job.X-Robots-Tag.Search engines other than Google and Bing may treat these rules differently, so a page can linger elsewhere for longer.
<head> with the value you expect.
curl -I https://yoursite.com/files/manual.pdf
X-Robots-Tag: noindex, nofollow
robots.txt.
Please donate if helpful