Pacific Blue Software Logo

How to Unindex Pages from Search Engines

How to Keep Pages Out of Search Engines with noindex

Not every page on a site belongs in search results. A search engine will happily list anything it can reach, including pages that only make sense after a visitor has done something first, or pages that were never meant for the public at all.

Typical candidates are:


Crawling is Not Indexing

Before choosing a method it helps to separate two things search engines do:

robots.txt controls crawling. noindex controls indexing. They are not interchangeable, and using the wrong one is the most common reason a page refuses to leave the results.


Method 1 - The Robots Meta Tag

For an HTML or PHP page, add a robots meta tag inside <head>:

<meta name='robots' content='noindex'>

The content value can combine rules, separated by commas:

To target only Google, use name='googlebot' instead of name='robots'.

If your pages share one head file, the cleanest approach is a flag that each page sets before including it, so the tag is written in one place only:

// in the page
$No_Index = true;
require_once $Lib.'head.php';

// in head.php
if (isset ($No_Index))
   {$Robots = 'noindex';}
else
   {$Robots = 'all, index';}
...
<meta name='robots' content='<?php echo $Robots;?>'>

Method 2 - The X-Robots-Tag Header

A PDF or a RAR file has no <head> to put a meta tag in. For these, send the same rule as an HTTP response header instead. Search engines treat X-Robots-Tag exactly like the meta tag.

With .htaccess (Apache)

Place this in the .htaccess of the folder holding the files. It needs Apache's mod_headers module enabled:

<IfModule mod_headers.c>
  <FilesMatch "\.(pdf|rar|zip|docx?)$">
    Header set X-Robots-Tag "noindex, nofollow"
  </FilesMatch>
</IfModule>

Leave out the FilesMatch lines to apply the header to everything in the folder.

With PHP

If a PHP script delivers the file, send the header before any output:

header ('X-Robots-Tag: noindex, nofollow');

Method 3 - robots.txt, and Why It Is Not Enough

A Disallow line in robots.txt asks crawlers not to fetch a URL:

User-agent: *
Disallow: /test/

That stops the page being read, but not being listed. If other pages link to the URL, Google can still show it in results - usually with no description, because it was never allowed to read the page.

Worse, a blocked page can never be seen carrying a noindex tag. Google states plainly that if a page is blocked by robots.txt, the crawler will never see the noindex rule and the page can still appear in results. So:


Pages That Should Not Exist at All

noindex is for pages that stay online but stay out of search. If the page is deleted, return 410 Gone (or 404) instead - search engines drop it on their next visit. If it is private, put it behind a password - a crawler that cannot log in cannot index it. Both are covered in the related articles below.


Speeding Up Removal

Nothing changes until the search engine visits the page again, which can take days or weeks for a page that rarely changes. To hurry it along:

Search engines other than Google and Bing may treat these rules differently, so a page can linger elsewhere for longer.


Testing


Summary


Useful References


Stop Search Engines Listing Pages You Want Kept Private


Back to Articles for Developers
Back to Articles on Websites
How to Generate HTTP Errors with htaccess
Protect Directories with XAMPP / Apache
How to Hide Download Links with PHP

If you found this useful, then please consider making a donation.

paypal
QR Code for donation Please donate if helpful