Resource Usage & CloudLinux Limits

Block Googlebot from Crawling Unnecessary URLs for Bandwidth Optimization

By the Domain India teamPublished 8 min read
Knowledge base article
Contents (9 sections)

Search engine crawlers can use a surprising share of a site's bandwidth and server capacity, especially on shops and blogs where filters, sorting options, tags and search pages create thousands of near-identical URLs. The fix is to tell crawlers which URLs are not worth fetching, without hiding the pages you want in Google. This guide shows how to do that correctly with robots.txt, when to use noindex instead, how to tell the real Googlebot from impostors, and how to check the result.

Key takeaways

Use robots.txt to stop Googlebot crawling URLs with no search value: internal search results, filter and sort parameters, cart, checkout and account pages. Don't block CSS, JavaScript or pages you want ranked, and remember that robots.txt stops crawling, not indexing. Use noindex only on pages Google is still allowed to crawl. Much heavy "Googlebot" traffic is fake, so verify it before blaming Google, and block impostors by IP address.

1. First, find out who is really using the bandwidth

Before changing anything, check which bots and URLs are actually heavy. On cPanel hosting, Metrics › Awstats lists robots and the bandwidth each one used, and the raw logs in your home directory's logs folder show every request. On our cPanel servers, use the files ending in _NGINX.gz: most requests are answered by the nginx proxy in front of Apache and never appear in the Apache log. The step-by-step method is in the definitive guide to log analysis and bandwidth optimisation.

Typical findings fall into three groups:

Real Googlebot on junk URLs
Endless filter, sort, calendar or session URLs. Fix it with robots.txt (section 3).
Fake "Googlebot"
Scrapers using Google's name. Verify, then block by IP (section 6).
Other crawlers
SEO tools and AI crawlers you may not want at all. Block them by user agent in robots.txt (section 5).

2. Crawling and indexing are different things

  • Crawling is Googlebot downloading a URL. This is what uses bandwidth.
  • Indexing is Google storing a page so it can appear in search results.

robots.txt controls crawling. A URL that is blocked in robots.txt can still appear in results, usually without a description, if other sites link to it. A noindex tag controls indexing, but Google can only see it if it is allowed to crawl the page.

Don't combine Disallow and noindex on the same page

If a page is blocked in robots.txt, Googlebot never sees its noindex tag, so the page can stay in the index. To remove a page from search, allow crawling and add noindex. To save bandwidth on pages that don't matter for search, use Disallow.

3. Write a robots.txt that blocks the waste

robots.txt is a plain text file at the root of your site, for example https://example.com/robots.txt. On Domain India hosting it goes in the document root of the domain: public_html for your main domain, or the folder you chose for an addon domain. Create or edit it with the File Manager. WordPress serves a virtual robots.txt when there is no real file; an uploaded file replaces it.

A starting point for a WordPress or WooCommerce site:

text
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Disallow: /?s=
Disallow: /search/
Disallow: /cart/
Disallow: /checkout/
Disallow: /my-account/
Disallow: /*?*add-to-cart=
Disallow: /*?*orderby=
Disallow: /*?*filter_

Sitemap: https://example.com/sitemap_index.xml

How the rules work:

  • User-agent: * applies to all crawlers that obey robots.txt. A crawler follows only the most specific group that names it, so a separate User-agent: Googlebot group replaces the * rules for Googlebot rather than adding to them.
  • * matches any characters and $ marks the end of the URL, so Disallow: /*.pdf$ blocks PDF files.
  • When an Allow and a Disallow both match, Google follows the more specific (longer) rule. That is why admin-ajax.php stays reachable.
  • Rules are case-sensitive and match from the start of the path.

4. What you must not block

Safe to block
  • Internal site search results
  • Filter, sort and "view" parameters that create duplicate listings
  • Cart, checkout, account and login pages
  • Admin areas, except files the front end needs
  • Endless calendar or session-ID URLs
Never block
  • CSS, JavaScript and image files your pages use
  • Product, category, article and landing pages you want in search
  • Your sitemap
  • Pagination you rely on for Google to find older posts or products

Google renders pages like a browser. If it can't fetch your CSS or JavaScript, it may misjudge the page and rank it lower. Tags and paginated archives are a judgement call: block them only if they add nothing, and never with rules so broad that they catch real content.

5. Crawl rate, other crawlers and AI bots

  • Google ignores Crawl-delay. Bing and some other crawlers honour it, so it can still help with them. Google adjusts its crawl rate on its own when your server slows down or returns errors.
  • Urgent overload from the real Googlebot: returning 503 or 429 to Googlebot for a short time makes it slow down. Google advises against doing this for more than a day or two, because longer periods can cause URLs to drop from the index.
  • Blocking other crawlers: name them in their own group. For example, to keep AI-training crawlers out while leaving search alone:
text
User-agent: GPTBot
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: Google-Extended
Disallow: /

Google-Extended is not a separate crawler; it tells Google not to use your content for its AI models, and does not affect Google Search. Well-behaved bots obey robots.txt; abusive ones don't, and need blocking at the server (section 6).

Avoid the old trick of blocking every crawler except Googlebot (User-agent: * with Disallow: /): it also removes you from Bing and other search engines.

6. Check that "Googlebot" is really Google

Scrapers often send Googlebot's user agent. Real Googlebot comes from Google's own IP addresses. To check an IP from your logs, do a reverse lookup and then a forward lookup:

bash
host 66.249.66.1
# 1.66.249.66.in-addr.arpa domain name pointer crawl-66-249-66-1.googlebot.com.
host crawl-66-249-66-1.googlebot.com
# crawl-66-249-66-1.googlebot.com has address 66.249.66.1

The name must end in googlebot.com, google.com or googleusercontent.com, and the forward lookup must return the same IP. Google also publishes its crawler IP ranges in JSON files linked from its "Verifying Googlebot" documentation. If the IP fails the check, it is not Google: block that IP or range in .htaccess rather than blocking the user agent:

apacheconf
<RequireAll>
  Require all granted
  Require not ip 203.0.113.0/24
</RequireAll>

Don't return 403 to the real Googlebot with user-agent rules. Robots.txt does the same job without the risk of blocking pages you want indexed. If a flood of bad-bot traffic is more than .htaccess can handle, open a support ticket with the IPs and times from your logs.

7. Test and monitor

  1. Test the file.
    Open https://yourdomain.com/robots.txt in a browser to confirm the live version. In Google Search Console, the robots.txt report (under Settings) shows the version Google fetched and any errors.
  2. Test important URLs.
    Use Search Console's URL Inspection on a key product or article page to confirm it is still crawlable.
  3. Watch Crawl Stats.
    Search Console's Crawl Stats report (Settings › Crawl stats) shows requests per day, download size and response codes. Expect the change to show over days, not hours.
  4. Recheck your logs
    a week later and compare the bandwidth used by crawlers.

8. Where Domain India fits

Every Domain India shared hosting plan lets you upload your own robots.txt and .htaccess rules, and cPanel gives you Awstats and raw logs to see crawler traffic. If crawlers are pushing you towards your monthly transfer allowance, see why does my website say bandwidth limit exceeded and facts on bandwidth for allowances and upgrades.

Frequently asked questions

Does blocking a URL in robots.txt remove it from Google?

No. robots.txt stops Googlebot from crawling the URL, but the URL can still be indexed without its content if other pages link to it. To remove a page from search results, allow crawling and add a noindex meta tag or X-Robots-Tag header.

Will blocking URLs in robots.txt hurt my SEO?

Not if you block only pages with no search value, such as internal search results, filter and sort URLs, carts and account pages. It hurts if you block CSS, JavaScript, or pages you want to rank.

Does Google follow the Crawl-delay directive?

No. Googlebot ignores Crawl-delay. Bing and some other crawlers honour it. Google slows down on its own when your server responds slowly or returns 503 or 429 errors.

How do I know if a visitor claiming to be Googlebot is genuine?

Run a reverse DNS lookup on its IP address. A genuine Googlebot resolves to a name ending in googlebot.com, google.com or googleusercontent.com, and a forward lookup of that name returns the same IP.

Where do I put robots.txt on Domain India hosting?

In the document root of the domain: public_html for your main domain, or the addon domain's own folder. It must be reachable at yourdomain.com/robots.txt.

How long does a robots.txt change take to work?

Google generally refreshes its copy of robots.txt within about a day, and the effect on crawling builds up over the following days. Watch the Crawl Stats report in Search Console.

Ready to cut crawler waste? Check your logs with the log analysis guide, update your robots.txt, and compare cPanel hosting plans if you need more room. For help with abusive bots, open a support ticket.

Hosting that shows you your traffic

Awstats, raw access logs and full control of robots.txt and .htaccess on every cPanel plan.

See cPanel hosting

Ready when you are

Get cPanel hosting from ₹125/mo + GST

See plans

Was this article helpful?

Your answer helps us decide what to improve next.

Still need help? Open a support ticket and our team will reply.

Prefer an app? Add this site to your home screen.Get the app