Optimized Custom robots.txt for Blogger to boost Blog SEO in 2023

Every search engine crawling bot first interacts with a website’s robots.txt file and its crawling rules. This means that the robots.txt file plays a pivotal role in the search engine optimization (SEO) of a Blogger blog. This article will guide you on how to create a well-optimized custom robots.txt file for Blogger and how to understand the implications of blocked pages reported by Google Search Console.

What are the functions of the robots.txt file?

The robots.txt file communicates to the search engine which pages should and shouldn’t be crawled. This allows us to control the functioning of search engine bots. In the robots.txt file, we declare user-agent, allow, disallow, and sitemap functions for search engines like Google, Bing, Yandex, etc.

Usually, we use robots meta tags to index or noindex blog posts and pages throughout the web. And robots.txt to control the search engine bots. You can allow the complete website to crawl, but it will exhaust the crawling budget of the website. To save the crawling budget of the website, you have to block the website’s search, archive, and label sections.

Default Robots.txt file of the Blogger Blog

To create an optimal custom robots.txt file for a Blogger blog, we first need to understand the Blogger blog structure and analyze the default robots.txt file. By default, this file looks like this:

User-agent: Mediapartners-Google
Disallow: 

User-agent: *
Disallow: /search
Allow: /

Sitemap: https://www.example.com/sitemap.xml

The first line (User-Agent) of this file declares the bot type. Here it’s

Google AdSense, which is disallowed to none(declared in 2nd line). That means the AdSense ads can appear throughout the website.
The following user agent is *, which means all the search engine bots are disallowed to /search pages. That means disallowing all search and label pages(same URL structure).
And allow tag define that all pages other than disallowing section will be allowed to crawl.
The following line contains a post sitemap for the Blogger blog.

This is an almost perfect file to control the search engine bots and provide instructions for pages to crawl or not crawl. But this file allows for indexing the archive pages, which can cause a duplicate content issue. That means it will create junk for the Blogger’s blog.

Creating an Optimal Custom Robots.txt File for a Blogger Blog

We understood how to default robots.txt file performs its function for the Blogger blog. Let’s optimize it for the best SEO.

The default robots.txt allows the archive to index, which causes the duplicate content issue. We can prevent this issue by stopping the bots from crawling the archive section. For this, /search* will disable crawling of all search and label pages.

Applying a Disallow rule /20* into the robots.txt file will stop the crawling of archive sections. The /20* rule will block the crawling of all posts, so to avoid this, we have to apply a new Allow rule for the /*.html section that allows the bots to crawl posts and pages.

The default sitemap includes posts, not pages. So you have to add a sitemap for pages located under https://example.blogspot.com/sitemap-pages.xml or https://www.example.com/sitemap-pages.xml for the custom domain. You can submit Blogger sitemaps to Google Search Console for good results.

So the new perfect custom robots.txt file for the Blogger blog will look like this.

User-agent: Mediapartners-Google
Disallow: 

#below lines control all search engines, and blocks all search, archieve and allow all blog posts and pages.

User-agent: *
Disallow: /search*
Disallow: /20*
Allow: /*.html

#sitemap of the blog
Sitemap: https://www.example.com/sitemap.xml
Sitemap: https://www.example.com/sitemap-pages.xml

/search* will disable crawling of all search and label pages.
Apply a Disallow rule /20* into the robots.txt file to stop the crawling of archive sections.
The /20* rule will block the crawling of all posts, So to avoid this, we’ve to apply a new Allow rule for the /*.html section that allows the bots to crawl posts and pages.

You’ve to replace www.example.com with your Blogger domain or custom domain name. For example, suppose your custom domain name is www.iashindu.com; then the sitemap will be at https://www.tagbot.site/sitemap.xml. In addition, you can check the current robots.txt at https://www.example.com/robots.txt.

Above file, the setting is the best robots.txt practice for SEO. This will save the website’s crawling budget and help the Blogger blog to appear in the search results. You have to write SEO-friendly content to appear in the search results.

Effects in Search Engine Console after implementing these rules in robots.txt

It’s important to note that Google Search Console may report that some pages are blocked by your robots.txt file. However, it’s crucial to check which pages are blocked. Are they content pages or search or archive pages? We can’t display search and archive pages, which is why these pages are blocked.

But if you want to allow bots to crawl the complete website, the best possible setting for robots.txt and robots meta tag, try robots meta tag and robots.txt file. The combination may exhaust the crawling budget, but the better alternative to boost the SEO of the Blogger blog.

How to Implement the Custom Robots.txt File to Blogger?

The Robots.txt file is located at the root level of the website. And in Blogger, there is no access to the root, so how to edit this robots.txt file? You can access root files like robots.txt and X-header Tags under the setting section of Blogger.

How to Edit Blogger robots.txt file — Provide custom robots.txt

How do I enable custom robots txt content in Blogger?