How Do Removals Really Work in Google Search? | Search Off the Record

4.4K views
•
August 4, 2022
by
Google Search Central
YouTube video player
How Do Removals Really Work in Google Search? | Search Off the Record

TL;DR

Removing content from Google Search requires choosing a method that matches one of three stages: crawling, indexing, or serving. A robots.txt disallow controls crawling but can leave a popular URL indexed without a snippet, while noindex, 404 or 410 responses, and the removals tool address different parts of the process. Read on to understand which removal method fits each situation and why some combinations fail.

Transcript

[MUSIC PLAYING] JOHN MUELLER: Hello and welcome to another episode of "Search Off the Record," a podcast coming to you from the Google Search team, discussing all things Search and maybe having some fun along the way. My name is John, and I'm joined today by Lizzi and Gary from the Search Relations team, of which I'm also a part of. So how are you ... Read More

Key Insights

  • Removals in Google Search involve three main stages: crawling, indexing, and serving.
  • Robots.txt is primarily used for controlling crawling and does not affect indexing directly.
  • The noindex meta tag is used during the indexing stage to prevent a page from appearing in search results.
  • HTTP status codes like 404 and 410 can signal to remove a page from the index.
  • The removals tool can temporarily hide a page from search results, but additional actions are needed for permanent removal.
  • The x-robots-tag HTTP header can be used for non-HTML files like PDFs to control indexing.
  • Unavailable after tags can be used to schedule content removal, but Google may still verify content availability.
  • The cache removal tool can remove outdated content from appearing in search results, ensuring updated information is shown.

Explore YouTube Video Summarizer or Get YouTube Transcript Extractor

Questions & Answers

Q: How do removals work in Google Search?

Google Search removals involve three main systems: crawling, indexing, and serving. Tools such as robots.txt, robots meta tags, HTTP status codes, and the removals tool map to different systems, so the correct method depends on which stage you need to control.

Q: Does robots.txt remove a page from Google Search?

No. A robots.txt disallow rule blocks crawling, but Google may still index the URL of a popular page and show its title link without a snippet. The transcript says such robot URLs form a tiny minority of the index, but they can still appear.

Q: Why can combining robots.txt with a noindex directive fail?

If robots.txt prevents Google from crawling a page, Google cannot see a meta robots noindex directive in its HTML or HTTP header. This means the URL may remain in search results even though the site owner added noindex.

Q: What does a noindex meta tag do?

A noindex meta tag instructs search engines not to include a page in search results during the indexing stage. Google must be allowed to crawl the page to discover and process that directive.

Q: How do 404 and 410 status codes help remove content?

A 404 Not Found or 410 Gone response signals that a page no longer exists. These HTTP status codes can cause the page to be removed from Google's index.

Q: Does the Google Search removals tool permanently remove a URL?

No. The removals tool can quickly and temporarily hide a URL from search results, but it does not permanently remove the URL from the index. Permanent removal requires another action, such as adding noindex or returning a 404 response.

Q: How can a PDF or another non-HTML file be marked noindex?

Use an X-Robots-Tag HTTP header with a noindex directive. This provides indexing control for non-HTML content such as PDF files, where an HTML meta tag is not available.

Q: Can content be scheduled for removal from Google Search?

An unavailable_after directive can indicate a date and time after which content should be removed from search results. Google may still verify whether the content remains available, so another method may be needed when immediate removal is required.

Summary & Key Takeaways

  • Removals in Google Search are managed through a series of processes involving crawling, indexing, and serving. Each stage has specific tools and methods, such as robots.txt for crawling and noindex tags for indexing, to control what content appears in search results. Understanding these mechanisms helps in effectively managing and removing content from Google Search.

  • The removals tool offers a quick way to hide content from search results, but it does not permanently remove it from the index. To ensure content is completely removed, methods like noindex tags or returning a 404 status code should be used. The process is designed to maintain accurate and updated search results.

  • Special considerations are needed for non-HTML content and multilingual pages. The x-robots-tag HTTP header allows for indexing control of files like PDFs, while hreflang clusters in multilingual setups may require additional attention to ensure the desired removal outcomes. These tools and strategies help maintain control over what content is visible in Google Search.


Read in Other Languages (beta)

Share This Summary 📚

Explore More Summaries from Google Search Central 📚