Interpreting the Crawl Stats Report in Google Search Console
AI generated
SERP
SEO · Technical SEO · Search Console
The Crawl Stats Report
Reading Search Console's crawl statistics correctly

The crawl stats report sits tucked away in Search Console's settings and gets used far less often than the performance report or the index coverage overview, even though it is the only report that shows directly how Googlebot actually crawls a website. Understanding how to break crawl requests down by response code, file type, and purpose often surfaces technical problems weeks before they show up in rankings or index status.

14 min read Reading response codes correctly The crawl budget connection

1. Where the crawl stats report lives and what it shows

The crawl stats report sits inside Search Console under Settings, then Crawl Stats, though it is only reachable at the domain property level since it reports on the entire host rather than a single URL-prefix property. It shows total crawl requests, total downloaded data volume, and the average server response time to Googlebot requests over the last 90 days.

Unlike the performance report, which reflects how users perceive the website within search, the crawl stats report exclusively documents Googlebot's own technical behavior. That makes it the single most important tool for determining whether Google can efficiently access a website's relevant content at all, independent of how that content subsequently performs in search results.

2. Breaking down crawl requests by response code

The report's first key breakdown sorts every crawl request by the HTTP status code returned by the server. In a healthy baseline, status code 200 dominates, often complemented by a visible share of 301 redirects, which is completely normal for well-maintained websites with regular URL changes. A healthy ratio usually sits well above 80 percent successful 200 responses.

If the share of 4xx or 5xx responses climbs noticeably, Googlebot wastes a growing portion of its allocated crawl budget on requests that deliver no indexable content at all. The timeline matters here: a single, brief spike often traces back to one isolated deployment error, while a persistently elevated level usually points to structural problems such as broken internal linking or a faulty sitemap.

3. Breaking down crawl requests by file type

The second breakdown shows which file types Googlebot actually fetches, split among categories including HTML, image, JavaScript, CSS, and others. For a classic content-driven website, HTML should make up the overwhelming majority of crawl requests, since these are exactly the files carrying indexable content, while CSS and JavaScript files are only needed for rendering.

An unusually high share of JavaScript or CSS requests relative to HTML can indicate that Googlebot has to load a disproportionate number of extra resources per rendered page, for instance due to uncached, dynamically generated asset URLs with shifting versioning parameters. That ties up crawl budget without delivering additional indexable content, and it can often be reduced substantially through consistent browser caching and stable, versioned asset filenames.

4. What a spike in 404 crawls actually means

A sudden increase in 404 responses in the crawl stats report generally means Googlebot is hitting URLs that no longer exist, discovered either through internal links, a sitemap, or old external backlinks. The first step in tracking down the cause is checking the indexing report filtered to not-found pages, combined with checking which source Googlebot used to discover the affected URL.

If the URLs are old, long-removed product pages still being reported through the XML sitemap, the problem usually lies in outdated sitemap generation that is not being refreshed automatically. If instead the URLs come from internal links in your own navigation or category pages, that points to a structural linking problem that should be fixed with high priority, since it continuously burns real crawl budget on dead links instead of new or updated content.

5. Correctly classifying server errors in the report

A rise in 5xx status codes should be treated more seriously than 404 errors, since it points to genuine server problems such as overload, faulty deployments, or timeouts on particularly compute-heavy pages. Google automatically lowers the crawl rate for a domain that repeatedly returns server errors, in order to avoid straining an already burdened server further, which helps in the short term but noticeably slows the discovery and refresh of new content over the medium term.

It becomes especially critical when 5xx errors coincide with elevated average response times in the same report, since that points to a fundamental server capacity limit rather than an isolated one-off error. In that case, fixing the individual error is not enough, it is worth checking whether the underlying infrastructure, such as PHP-FPM worker limits or database connection pools, is actually sized appropriately for the current crawl volume plus concurrent user traffic.

6. Understanding host status warnings correctly

Above the actual charts, the report shows a host status with a simple traffic-light rating for robots.txt fetch availability, DNS resolution, and general server connectivity issues over the last 90 days. A red marker on robots.txt fetch availability is especially critical, since after repeated failures to fetch that file, Googlebot may cautiously pause crawling of the entire domain, to avoid accidentally crawling sections that are supposed to be disallowed.

DNS resolution problems often point to an unstable or misconfigured nameserver setup and should be resolved immediately with the hosting or DNS provider, since they potentially affect reachability of the entire domain, not just individual pages. Server connectivity issues, in turn, can indicate firewall rules that unintentionally block Googlebot's IP ranges, which in practice happens surprisingly often after security hardening measures or a hosting provider switch.

7. How crawl statistics relate to crawl budget

Crawl budget refers to the limited number of requests Googlebot performs on a domain within a given time period, determined by a combination of crawl rate limit, meaning the maximum server load Google considers tolerable, and crawl demand, meaning how eager Google is to keep the domain's content refreshed frequently. The crawl stats report is the only official Google source that translates this abstract concept into concrete, observable numbers.

For most small and mid-sized websites, crawl budget is not a limiting factor in practice, since Google typically allocates far more capacity than actually gets used. It becomes relevant mainly for very large websites with several hundred thousand URLs, such as extensive e-commerce catalogs with many filter combinations, where inefficient crawl behavior can genuinely delay the discovery of new or updated pages noticeably.

8. Crawl purpose: discovery versus refresh, and Googlebot type

The report additionally distinguishes between crawling for discovery, when Googlebot visits a URL for the first time, and crawling for refresh, when an already known URL gets fetched again to detect changes. A healthy ratio depends heavily on a website's maturity and update frequency: newer websites naturally show a higher discovery share, while established, frequently updated blogs usually show a high refresh share.

Traffic can also be broken down by Googlebot type, including Smartphone Googlebot as the primary crawler since the complete shift to mobile-first indexing, plus separate crawlers for images, video, and AdsBot. A surprisingly high share of Desktop Googlebot requests on a website with no separate desktop targeting can hint at a technical problem in how the primary mobile crawler is being detected, and deserves a closer look at server configuration.

9. A practical action plan for anomalies

For any noticeable change in the crawl stats report, a structured approach pays off: first check whether the change coincides with a known deployment, a migration, or a robots.txt edit, since many anomalies get fully explained that way. Only once no obvious internal trigger turns up does it make sense to dig deeper through server logs, which, unlike the aggregated Search Console report, name every single affected URL concretely.

Long term, it is worth checking the crawl stats report not only reactively when problems appear, but establishing it as a fixed part of a monthly technical SEO check, similar to how many teams already handle the index coverage report. Anyone who knows their own website's typical response time, response code distribution, and file type distribution spots deviations far faster than someone who only consults the report sporadically without a baseline for comparison.

Signal in the report Likely cause Urgency First check
Spike in 404 responses Outdated sitemap or dead internal links Medium Filter the indexing report for not-found pages
Spike in 5xx responses Server overload or faulty deployment High Check server logs and response times
Red robots.txt availability File unreachable or server blocking access Very high Fetch robots.txt manually and via a test tool
High JS/CSS share Missing caching or unstable asset URLs Low Check caching headers and versioning
DNS resolution errors Unstable nameserver configuration Very high Contact the DNS provider

Mironsoft

Technical SEO, content strategy, and sustainable ranking

Visibility that doesn't disappear with the next Google update?

We review existing websites for technical SEO issues, weak content structure, and missing structured data, then build a foundation that supports sustainable, not just short-term, organic growth.

Technical SEO Audit

Systematically checking crawling, indexing, Core Web Vitals, and structured data.

Content Strategy

Building search-intent-based content instead of keyword stuffing for real relevance.

Onpage Optimization

Shaping meta data, internal linking, and page structure consistently and scalably.

10. Summary

Crawl Stats Report: Key Takeaways

Response codes

Above 80 percent 200 responses counts as a healthy baseline, spikes in 4xx/5xx are warning signs.

File types

HTML should make up the bulk of crawl requests on content-driven websites.

Host status

Red markers on robots.txt or DNS require immediate action, not just observation.

Crawl budget

Mainly relevant for very large websites, not a limiting factor for most smaller sites.

11. FAQ: Crawl Stats Report: Key Takeaways

1Where do I find the crawl stats report in Search Console?
Under Settings, then Crawl Stats, though only for a domain property, since the report covers the entire host and is not available for a plain URL-prefix property.
2What share of 200 responses counts as healthy?
A good rule of thumb is well above 80 percent successful 200 responses, complemented by a normal, moderate share of 301 redirects on well-maintained websites.
3What should I do about a sudden spike in 404 crawls?
First check the indexing report to see which source Googlebot used to discover the affected URLs. If they come from the sitemap, its generation needs updating; if they come from internal links, those links need removing or fixing.
4Why are 5xx errors more serious than 404 errors?
5xx errors point to genuine server problems such as overload or faulty deployments, while 404 errors usually just involve outdated links. Google also automatically lowers the crawl rate after frequent 5xx errors.
5What does a red marker on robots.txt availability mean in the host status?
It means Googlebot repeatedly failed to fetch robots.txt, which can lead to a cautious, temporary pause of crawling for the entire domain.
6Is crawl budget a relevant topic for every website?
No, for most small and mid-sized websites the crawl budget Google allocates is more than sufficient. It becomes relevant mainly for very large websites with several hundred thousand URLs.
7What is the difference between discovery crawling and refresh crawling?
Discovery crawling refers to the first visit to a previously unknown URL, while refresh crawling refetches an already known URL to detect changes to its content.
8Why can a high JavaScript and CSS share of crawl requests be a problem?
Because that share ties up crawl budget without delivering additional indexable content. Common causes are missing caching or dynamically generated asset URLs with shifting versioning parameters.
9How do I recognize that a deployment caused an anomaly in the report?
By the anomaly coinciding in time with the known deployment date. That is why it always pays to check against your own deployment and change log first, before starting a deeper investigation.
10How often should the crawl stats report be checked?
Ideally as a fixed part of a monthly technical SEO check, not only reactively once problems are already visible, since that makes deviations from a website's typical values much easier to spot.