Webizona

    Article

    Fixing Crawl Errors: A Technical SEO Checklist for Search Console

    · Webizona · Digital Strategy

    Google search bar on a screen, representing search engine optimisation

    In short: Crawl errors are the gap between what you publish and what Google can fetch and index. Work through them in order of impact: server errors and blocked resources first, then 404s on pages that earned traffic or links, then redirect chains, soft 404s and crawl waste. Search Console’s Pages report and the URL Inspection tool show you exactly which URL is in which state.

    Start with the Pages report

    In Search Console, Indexing → Pages lists every URL Google knows about and why it is or is not indexed. The “Why pages aren’t indexed” table is your worklist. Export each reason, because the counts hide which URLs matter: a 404 on a page that had backlinks is urgent; a 404 on a tag archive from 2019 is not. Google documents each status in its Page indexing report help.

    Server errors (5xx) and fetch failures

    Any 5xx response tells Googlebot the site is unreliable and crawling slows down. Common causes on WordPress hosting: PHP memory limits, a plugin timing out on specific URLs, overly aggressive security rules returning 406 or 403 to unfamiliar user agents, and rate limiting during a crawl spike. Check server logs for Googlebot’s IP ranges, test the failing URL with the URL Inspection tool’s live test, and fix the root cause rather than retrying.

    Not found (404) and soft 404

    A genuine 404 for content you removed deliberately is fine; Google drops it over time. The problem is 404s for content that still exists elsewhere (after a redesign or URL change) or that earned links and impressions. For each of those, add a permanent 301 redirect to the closest equivalent page, not the homepage; bulk redirects of unrelated pages to the home page get treated as soft 404s. A soft 404 is a page that returns 200 but looks empty or like an error: an archive with no posts, a product that is out of stock with no content, a search page with no results. Either give the page real content, return a true 404 or 410, or noindex it.

    Redirects: chains, loops and the wrong code

    Every redirect hop costs crawl budget and leaks a little signal. Audit with a crawler and collapse chains so each old URL goes directly to its final destination. Use 301 for permanent moves. Make sure the http→https and www→non-www redirects happen in one hop, and that internal links point at the final URL rather than relying on the redirect.

    Blocked by robots.txt and noindex conflicts

    “Blocked by robots.txt” for pages you want indexed is usually a leftover staging rule or an over-broad Disallow. “Indexed, though blocked by robots.txt” means Google found links to a page it cannot read; if you want it gone, unblock it and add a noindex tag, because robots.txt prevents Google from seeing the noindex. Check that CSS and JavaScript files are crawlable, otherwise rendering fails and the page may be judged thin.

    Crawl budget and discovered-but-not-indexed

    “Discovered – currently not indexed” means Google knows the URL but has not prioritised crawling it. On small sites it is almost always a quality or internal-linking signal rather than a budget limit: the page is orphaned, duplicated, or thin. Link to it from relevant pages, add it to the sitemap, strengthen the content and request indexing once. Remove crawl waste that competes for attention: faceted URLs with parameters, paginated archives of noindexed content, attachment pages, and tag archives with one post each.

    Keep the sitemap honest

    The XML sitemap should contain only canonical, indexable, 200-status URLs with accurate lastmod dates. Resubmit it after structural changes. If Search Console shows “Couldn’t fetch”, check that the sitemap URL itself is not blocked or rate-limited for Googlebot.

    A repeatable monthly routine

    1. Export the Pages report and diff it against last month.
    2. Crawl the site and compare status codes against the sitemap.
    3. Fix 5xx and blocked resources the same day.
    4. Redirect or 410 any new 404s that had links or traffic.
    5. Review Core Web Vitals in the Experience report; slow pages get crawled less.

    If the list is long or the causes sit in hosting and templates, Webizona’s technical SEO work starts with exactly this audit.

    Working on something similar? Talk to the team →