Prompt
How do I troubleshoot crawl retries and duplicate pages?
Latest observation
To troubleshoot crawl retries and duplicate pages, work through these in order:
1) Check whether retries are caused by fetch failures
Look at the crawler logs or crawl report for:
- 4xx / 5xx responses
- timeouts
- DNS / connection errors
- robots.txt blocks
- redirect loops
- server overload / rate limiting
If retries happen on the same URLs repeatedly, the issue is usually one of:
- unstable server responses
- slow page loads
- intermittent network issues
- crawler being blocked or throttled
2) Identify the exact URLs being retried
Compare:
- the original URL
- the final fetched URL
- any redirect chain
A retry may actually be caused by:
http↔httpswww↔ non-www- trailing slash differences
- uppercase/lowercase variants
- query parameters creating many URL versions
3) Check for duplicate URL variants
Duplicate pages often come from multiple URLs serving the same content:
/page/page//page?ref=.../page?utm_source=.../index.html/PAGE
Fix this by enforcing:
- one canonical URL format
- consistent redirects
rel="canonical"tags- parameter handling rules in your crawler or search platform
4) Verify canonical tags
Make sure each page has:
- a self-referencing canonical, or
- a canonical pointing to the preferred version
Common problems:
- canonical points to a different page accidentally
- canonical differs from redirected URL
- canonical chains across multiple pages
- missing canonicals on duplicate templates
5) Check sitemap and internal links
Your sitemap and internal links should only reference canonical URLs.
If the sitemap includes duplicates, the crawler may:
- revisit the same content under different URLs
- waste crawl budget
- report duplicate pages
6) Review redirect behavior
Bad redirect setups can trigger retries and duplicates:
- redirect chains
- redirect loops
- temporary redirects where permanent redirects should be used
- inconsistent redirects for mobile vs desktop
Best practice:
- use 301 for permanent canonicalization
- keep redirects to one hop when possible
7) Inspect pagination, faceted navigation, and filters
These often create large numbers of near-duplicate pages:
- category pages with sort/filter parameters
- search result pages
- pagination URLs
- session IDs in URLs
If these are not meant to be indexed, consider:
- canonical tags
- noindex rules
- parameter filtering
- robots.txt only when appropriate
8) Check server performance
If crawl retries are due to slow response times:
- inspect TTFB
- check app/server logs
- look for CPU, memory, or database bottlenecks
- review CDN/cache effectiveness
Slow pages can cause:
- crawler timeouts
- incomplete fetches
- repeated retry attempts
9) Compare content hashes or page similarity
If duplicate pages are not obvious from URLs, compare:
- page titles
- meta descriptions
- main content
- HTML fingerprints or hashes
This helps detect:
- template duplicates
- printer-friendly pages
- language/country copies
- content syndicated across URLs
10) Use crawl diagnostics tools
Depending on your platform, check:
- crawl stats
- failed fetch logs
- URL inspection
- duplicate detection reports
- canonical reports
- redirect reports
Quick fix checklist
- Resolve 4xx/5xx and timeout errors
- Standardize one preferred URL version
- Add correct canonical tags
- Remove duplicate URLs from sitemap
- Fix redirect chains/loops
- Reduce parameter-based duplicates
- Improve server response time
If you want, I can also give you:
- a step-by-step troubleshooting flowchart, or
- a Google Search Console / crawler-specific checklist.
Brands