-
Crawling and indexation
robots.txt, meta robots and X-Robots-Tag directives, sitemap.xml, the Search Console coverage report, orphan pages, crawl budget.
important pages out of the index, junk URLs in it, sitemaps with errors or 404s.
-
Status codes and redirects
3xx chains and loops, broken internal links, soft 404s, 5xx errors, correctness of 301s after migrations, redirects for old URLs that still have backlinks.
redirects longer than one hop, internal links to 3xx/4xx, lost equity from old URLs.
-
Duplicates and canonicalization
Parameters, pagination, http/https, www, trailing slash, letter case, rel=canonical, hreflang between language versions.
the same content on several URLs, canonicals pointing to a 404 or another language, canonical and hreflang conflicts.
-
Site structure and internal linking
Click depth, link equity distribution, navigation, breadcrumbs, anchors, linking between categories, products and articles.
money pages sitting 4–5 clicks deep, key pages getting fewer links than utility pages.
-
Speed and Core Web Vitals
LCP, INP and CLS from CrUX field data and lab measurements; image weight, scripts, fonts, caching, server TTFB.
page templates that fail the “good” threshold for real users, with the specific cause of each.
-
Rendering and JavaScript
What Googlebot sees without executing scripts, content and links that only appear after JS, lazy loading, SSR/SPA, render caching.
text, links or metadata missing from the rendered HTML, and pages that depend on clicks.
-
Mobile version
Mobile-first indexing, content and link differences between desktop and mobile, viewport, tap targets, hidden blocks.
content available only on desktop and elements that break the mobile page.
-
Metadata and headings
Title, description, H1–H3 per page template; duplicate, empty and truncated values; auto-generation logic for large catalogs.
templates that give hundreds of pages the same title, pages without an H1 or with several.
-
Structured data
Schema.org types for each page type, validity, whether the markup matches visible content, rich result eligibility.
validation errors, markup with data that is not on the page, missed rich result opportunities.
-
Server and security
HTTPS and mixed content, HSTS, response headers, staging environments visible to search, response stability under load.
an exposed staging subdomain, a site copy on the IP or http, periodic 5xx in Googlebot logs.
-
Analytics and tracking
Search Console, GA4, GTM: whether data is collected correctly, whether tags fire twice, whether you can measure the effect after implementation.
events that do not fire and filters that distort traffic in reports.
-
Server logs
For large sites: how Googlebot really moves through the site, which sections it crawls most, where it spends budget, what it ignores for months.
sections the bot never visits, and thousands of requests to pages that should not be in the index.