A step-by-step audit workflow for finding and fixing the technical issues that block search visibility

Most technical SEO problems are invisible in the browser. A page can look perfect and still be excluded from Google's index because of a stray noindex tag, a canonical pointing at the wrong URL, or a crawl budget quietly wasted on parameter-laden duplicate pages. A technical SEO audit is how these issues get found before they cost months of lost traffic.
This is a working checklist, not a theory guide. It follows the order an experienced technical SEO actually works through a site, with the specific thresholds and tools used in 2026.
You need four things running before the first check:
Run the crawler with JavaScript rendering enabled at least once, and compare it against a raw-HTML crawl. A gap between the two usually means content or links are missing from what Google's renderer sees on the first pass.
Crawl budget matters most on sites above roughly 10,000–50,000 URLs; smaller sites rarely hit real budget constraints.
Checklist:
Common fix: block low-value parameter combinations in robots.txt, consolidate faceted navigation with a defined canonical pattern, and prune or noindex thin auto-generated pages rather than letting them dilute crawl budget.
robots.txt checks:
robots.txt returns a 200 status at /robots.txt and is under 500KB (Google ignores rules beyond that size)Disallow rule is accidentally blocking CSS/JS files needed for renderingSitemap: directive points to the correct, live sitemap URLXML sitemap checks:
200 status — no redirects, no 404s, no noindex pages included<lastmod> dates are accurate and update when content actually changes (fabricated timestamps are increasingly ignored by Google)This is where most lost traffic hides. Run the crawl and cross-reference against Search Console's Index Coverage / Pages report.
Checklist:
noindex, Duplicate without user-selected canonical, Crawled – currently not indexed, Discovered – currently not indexed, Alternate page with proper canonical tag<meta name="robots" content="noindex"> tag (a common cause: a staging-environment robots meta tag left in production)X-Robots-Tag HTTP headers, since these override in-page meta tags and are easy to misssite: search estimates and internal page counts to sanity-check the scale of the gapwww vs. non-www, versions consistently resolve to one canonical version site-wideSince Google has used mobile-first indexing for years, the mobile version of a page is effectively the only version that matters for ranking.
Audit using field data (real user data from CrUX, visible in Search Console and PageSpeed Insights) as the primary signal — lab data (Lighthouse) is for diagnosis, not final scoring, since only field data determines the ranking signal.
| Metric | Good | Needs Improvement | Poor |
|---|---|---|---|
| LCP (Largest Contentful Paint) | ≤ 2.5s | 2.5s – 4.0s | > 4.0s |
| INP (Interaction to Next Paint) | ≤ 200ms | 200ms – 500ms | > 500ms |
| CLS (Cumulative Layout Shift) | ≤ 0.1 | 0.1 – 0.25 | > 0.25 |
Checklist:
Server logs are the only source that shows exactly what Googlebot did, as opposed to what a crawler simulation predicts it would do.
| Issue Found | Typical Fix | Relative Effort |
|---|---|---|
| Blocked CSS/JS in robots.txt | Remove disallow rules for render-critical assets | Low |
| Duplicate content without canonical | Add self-referencing or cross-page canonical tags | Low |
| Bloated sitemap with non-200 URLs | Regenerate sitemap dynamically, exclude non-indexable URLs | Medium |
| Crawl budget wasted on faceted URLs | Disallow parameter patterns, add canonical rules | Medium |
| Poor LCP from unoptimized hero image | Compress, preload, and serve responsive images | Medium |
| High INP from third-party scripts | Defer/async non-critical scripts, audit tag manager load | High |
| Thin/duplicate templated pages | Consolidate, add unique content, or noindex | High |
Search engines allocate finite crawling and indexing resources to every site. Technical issues do not just look untidy — they actively prevent content from being discovered, indexed, and ranked, no matter how strong the content itself is. A page that never gets indexed cannot rank, and a page that loads slowly on real user connections loses both rankings and conversions. Running this checklist on a recurring basis (quarterly for most sites, monthly for large or fast-changing ones) catches regressions before they compound.
How often should a technical SEO audit be run?
A full audit quarterly is reasonable for most sites. Large sites (100,000+ URLs) or sites that ship frequent template changes benefit from lighter monthly checks on crawl stats, indexation, and Core Web Vitals, with a deep audit twice a year.
Do Core Web Vitals directly affect rankings?
Yes, but as one of many ranking factors, and Google has said it is not typically a decisive one for well-optimized, relevant content. It matters more as a tiebreaker between pages of similar relevance, and poor Vitals correlate strongly with poor user engagement regardless of ranking impact.
What's the difference between "Crawled – currently not indexed" and "Discovered – currently not indexed"?
"Discovered" means Google knows the URL exists but hasn't crawled it yet, often due to crawl budget prioritization. "Crawled – currently not indexed" means Google fetched the page but chose not to index it, usually because of quality or duplication concerns — this status requires content-level fixes, not just technical ones.
Is a low crawl budget ever the real bottleneck for a small site?
Rarely. Crawl budget constraints generally only bind on very large or rapidly-changing sites. For most small-to-mid-sized sites, indexation and content-quality issues explain far more missing pages than crawl budget ever does.
Should robots.txt be used to prevent duplicate content from being indexed?
No — disallowing a URL in robots.txt only stops crawling, not indexing; a blocked URL can still appear in search results without a description if it's linked elsewhere. Use canonical tags or noindex meta tags for duplicate-content control instead, and reserve robots.txt for crawl-budget management.
A technical SEO audit is not a single scan — it's a layered investigation that starts with whether pages can be crawled, moves to whether they get indexed, and ends with whether they perform well enough to satisfy both search engines and real users. Working through crawl budget, robots.txt, sitemaps, indexation, canonicalization, structured data, mobile usability, Core Web Vitals, and log files in that order catches the majority of issues that quietly suppress organic visibility, well before they show up as a traffic drop in analytics.