The Technical SEO Audit Checklist We Run Before Any Campaign
Digital Marketing

The Technical SEO Audit Checklist We Run Before Any Campaign

Hana Suzuki19 July 2024 14 min read

Before touching a single keyword or writing a word of new content, every technical SEO audit checklist we run before starting a campaign begins in the same unglamorous place: can search engines actually crawl and index this site properly, and is anything currently getting in the way of that. It sounds like a low bar, but a surprising share of underperforming websites, including some with genuinely good content, are quietly sabotaged by a crawl-blocking robots.txt rule left over from a staging environment, a misconfigured canonical tag pointing every page to the homepage, or a sitemap that has not been regenerated since a site migration eighteen months ago. Skipping straight to content or link building on a site with foundational crawl or indexation problems is like advertising a shop whose front door is locked, and no amount of additional marketing spend fixes a technical block sitting upstream of everything else. This checklist reflects the order we actually work through on a real audit, not a randomly sorted list of SEO best practices, because sequence matters: crawlability and indexation issues should be diagnosed and fixed first, since they can mask or distort the results of every other check that follows.

The first practical step is pulling the robots.txt file directly and reading it line by line rather than trusting that it looks fine at a glance. It is common to find a Disallow rule still blocking an entire staging subdirectory that got copied into production during a redesign, or overly broad rules blocking parameterised URLs that were meant to target a narrow pattern but instead catch legitimate pages. Cross-referencing this against Google Search Console's Page Indexing report reveals the practical impact, specifically the "Blocked by robots.txt" and "Excluded by noindex tag" categories, which name the exact URLs affected rather than leaving it to guesswork. From there, the XML sitemap needs its own check: does it exist, is it referenced correctly in robots.txt, does it actually validate as proper XML, and critically, does it match the current site structure rather than listing thousands of URLs that now 404 or redirect. A sitemap full of dead URLs does not just fail to help, it actively signals a poorly maintained site to search engines and wastes a portion of the crawl budget allocated to the domain, an issue that matters increasingly as site size grows into the tens of thousands of pages.

Indexation auditing goes a layer deeper than the sitemap check, comparing how many pages the site actually has against how many Google has indexed, using a site: search alongside the indexed page count in Search Console, since the two sources occasionally disagree and both are useful signals. Large gaps in either direction point to real problems: far fewer indexed pages than exist on the site suggests crawl or quality issues suppressing indexation, while far more indexed pages than the site is supposed to have often points to duplicate content being generated by faceted navigation, URL parameters, or a staging environment that accidentally got crawled and indexed. Canonical tag auditing catches a related but distinct problem, checking that every page's canonical tag correctly self-references or points to the genuinely preferred version of duplicate or near-duplicate content, since a broken canonical implementation, self-referencing canonicals missing entirely, or canonicals pointing to the wrong page entirely, can quietly cause Google to index the wrong version of a page or consolidate ranking signals in an unintended direction. This step alone regularly surfaces issues that have been silently capping a site's organic performance for months or years without anyone noticing, because nothing about a canonical misconfiguration produces an obvious error message anywhere in a typical analytics dashboard.

Site architecture and internal linking come next, evaluated primarily through a full crawl using a tool like Screaming Frog or Sitebulb, which builds a complete map of the site as search engines actually see it rather than as a human navigating the visible menu experiences it. This surfaces orphan pages, pages that exist and may even rank for something, but have no internal links pointing to them, making them harder for search engines to discover and reinforcing to Google that the site itself does not consider them important. It also reveals click depth, how many clicks from the homepage it takes to reach a given page, since pages buried five or six clicks deep tend to get crawled less frequently and rank less well even when the content itself is strong, simply because internal link structure is one of the clearest signals of relative page importance available to a search engine. Redirect chains, multiple hops before a URL reaches its final destination, and redirect loops both waste crawl budget and degrade user experience, and a full crawl audit catches these systematically rather than relying on spot-checking individual URLs a client happens to mention.

Page speed and Core Web Vitals form the next major section of the checklist, and this is an area where the specific metrics Google actually uses have shifted meaningfully over the past few years and where generic "make it faster" advice is not specific enough to act on. Largest Contentful Paint, measuring how quickly the main content of a page becomes visible, Interaction to Next Paint, which replaced First Input Delay as the responsiveness metric in March 2024 and measures how quickly a page responds to actual user interaction throughout the page's lifecycle rather than just the first click, and Cumulative Layout Shift, measuring visual stability as a page loads, together make up the Core Web Vitals Google uses as a ranking signal and, more importantly, as a genuine proxy for real user experience. Running PageSpeed Insights and, ideally, pulling real-world field data from the Chrome User Experience Report rather than relying solely on lab data gives a more accurate picture, since lab tests run under artificial, consistent conditions can miss the real-world variability caused by different devices, network conditions and geographic distance from the server that actual visitors experience. Common culprits behind poor scores include unoptimised images served at far larger dimensions than displayed, render-blocking JavaScript and CSS loaded before critical content, and third-party scripts, chat widgets, tracking pixels, ad tags, that individually seem harmless but collectively add seconds of load time.

Mobile usability gets its own dedicated check even though Core Web Vitals partially overlaps with it, because Google has run mobile-first indexing since 2019, meaning the mobile version of a site is what actually gets crawled and evaluated for ranking purposes in the overwhelming majority of cases, regardless of how much desktop traffic a business believes it receives. This means checking that mobile navigation is genuinely usable rather than a cramped, shrunk-down version of desktop, that tap targets are large enough to hit reliably with a thumb rather than requiring pixel-perfect precision, that text is legible without pinch-zooming, and that any content hidden behind accordions or tabs on mobile is still fully accessible to and crawlable by search engines, since content that requires a user interaction to reveal can sometimes be treated differently than immediately visible content depending on implementation. It is also worth specifically testing whether any content available on desktop is missing entirely from the mobile version, a genuine and recurring problem on responsive sites that hide certain elements at smaller breakpoints for design reasons without realising that mobile-first indexing means Google may only ever see the mobile version of that page.

JavaScript rendering issues deserve specific attention on modern sites built with frameworks like React, Vue or Angular, because while Google's crawler can execute JavaScript, it does so in a separate rendering step that is more resource-intensive and occasionally less reliable than parsing static HTML directly. Using Google Search Console's URL Inspection tool to view how Googlebot actually renders a given page, and comparing that rendered output against what a human visitor sees in a browser, catches cases where critical content, navigation links, or metadata are not appearing in the rendered version Google actually indexes. This is a particularly common and particularly damaging issue on single-page applications built without proper server-side rendering or static generation, where content that loads perfectly for a human visitor after a JavaScript execution delay may never render in time for, or at all during, Google's rendering pass, effectively making that content invisible to search despite being fully visible to any human tester. Sites experiencing an unexplained gap between apparently good content and poor organic performance should check this specifically, since it is one of the harder problems to diagnose without deliberately looking for it, and one of the easier ones to fix once correctly identified, typically through server-side rendering, static site generation, or dynamic rendering for crawlers.

Structured data implementation gets audited both for technical correctness and for genuine relevance to the content it describes. Running pages through Google's Rich Results Test and the Schema Markup Validator catches syntax errors, but the more valuable check is whether the structured data actually matches the visible content on the page, since Google has become increasingly strict about penalising or simply ignoring structured data that misrepresents what a page actually contains, a practice sometimes attempted to game rich result eligibility. Common structured data types worth checking depending on the site's content include Organization and LocalBusiness markup for company and location information, Product and Review schema for e-commerce, Article and Author markup for content-heavy sites, FAQPage for genuine frequently-asked-question content, and BreadcrumbList for navigation context. It is worth noting that Google has scaled back which structured data types actually trigger visible rich results in search over time, removing FAQ and HowTo rich results for most sites in 2023, so part of a proper audit is checking current documentation on what structured data still produces a visible search result enhancement versus what is now purely a semantic signal without a direct visual payoff, since implementation effort should be weighted accordingly rather than assumed to always produce the same visible benefit it once did.

Duplicate content auditing covers more ground than most site owners expect, extending well beyond obviously copied text to include near-duplicate issues created by URL parameters generating multiple versions of the same page, HTTP and HTTPS versions of a site both resolving without a proper redirect, www and non-www versions both live simultaneously, and printer-friendly or session-ID-appended URL variants that all effectively serve identical content under different addresses. Each of these dilutes ranking signals across multiple URLs that should be consolidated into one authoritative version, and the fix is almost always some combination of proper 301 redirects, correctly implemented canonical tags, and, where appropriate, parameter handling configuration. E-commerce sites face a particularly acute version of this problem through faceted navigation, filtering products by size, colour or price range, which can generate an enormous number of crawlable URL combinations, most of which offer no unique value to a search engine and simply fragment crawl budget and ranking signal across near-identical pages. A proper audit maps out exactly how many of these parameter combinations are currently crawlable and indexed, and recommends a deliberate policy, canonicalising to the base category page, blocking via robots.txt, or selectively allowing genuinely valuable filtered views, rather than leaving the default behaviour of whatever platform generated the site unexamined.

International SEO checks apply specifically to sites serving multiple countries or languages, and hreflang implementation is the area most commonly found broken even on otherwise well-built multilingual sites. Hreflang tags tell search engines which language and regional version of a page to show to users in a given locale, but they require exact reciprocity, meaning if a French page references an English version via hreflang, that English page must reference the French version back, and any broken or missing reciprocal link can cause the entire hreflang cluster for that page to be ignored or misinterpreted. Auditing this properly means checking every hreflang tag against its claimed counterpart pages, verifying correct language and region codes following the ISO 639-1 and ISO 3166-1 standards, since a surprisingly common error is using an incorrect or invalid country code that silently fails rather than throwing an obvious error. For sites not operating internationally, this entire section is simply skipped, but it is worth explicitly confirming during the audit that no dangling or accidental hreflang implementation exists from a previous agency or plugin, since leftover broken hreflang tags from an abandoned expansion attempt can quietly confuse search engines about which version of a page to serve in which market long after the original international plans were shelved.

Log file analysis is the more advanced, often-skipped section of a thorough technical audit, but it is the only method that shows exactly how search engine crawlers actually behave on a site rather than inferring behaviour from indirect signals like Search Console reports. Pulling raw server logs and filtering for verified Googlebot and Bingbot requests reveals which pages actually get crawled, how frequently, and whether crawl budget is being disproportionately spent on low-value pages, filtered category variations, old paginated archives, parameter-heavy URLs, at the expense of newer or more important content that gets crawled rarely or not at all. This is particularly valuable for larger sites, generally above a few thousand pages, where crawl budget genuinely becomes a limiting factor, since Google does not crawl every page on every site with equal frequency or thoroughness, and understanding the actual pattern of crawler behaviour reveals optimisation opportunities that no other single audit technique surfaces as directly. Smaller sites under a thousand pages rarely have genuine crawl budget constraints and can usually skip detailed log analysis in favour of the other checks, but it is worth knowing this technique exists and escalating to it specifically once a site has grown large enough, or complex enough in its URL structure, for crawl efficiency to plausibly be limiting organic performance.

Security and HTTPS implementation get checked even on sites that migrated to HTTPS years ago, because partial or degraded HTTPS implementations are more common than expected. This includes checking for mixed content warnings, where a page loaded over HTTPS still calls some resources, images, scripts, over unencrypted HTTP, which browsers flag and which can affect both user trust signals and page functionality. It includes verifying that HTTP versions of every URL properly 301 redirect to their HTTPS equivalent rather than returning a 200 status or, worse, different content, and checking that the SSL certificate itself is current, correctly configured for all subdomains that need coverage, and not approaching an expiration date that could cause an unexpected outage. Security headers, including HTTP Strict Transport Security, Content Security Policy and X-Content-Type-Options, are increasingly relevant both for genuine security posture and because some of these signals feed indirectly into broader trust and safety evaluations search engines make about a domain. None of this is exotic or expensive to fix once identified, but an audit that skips checking HTTPS implementation quality on the assumption that "the site is already on HTTPS so this box is ticked" misses a real and fixable category of issues.

Image optimisation deserves a dedicated pass beyond its role in page speed scoring, covering alt text completeness and accuracy for both accessibility and image search visibility, appropriate file format selection, with modern formats like WebP and AVIF offering substantially better compression than legacy JPEG or PNG for equivalent visual quality, and correct implementation of responsive images so that a phone downloads an appropriately sized file rather than the same large image served to a desktop monitor. Missing or generic alt text, "image1.jpg" or a blank attribute, represents a missed opportunity for image search traffic on sites where that channel matters, product photography for e-commerce being the clearest example, and it is also a genuine accessibility failure for users relying on screen readers, which is reason enough to fix it independent of any SEO benefit. Lazy loading implementation is worth checking too, ensuring images below the fold load only as a user scrolls to them to improve initial page load speed, while confirming that above-the-fold images critical to Largest Contentful Paint are not being lazy-loaded in a way that actually delays their appearance and hurts the very Core Web Vitals metric the site is trying to improve.

The tool stack behind all of this matters less than the discipline of actually running through every check consistently, but it is worth naming what we actually reach for, since the right combination genuinely saves hours over trying to do this manually. Screaming Frog or Sitebulb handle the full-site crawl, redirect mapping, and canonical and hreflang auditing described above, and both scale to hundreds of thousands of URLs with the right configuration. Google Search Console remains the single most important free source of truth for how Google itself sees indexation, crawl stats and Core Web Vitals field data, and should always be cross-referenced against third-party crawl data rather than treated as optional. Ahrefs or Semrush add their own site audit modules that overlap somewhat with a dedicated crawler but are useful for tracking issues over time and correlating technical fixes against subsequent ranking or traffic changes. For log file analysis specifically, tools like Screaming Frog's Log File Analyser or JetOctopus make the raw server log data manageable without needing custom scripting, though larger enterprise sites sometimes justify a more bespoke data pipeline given the sheer volume of log data involved. None of these tools substitute for understanding what each check is actually revealing and why it matters, which is the point of running through the checklist in a fixed, deliberate order rather than clicking through a tool's default report and reacting to whatever it happens to flag first.

Pulling all of this together into an actual audit report is where a checklist approach earns its value over an unstructured review, because the output needs to prioritise findings by actual impact and effort rather than presenting forty issues in the order they happened to be discovered. A well-structured audit groups findings into critical issues actively suppressing organic performance, meaningful improvements worth scheduling into the next development sprint, and minor polish items that can wait, rather than treating a missing alt tag on a rarely-visited page with the same urgency as a robots.txt rule blocking the entire product catalogue. It is also worth building the audit as a living reference rather than a one-time deliverable, since technical SEO is not a problem that gets solved once and stays solved, platform updates, new content types, plugin changes and site redesigns all reintroduce the same categories of issue over time, which is why we re-run this full checklist before starting any new campaign rather than assuming a client's previous audit, however thorough at the time, still accurately reflects the current state of the site.