CRO Audit Checklist: Running a Conversion Audit End to End
A CRO audit runs in five phases, in a fixed order. Verify the conversion tracking and set the baseline first, because every later number depends on it. Use quantitative data to find where visitors leave, segmented by device and traffic source. Use qualitative research — recordings, maps, on-page surveys, the sales team — to explain why. Turn each finding into a hypothesis with its evidence attached. Then rank the backlog, check which items can actually reach significance, and sequence a 90-day plan. Budget about a day and a half of hands-on work across a two-week window.
Who this is for
Written for whoever owns the conversion number: an in-house marketer, a founder running growth directly, a product or UX lead, or a freelancer auditing a client site. You need analytics access with a working conversion event, permission to install a recording or heatmap script, and enough traffic that the numbers stabilize — as a rough floor, a few hundred conversions a month at the step you are auditing, or the quantitative half of this checklist will produce noise you can mistake for signal. Below that threshold the checklist still works, but run it as qualitative research and say so, rather than dressing thin data up in percentages. This is not a UX or accessibility audit, and it is not a redesign brief: the output is a ranked, evidence-backed list of things to change and test, not a new set of page comps.
Phase 0 — Before you look at a single chart
Four setup steps. Skipping any of them does not make the audit faster, it makes the findings unusable — and unusable findings are worse than none, because people act on them.
Verify the conversion tracking before you trust any number
Complete the conversion yourself and watch the event fire. Confirm it fires exactly once — reload the confirmation page and check it does not fire again. Then compare the conversion count in analytics against the CRM or store admin over the same date range and time zone, and write down the gap as a percentage.
Why it matters: If the conversion event double-fires or misses a path, every drop-off rate downstream is wrong in a way that looks entirely plausible. An audit on broken tracking does not fail loudly; it produces a confident recommendation pointing at the wrong step.
Name the one conversion you are auditing, plus a guardrail metric
Pick the primary metric as far down the funnel as your data reaches — qualified leads rather than form submits, completed orders rather than carts started. Then name one guardrail a "win" is not allowed to damage: lead quality, average order value, refund rate, or sales-accepted rate. Write both down before you look at anything.
Why it matters: Removing form fields reliably lifts submissions and can lower qualified leads at the same time. Without a guardrail declared in advance, that trade gets discovered by the sales team two months later, not by the audit.
Set the window and record the baseline
Pull at least one full business cycle — four weeks minimum, longer for seasonal or long-cycle B2B. Record, in writing, the current conversion rate overall, by device, and for the top five traffic sources. Note any promotion, launch, outage or seasonal event inside the window.
Why it matters: The baseline is what makes every later claim checkable, and writing it down before you form an opinion is the only reliable defense against reading the data to fit a conclusion you already had.
Instrument the qualitative tools now, then let them run
Install session recording and heatmap tracking on the funnel pages today, before you start the quantitative work, and configure them to capture the pages you suspect rather than the whole site. Give them one to two weeks of traffic while you do phase 1.
Why it matters: Recordings are not retroactive. Installing them in week two, when the funnel data has told you exactly which step to watch, means waiting another two weeks — or watching the wrong sample because you would not.
Phase 1 — Quantitative: find out where they leave
This phase produces one output: a ranked shortlist of two or three steps that lose the most people who could reasonably have converted. Resist explaining anything yet.
Build the funnel step by step and get the drop-off rate at each one
Lay out the real sequence — entry page, key content or product page, form or cart start, form or cart completion, then the qualified downstream step. Use a funnel exploration and record the completion and abandonment rate at each transition. Use an open funnel so you can see where people re-enter partway through, which is common and quietly distorts closed-funnel numbers.
Segment the funnel by device
Run the same funnel for mobile and desktop separately. Compare the shape, not just the final rate: the step where mobile loses people is frequently not the step where desktop loses them.
Why it matters: A blended rate is an average of two different failures, and it can stay flat while both halves get worse. Most real findings start by splitting a number somebody had been reading as a single thing.
Segment by traffic source and by new versus returning
Break the funnel out by the top five sources and by first-time versus returning visitors. Look specifically for a source whose entry volume grew while its conversion rate fell.
Why it matters: A broad-match campaign, a new placement or a cheap traffic source can move the site-wide rate more than anything on the page. If the mix changed, the page is not the problem and no page test will fix it — that finding belongs in the audit as prominently as any layout issue.
Check field data for page performance on the funnel pages
Pull real-user Core Web Vitals — largest contentful paint and interaction to next paint — for each page in the funnel, on mobile, from field data rather than a lab test on your own laptop. Note any page failing on mobile.
Why it matters: Slow interaction on a form or checkout step reads in recordings as hesitation and repeated tapping, and it gets misdiagnosed as a copy or trust problem. Check the mechanical explanation before you reach for a psychological one.
Add field-level analysis to any form or checkout in the funnel
For each field: how many people start it, how many abandon inside it, how long it takes, and how often it errors. Most form tools and recording tools report this; if yours does not, a per-field focus and blur event pair will.
Why it matters: "The form converts badly" is not actionable. "Sixty percent of abandonment happens in one field, and a third of the people who reach it trigger a validation error" names the fix.
Check what the drop-off costs, in money
Estimate revenue or pipeline lost at each of the shortlisted steps: traffic reaching the step, times the drop-off, times the average value of a conversion. For ecommerce, use contribution margin rather than revenue.
Why it matters: This is what turns an ordered list of leaks into a prioritized one. The step with the worst percentage is often not the step with the largest amount of money behind it, and prioritizing by percentage sends the whole quarter to the wrong page.
Phase 2 — Qualitative: find out why
Only look at the two or three steps phase 1 shortlisted. Auditing the whole site qualitatively is how a two-week audit becomes a two-month one that concludes everything could be better.
Watch abandonment sessions, with converters as a control
Sample ten to fifteen recordings of sessions that reached the shortlisted step and left, and five that converted. Watch on the device split that phase 1 flagged. Log what you see as observations, not conclusions — "scrolled past the pricing table twice, then opened the FAQ" rather than "confused by pricing."
Why it matters: Converted sessions are easier to find and far more pleasant to watch, and they contain almost no information. The audit lives in the abandonments; the converters exist only to tell you what normal looks like.
Read the click, scroll and attention maps for the flagged pages
Look for three specific things: rage clicks and dead clicks on elements that are not interactive, the scroll depth at which half the visitors are gone, and whether the primary call to action sits above or below that line. Compare mobile and desktop maps separately.
Run a one-question survey on the flagged step
Put a single open question on the page or on exit: "What almost stopped you from [the action] today?" is more productive than asking what people liked. Leave it up until you have twenty or so answers, then read them all at once and group them.
Why it matters: Twenty answers is usually enough for the same two objections to appear repeatedly. Those two objections are the highest-quality hypothesis material in the entire audit, and they cost almost nothing to collect.
Do a heuristic pass against the page's job, and ask the people who talk to buyers
For each flagged page, check whether it answers — visibly, above the fold, in the visitor's own words — who it is for, what happens next, what it costs, and what the risk of proceeding is. Then ask sales and support what people ask them before buying, and what they say to close it.
Why it matters: Sales and support hear the real objection dozens of times a week for free. Most audits collect expensive behavioral data and never open the cheapest, richest source of qualitative evidence in the building.
Benchmark three competitors on the same step only
Go through the equivalent step on three competitor sites and note what they make easy that you make hard: fewer fields, a price shown before the form, a visible next step, no account required. Record it as a question to test, not a design to copy.
Why it matters: You are looking at either their tested winner or their untested guess, and you cannot tell which from outside. Treating a competitor's page as evidence is how teams import someone else's mistake with total confidence.
Phase 3 — Turn findings into testable hypotheses
Write each finding in one sentence, in hypothesis form
Use a fixed structure: "Because [evidence], we believe that changing [X] for [audience] will cause [metric] to [direction]. We will know we are right when [threshold]." One coherent change per hypothesis — a headline and a layout and a call to action changed together is a redesign, and a redesign teaches you nothing you can reuse.
Attach the evidence line to every hypothesis
Name the specific funnel number and the specific recording, map or survey quote behind it. If a hypothesis has no evidence line, move it to a separate list called "assumptions" and be honest that it came from taste.
Why it matters: The evidence line is what lets someone re-examine a losing test six months later and work out whether the idea was wrong or the execution was. Without it, every inconclusive result is unlearnable.
Check whether each test can actually reach significance
For each hypothesis, take the current rate at that step, decide the smallest lift worth detecting, and calculate the sample size needed per variant. Divide by the weekly traffic reaching that step. If the answer is longer than about four to six weeks, mark it now as one of three things: ship-without-testing, a cheaper test of the same question, or not testable at this traffic level.
Why it matters: This is the step that separates an audit that gets executed from one that stalls in week three. Ranking tests that need five months each produces a roadmap that looks rigorous and cannot be run.
Phase 4 — Prioritize and sequence
Score the backlog once, with one model, in one sitting
Use a consistent scoring model — impact, confidence and effort, or potential, importance and ease — and score every item with the same people in the same session. Aim for a backlog of roughly ten to fifteen items; a hundred-item list is an inventory, not a plan.
Why it matters: The value of the score is the forced comparison between items, not the number itself. Scored on different days by different people, the numbers are not comparable and the ranking is theater.
Split the backlog into fixes, tests, and open questions
Fixes are things that are broken or plainly wrong — a validation error, a dead link, a call to action below the fold on mobile. Ship those immediately without testing. Tests are hypotheses with enough traffic to resolve. Open questions are things you cannot yet answer and need more research to reach.
Why it matters: Spending three weeks of traffic A/B testing a bug is a real and common way to waste an audit. You do not need an experiment to prove that a broken thing should work.
Sequence the tests against a calendar, one experiment per step at a time
Lay the tests across ninety days so that two experiments never touch the same funnel step in the same window. Mark promotions, launches, seasonal peaks and any planned site release on the same calendar, and avoid starting a test that will be read across one of them.
Why it matters: Overlapping experiments on the same step produce interaction effects nobody can untangle afterward, and a promotional email mid-test moves the traffic mix underneath the result. Both produce a number that looks clean and means nothing.
Phase 5 — Ship the deliverable
Write it up as one document with five parts
Baseline numbers and the audit window; the funnel map with drop-off by step and segment; each finding with its evidence; the ranked backlog split into fixes, tests and open questions; and the ninety-day sequence with an owner per item. Findings that exist only in someone's memory are not findings.
Set the re-audit trigger before you close the document
Name the condition that starts the next audit: a date, the backlog dropping below a set number of runnable items, a redesign, or a shift in traffic mix past a threshold you specify. Put it in a calendar with an owner.
Why it matters: An audit is a snapshot of a site that keeps changing. Without a trigger, the backlog decays quietly until someone commissions a whole new audit to rediscover the three things the last one already found.
Common mistakes this guide prevents
- Auditing on tracking nobody verified. If the conversion event double-fires on a reloadable confirmation page, every drop-off rate below it is wrong — and the audit still reads as coherent, which is what makes it dangerous.
- Reading a blended conversion rate. Mobile and desktop, paid and organic, new and returning are different funnels with different failures, and averaging them produces a number that describes nobody and can hold steady while both halves deteriorate.
- Starting session recordings the week you need them. Recordings are not retroactive: the two weeks of behavior you wanted to review were never captured, and you either wait or watch the wrong sample.
- Watching converted sessions because they are easier to find and more enjoyable. The information is in the abandonments; converters are a control, not the study.
- Blaming the page when the traffic changed. A new broad-match campaign, placement or promotional cohort moves the rate more than any headline. Check the source mix before touching a word of copy — and put that finding in the audit, since no page test can fix it.
- Writing hypotheses that are really to-do items. "Redesign the pricing page" cannot be evaluated afterward; it names no evidence, no audience, no metric and no expected direction, so whatever happens next teaches nobody anything.
- Ranking a backlog without checking runnability. At typical B2B traffic, a large share of ranked items would need months each to resolve, and a roadmap built from them looks rigorous right up until it stalls in week three.
- Optimizing the conversion event instead of the qualified outcome. Cutting form fields lifts submissions and can lower qualified leads at the same time — which is why the guardrail metric gets named in phase 0, not discovered by sales in month three.
- Testing a bug. Broken validation, a dead link, a call to action below the mobile fold: fix it and move on. Three weeks of traffic spent proving that a broken thing performs worse than a working one is three weeks gone.
- Copying a competitor's layout on the grounds that it must have been tested. From outside you cannot distinguish their proven winner from their untested guess, and importing the second one is indistinguishable from research right up until the result comes back.
- Auditing a step with too little traffic and reporting it in percentages anyway. Below roughly a few hundred conversions a month at the audited step the quantitative half will not stabilize; run it as qualitative research and label it honestly.
- Delivering the audit as a document instead of a backlog. Findings with no ranking, no owner and no sequence get read once, agreed with warmly, and never executed.
Common Questions
As a working floor, a few hundred conversions a month at the step you are auditing. Below that, funnel percentages swing enough week to week that you can read noise as a trend, and almost nothing will reach statistical significance in a testable window. The checklist still has value at lower volume — the qualitative phase, the heuristic pass and the field-level form analysis all work — but report it as research rather than as measured rates.
Roughly a day and a half of hands-on work, spread across about two weeks of calendar time. The gap is not padding: session recordings and an on-page survey need real traffic to accumulate before there is anything to review, which is exactly why they get instrumented in phase 0 rather than when you need them.
No. The quantitative half runs entirely on GA4 explorations plus whatever your CRM or store admin already reports. Most recording and heatmap tools have a free tier that comfortably covers a single funnel for two weeks, and a one-question on-page survey can be a plain form. Buy tools after the audit has told you which step you will be watching for the next year, not before.
A UX audit evaluates the experience against usability and accessibility principles and produces recommendations. A conversion audit starts from the money: it finds where measurable value leaks out of a specific funnel, explains it with evidence, and produces a ranked, testable backlog. There is real overlap in phase 2, but the deliverables differ — one ends in principles, the other in a prioritized sequence with owners.
Fix the broken things immediately and do not test them: validation errors, dead links, a call to action nobody can see on a phone. Test the things where reasonable people would disagree about the outcome. That split is exactly what phase 4's fixes-versus-tests division is for, and getting it wrong in either direction — testing bugs, or shipping opinions untested — wastes the same traffic.
Ten to fifteen ranked items is the useful range for a single audit. Fewer suggests the qualitative phase was rushed; many more means the list is an inventory rather than a plan, and the items at the bottom will never be reached before the next audit makes them stale anyway.
Then that is the finding, and it belongs at the top of the document. A conversion rate is a ratio, and changing the numerator's composition moves it more than most page changes ever will. Say so plainly, show the segmented data that demonstrates it, and do not spend the next quarter testing headlines against a mix problem.
Yes, and it is a good way to use it. The funnel shortens to entry, engagement, form start, form completion, and qualified outcome, but every phase applies unchanged. The one adjustment: on a single page you will hit the traffic ceiling on testing sooner, so phase 3's significance check tends to push more items into the ship-without-testing column.
Ready to see where your budget leaks?
Free 30-minute audit, written roadmap included. No contracts.
Get My Free Growth Audit