- Detecting a 20% lift on a 2% baseline needs roughly 4,000 conversions per variant, which rules out A/B testing for most small sites.
- On the last four accounts we reconciled, analytics under-reported enquiries by 20% to 60%.
- Broken basics, long forms, unlinked phone numbers, hidden errors, do not need a test before they are fixed.
- Twenty session recordings and five user tests find more at low traffic than any testing tool.
- Sequential comparison works if you change one thing, use equal whole-week windows, and write down what else changed.
Conversion rate optimisation is written about almost entirely by people with enormous traffic. Read it and you will conclude that CRO means running simultaneous multivariate tests. Then you open your own analytics: 400 sessions a week, eleven enquiries a month, and a test that would take fourteen months to reach significance.
This is what we do at that scale, which is the scale most businesses actually operate at.
Work out whether you can test at all
Before designing anything, do the arithmetic. To detect a 20% relative improvement on a 2% baseline conversion rate with reasonable confidence you need somewhere around 4,000 conversions per variant. At eleven conversions a month, that is not a test, it is a retirement plan.
Be honest about this early. A test that cannot conclude is worse than no test, because it produces a number that looks like evidence and is not.
As a rough threshold: under about 1,000 conversions a month, classic A/B testing answers only very large questions. Above it, test properly.
Instrument before you optimise
Everything downstream depends on the measurement being right, and in most accounts we inherit it is not.
The minimum: every form submission fires a named event with the form identified. Phone clicks and email clicks are tracked. The thank-you page is not the only conversion signal, because single-page forms often never reach one. Events fire once, not three times. Tags load after consent, and the consent state is recorded.
Then reconcile. For one month, compare the analytics conversion count against the actual number of enquiries in the inbox. On the last four accounts we did this for, analytics was under by between 20% and 60%. Fix that before optimising anything.
The most common single cause, by a distance, is a form plugin that redirects to a page the tag never sees. The second is a phone number that is a link on desktop and plain text on mobile.
For one month, count the enquiries that actually reached your inbox and compare with what analytics reports. If they differ by more than 10%, fix the tracking before changing anything on the site.
Fix the obviously broken first
There is a stage before testing, and most sites are still in it. These changes do not need a test, because there is no plausible argument for the current state.
Forms asking for eleven fields when four would do. Phone numbers that are not links on mobile. A thank-you page that says nothing about what happens next. Pricing pages with no numbers on them. Errors that appear only after submission, in red, above a fold you have already scrolled past. Load times over four seconds on mobile.
On a client site last year we removed four form fields and made the phone number a link. Enquiries rose 38% over the following two months. That is not a test result, there is no control, but nobody sensible would have argued for the previous state.
Use qualitative evidence, seriously
At low traffic, the highest-yield tool is watching people.
Session recordings. Watch twenty sessions that reached the enquiry page and did not convert. Twenty is enough to spot a pattern, and you will usually see the same hesitation three or four times.
Five-user tests. Give five people a task on your site and watch them do it. Old advice, still the fastest way to find serious usability problems.
Ask the people who did convert. One question on the thank-you page: what nearly stopped you getting in touch? The answers are startlingly direct and they cost nothing to collect.
Read the sales inbox. Every question a prospect asks by email is a question your page failed to answer. We keep a running list of those on client projects and it becomes the next quarter's page edits.
Sequential testing, done carefully
When you must compare and cannot split, compare over time, with discipline.
Run the current version for a full cycle, usually four weeks, and record the conversion rate. Change one thing. Run four more weeks. Compare, then check the obvious confounds: a seasonal change, a campaign that started, a competitor's promotion, a Google update.
This is a weaker method and we say so to clients. It is still much better than shipping changes and never looking. The rules that keep it honest: one change at a time, equal-length windows, whole weeks, and a written note of everything else that changed in the same period. That note is the part people skip, and it is the reason so many before-and-after claims fall apart under questioning.
The changes that pay most often
From our own client work, in rough order of how often they have helped: reducing form fields; adding specific proof next to the form, a named client, a number, a photograph of real work; stating price ranges instead of "contact for a quote"; making the primary action visible without scrolling on mobile; and rewriting the headline to match the ad or the search term that brought the visitor.
The last one is the most underrated. Someone who arrived searching shopify migration cost and lands on a page headed Digital Transformation Partners has already decided to leave. This is also where conversion work and paid search meet: the same mismatch shows up as a poor quality score, and our small-budget PPC guide covers the ad side of it.
We run conversion programmes for low-traffic sites: instrumentation, recordings, user tests and sequential changes with an honest read on what moved.
What we do not do at this scale
Multivariate testing. Personalisation engines. Heatmap-driven redesigns where the heatmap is decoration rather than evidence. AI-generated variant copy tested against itself, which optimises for a metric nobody checked the meaning of. And calling a test at 90% confidence because the client is impatient: that is a coin flip with a certificate.
When to bring us in
Two situations are worth paying for. The first is the reconciliation, because tracking that under-reports by a third makes every other decision wrong and it is fiddly to fix properly. The second is when the obvious problems are already fixed and enquiries are still flat, which usually means the problem is the offer or the page's argument rather than its buttons. Both are what our conversion rate optimisation work consists of, and if the site is slow as well, start with Core Web Vitals instead.
Questions
How much traffic do I need for A/B testing?
Roughly 1,000 conversions a month before classic split testing answers ordinary questions in a sensible timeframe. Below that, use sequential comparisons and qualitative research instead.
Are session recordings worth the privacy trade-off?
They can be, if you mask input fields, exclude sensitive pages, disclose it in your privacy policy and honour consent. Recording form contents without consent is not defensible in any jurisdiction we work in.
What is a good conversion rate?
Comparisons across businesses are close to useless, because traffic quality varies too much. The number that matters is your own: this quarter against last, with the traffic mix noted.
Should I test my headline or my button colour?
Headline, every time. Button colour tests are the reason CRO has a reputation for triviality. The message, the offer and the form length are where the differences live.
Can I trust a test that reached 90% confidence?
Treat it as a hint, not a result. At 90% you are wrong roughly one time in ten, and low-traffic tests are usually stopped early, which makes it worse. Either wait, or make the change on judgement and say that is what you did.

