Analytics & CRO
A/B testing that produces real answers
An A/B test shows two versions of a page to similar visitors and measures which converts better, replacing opinion with evidence. We run testing programs properly: research-backed hypotheses, correct statistics and honest reporting of losers as well as winners.
Who it is for: Sites with enough traffic to test and teams who want decisions settled by data, or who ran tests before and never trusted the results.
Everything in this service
- Test roadmap built from research, not brainstorms
- Hypotheses with predicted impact stated before launch
- Test builds, QA across devices and correct traffic splits
- Statistical analysis with honest confidence reporting
- Multivariate testing where traffic supports it
- A searchable archive of every result, including losers
What to expect
- Changes proven to help before rolling out to everyone
- A growing archive of what works for your audience
- An end to decisions by loudest opinion
How a/b testing actually works
What a test can tell you, and what it cannot
An A/B test answers one narrow question well: for visitors like the ones in the test, over the period it ran, did version B produce more of the measured outcome than version A? That is genuinely valuable and it is much narrower than how results usually get reported.
It does not tell you why, which is why tests should start from research rather than replacing it. It does not tell you whether the result holds for a different audience, a different season or a different traffic mix. And it does not tell you about effects beyond the measured metric. A change that lifts form submissions while lowering the quality of those submissions is scored as a win by the test and as a loss by your sales team.
So we define the success metric as close to revenue as the data allows, and we watch a guardrail metric alongside it. Measuring enquiries alone invites optimising for enquiries nobody wants.
Sample size, significance and the peeking problem
The single most common cause of unreliable testing is stopping early. Conversion rates fluctuate heavily at low volumes, and a test watched daily will, at some point, show a large apparent difference purely by chance. Stopping at that moment produces a confident number and a false conclusion.
The discipline is to calculate the required sample before launch, from your current rate and the smallest improvement that would be worth acting on, and then to run to that number regardless of what the dashboard shows in week one. If a test needs longer than you are willing to wait, that is information: the effect you are looking for is too small to detect at your volume, and the sensible move is to test something bolder.
Two further rules. Run in whole weeks, because weekday and weekend behaviour differ and a partial week skews the sample. And decide the success metric before launch, because a test with five metrics will show significance on one of them by chance alone, and choosing the winner afterwards is not analysis.
- Sample size calculated before launch, from a real minimum detectable effect
- One primary metric decided in advance, with a guardrail metric watched
- Whole business cycles, so weekday and weekend mix is represented
- No stopping the moment the graph looks good
Why previous tests showed wins that later vanished
This is the most common thing clients tell us about their earlier testing, and there are usually three causes.
The first is early stopping, which manufactures wins that were never real. The second is testing during an unrepresentative period: a sale, a campaign launch, a seasonal peak, or a week when a large referral source spiked. The audience during that window was not your normal audience, so the result does not generalise. The third is running many tests without correcting for the fact that some will look significant by chance.
There is also a genuine effect that gets mistaken for a fault. Some changes produce a short-term lift simply because they are new and returning visitors notice them, and that lift decays. Running long enough for novelty to settle, and re-checking a headline win a quarter later, separates a real improvement from a temporary one. We would rather report that a win faded than let it sit in a case study.
What to test, in what order
Small cosmetic tests are where most programmes stall. Button colours and microcopy tweaks produce effects too small to detect at typical traffic volumes, so the test either runs for months or gets stopped early, which brings back the first problem.
Bigger swings are more testable, because larger effects need smaller samples. Testing a genuinely different offer framing, a substantially shorter form, a restructured page that leads with a different message, or the removal of a step, all produce effects large enough to measure in reasonable time.
Order follows the same logic as the rest of the conversion programme: highest traffic multiplied by highest stake first. A test on a page with a hundred visitors a month is not a test, it is a guess with extra steps, and we will tell you when that is the situation rather than billing for it.
Multivariate, bandits, and when they earn their complexity
Multivariate testing varies several elements at once to find the best combination. It needs far more traffic than a straight A/B test, because every combination requires its own sample. On most sites it is the wrong tool, and the honest recommendation is a sequence of A/B tests instead.
Multi-armed bandit approaches shift traffic toward the better performer while the test runs, which reduces the cost of showing a losing version. That suits short-lived situations such as a campaign creative or a limited-time offer, where maximising results during the window matters more than a clean measurement afterwards.
For durable decisions about a page you will live with for years, a properly run A/B test with a fixed sample and a clear stopping rule gives a cleaner answer, and a cleaner answer is what you are actually buying. We use the simpler tool unless traffic and the situation genuinely justify the complex one.
A clear path, step by step
- 01
Hypothesise
Each test starts from research: because we observed X, changing Y should improve Z.
- 02
Build and QA
Variants built and checked on every device and browser, so bugs never masquerade as results.
- 03
Run to significance
Tests run until the sample is sufficient. No peeking, no calling winners early.
- 04
Decide and archive
Winners implemented, losers documented, and every learning feeds the next test.
Why choose us for this
We report losing tests, because that is half the value
Statistics done properly: no early calls, no cherry-picking
Tests QA-checked so broken variants never poison results
Common questions
How much traffic do we need to A/B test?
As a rough guide, a few hundred conversions per month on the page being tested. Below that, tests take too long to conclude honestly, and research-led CRO with before-and-after measurement serves you better. We will tell you which side you are on.
Why did our previous tests show wins that disappeared?
Usually from stopping tests early, testing during unusual periods or chasing too many metrics. A result at proper significance, run over full business cycles, holds. That discipline is most of what you are paying for.
Explore related work
Want this for your business?
Book a free visibility call and I will tell you honestly whether I can help.
How this is delivered
One person leads every project. Where a job genuinely needs a specialist, I bring in people I have worked with before and manage them, so you get one point of contact and one invoice rather than three suppliers blaming each other.
- You talk to the person responsible for the work, not an account manager
- Specialists are briefed and managed by me, and their work is checked before it reaches you
- One contract, one invoice, one place to chase