Website ab testing is the single most reliable way to stop guessing which version of a page or element actually earns more revenue — and it belongs in every serious CRO programme in South Africa. Run two variants simultaneously, split your traffic, measure the outcome, and keep what wins. That is the whole method, and everything else is execution detail.
The execution detail is where most local businesses fall short. Sample sizes are underestimated, tests are called too early, and the SA context — EFT preference over card, mobile-first traffic from townships and smaller metros, courier trust gaps outside Johannesburg and Cape Town — is ignored entirely. This post fixes that.
Quick Answer
Website ab testing works by splitting live traffic between two page variants and measuring which converts better at statistical significance. In South Africa, tests need larger samples than most tools suggest because mobile traffic is dominant, EFT and instant-EFT behaviour differs from card, and courier trust varies sharply by region. Run tests for a minimum of two full business weeks before drawing conclusions.
Not sure whether your traffic volumes are large enough to run meaningful tests?
Get a Free Traffic & Sample Size AssessmentWhat Website AB Testing Actually Is
Website ab testing is a controlled experiment that shows different visitors different versions of the same page or element, then measures which version produces a statistically significant improvement in a defined goal. It is not a poll, not a preference vote, and not a heatmap — it is a live revenue experiment.
The classic form is an A/B test: one control, one variant, one primary metric. Multivariate testing extends this to multiple elements simultaneously, but it demands proportionally more traffic — a constraint that rules it out for most South African SME sites below around 30,000 monthly sessions.
The metric you choose matters as much as the design change. Conversion rate is the obvious choice, but revenue per visitor is often more meaningful. A variant that converts 12% more visitors at a lower average order value can produce less revenue than the control. Define your success metric before the test starts, not after you see the numbers.
Why Most SA Businesses Get Website AB Testing Wrong
The most common failure is stopping a test the moment one variant looks better in the dashboard. This is called peeking, and it reliably produces false positives. A variant that appears to be winning at day three with 200 sessions per arm has almost certainly not reached the statistical power required to trust the result.
South African ecommerce sites face a compounding problem here. If you understand what conversion rate optimisation requires, you know that low baseline conversion rates — common on local sites where EFT introduces a payment delay — mean you need more conversions, not just more sessions, to reach significance. A site converting at 1.2% needs materially more traffic to detect a realistic uplift than a site converting at 3.5%.
A second failure is testing the wrong thing. Button colour is the most-tested element on the internet and among the least impactful. The Baymard Institute's research across large ecommerce sites found that the average documented cart abandonment rate is 70.22%, with extra costs, slow delivery, and checkout complexity driving the majority of avoidable drop-off. Those are the areas worth testing — shipping cost presentation, checkout step count, payment method prominence.
In the South African context, payment method order is a particularly high-value hypothesis. A checkout that leads with card may be suppressing conversions among users who prefer Ozow or PayFast's instant EFT. Reordering the payment options, or surfacing a "Pay by EFT" option earlier, is a meaningful test that button colour never is.
Key Insight
Testing payment method prominence — EFT, Ozow, Peach Payments instant EFT versus card — is a higher-leverage hypothesis for most South African ecommerce sites than any visual design change. Start with checkout friction, not aesthetics.
Choosing the Right Website AB Testing Tool
The right tool for split testing depends on your platform, your traffic volume, and how much engineering support you can commit. There is no universal answer, but the decision tree is short.
For WordPress and WooCommerce sites, Google Optimize was the dominant free option until Google deprecated it. The current free-tier choice is either VWO's limited free plan or Optimizely's starter tier, though both push hard toward paid plans once you need statistical significance reporting. Hotjar's A/B testing module integrates with their heatmap data, which is useful for understanding why a variant wins, not just whether it does.
For Shopify stores — and Shopify is the most common platform among the South African ecommerce businesses we work with — the native Shopify Markets and theme editor do not run A/B tests natively.
You need a third-party app such as Neat A/B Testing or Convert.com, or you instrument the test manually via Google Tag Manager. The GTM route gives you the most control and costs nothing in tool fees, but it requires someone who can write a clean experiment script without introducing flicker.
Klaviyo, which many SA retailers use for email, runs its own A/B tests on subject lines, send times, and email content. These are worth running in parallel with on-site tests — email-driven traffic behaves differently from organic or paid traffic, and a variant that wins for one segment may not win for another.
| Tool | Best For | Pricing Model | SA Consideration |
|---|---|---|---|
| VWO (free tier) | WordPress / WooCommerce up to 10k sessions/month | Free to ~R3,500/mo | Rand billing not available; USD costs fluctuate |
| Convert.com | Mid-size Shopify or custom sites | From ~R2,800/mo | GDPR/POPIA-friendly consent mode |
| GTM + GA4 custom experiments | Any platform with dev access | Free (dev time only) | No flicker if implemented correctly; most flexible |
| Hotjar A/B | Teams already using Hotjar heatmaps | Bundled with Hotjar plans | Combines behavioural and split-test data |
| Neat A/B Testing | Shopify stores | From ~R250/mo (app store) | Native Shopify integration; lowest setup friction |
| Klaviyo A/B | Email sequences and flows | Bundled with Klaviyo | Useful for EFT abandonment email sequences |
Running WooCommerce or Shopify and unsure which testing setup will give you clean data?
Get a Free Platform-Specific Testing RecommendationSample Size Calculations for Website AB Testing in South Africa
Sample size is where most split testing programmes fail silently. Running a test to "statistical significance" means nothing if you defined significance at 80% power with a 5% minimum detectable effect on a site seeing 400 sessions a week — the test will run for months before it is valid, and most teams call it early.
The standard inputs for a sample size calculator are: baseline conversion rate, minimum detectable effect (MDE), statistical significance threshold (usually 95%), and statistical power (usually 80%). In South Africa, the baseline conversion rate on ecommerce sites is typically lower than global benchmarks, partly because EFT payments add a delay between intent and confirmed order, and partly because mobile data costs still influence session depth in some segments.
A practical rule: if your site converts at around 1.5% and you want to detect a 15% relative improvement (taking you to about 1.73%), you need roughly 15,000 sessions per variant before the result is trustworthy at 95% confidence.
At 5,000 sessions per week, that is a six-week test. At 1,000 sessions per week, it is a 30-week test — by which point your site has likely changed for other reasons and the experiment is contaminated.
The implication is direct: if your traffic is below roughly 8,000 monthly sessions, split testing will produce unreliable results for anything less than a very large effect. At that traffic level, you are better served by user research, session recording, and iterative design changes guided by qualitative data. Running experiments for their own sake on low-traffic sites produces confident-looking numbers that mean nothing.
| Scenario | Before | After |
|---|---|---|
| Baseline conversion rate (WooCommerce checkout) | 1.4% | 2.1% |
| Monthly sessions | 12,000 | 12,000 (unchanged) |
| Monthly orders | 168 | 252 |
| Average order value (R) | R890 | R890 (unchanged) |
| Monthly revenue | R149,520 | R224,280 |
| Revenue uplift per month | — | R74,760 (+50%) |
| Test element changed | Generic "Checkout" CTA, card-first payment order | Specific CTA copy, EFT/Ozow surfaced first |
| Test duration | — | 4 weeks at 95% confidence |
The figures above illustrate the pattern we see, not a guaranteed outcome. Your baseline, traffic mix, and winning variant will differ.
The Six-Step Website AB Testing Process
A disciplined split testing process eliminates the most common failure modes before the experiment launches. Follow these steps in order — skipping any one of them produces contaminated data.
Step 1 — Define the hypothesis. Every test needs a specific, falsifiable hypothesis: "Surfacing Ozow as the first payment option will increase checkout completion rate because SA mobile users prefer instant EFT over card entry." Not "let's try a different button."
Step 2 — Calculate your required sample size before touching any code. Use a free calculator (Evan Miller's is widely used) with your actual baseline conversion rate and a realistic MDE. If the required runtime exceeds eight weeks, reconsider the test or consolidate it with a larger redesign.
Step 3 — Implement without flicker. Flicker — the visible flash of the control before the variant loads — is a known contaminant. Implement variants server-side where possible, or use a synchronous GTM tag that fires before page render. On Shopify, this means using a native app or Liquid-level changes rather than JavaScript injection.
Step 4 — Run for at least two full business weeks, regardless of what the early numbers show. South African shopping behaviour varies meaningfully between weekdays and weekends, and between salary-week traffic and mid-month traffic. A test that only captures one pay cycle has a biased sample.
Step 5 — Analyse correctly. Read the confidence interval, not just the point estimate. A result showing "Variant B converted 18% better" means nothing without the confidence interval. If that interval spans from +2% to +34%, the result is noisy. If it spans +14% to +22%, you have a clean win.
Step 6 — Document and ship. Every test result — win, loss, or no-result — goes into a test log with the hypothesis, dates, traffic split, result, and confidence level. Losing tests are often as valuable as winning ones. A checkout simplification that produces no uplift tells you friction was not the conversion barrier; now you know to look elsewhere.
Key Insight
A test log is not optional housekeeping — it is the compounding asset of a CRO programme. Each documented result narrows the hypothesis space for the next experiment, so the tenth test is faster to design and more likely to win than the first. Teams that skip documentation repeat the same experiments and learn nothing twice.
GPM's Approach to Website AB Testing
Most agencies treat split testing as a bolt-on service they run when a client asks for it. We built our CRO practice the other way: testing is the mechanism by which every design and copy decision earns its place, not a periodic audit.
When we work with a South African retailer or B2B lead-gen site through our conversion rate optimisation service, the first four weeks are spent establishing baselines — real conversion rates by traffic source, by device, by payment method, by region. Gauteng desktop traffic converting at 3.1% and KZN mobile traffic converting at 0.9% are two different problems. Running a single test across the blended average solves neither.
We prioritise hypotheses by expected value: (probability of winning) × (estimated uplift) × (revenue at stake). Checkout payment order, shipping cost framing, and form field reduction consistently rank higher than visual design changes in this framework — which is why our test roadmaps look different from most agency proposals.
We also account for the South African commercial calendar. Running experiments through month-end salary week, through Black Friday, or through the school-holiday period in December produces skewed results. We schedule test windows around these periods or flag the contamination explicitly in reporting, so clients always know whether a result is clean or qualified.
Key Insight
Segmenting by region before building a test roadmap is not a nice-to-have — it is the difference between solving a real problem and averaging it away. A site where Gauteng desktop converts at three times the rate of KZN mobile needs two different hypotheses, not one blended experiment that leaves both audiences underserved.
Who This Is NOT For
Sites below 8,000 monthly sessions. Website ab testing requires sufficient traffic to reach statistical significance within a reasonable timeframe. Below roughly 8,000 sessions per month, most tests will run for so long that external changes — seasonal shifts, a Takealot promotion, a Google algorithm update — will contaminate the result before you can trust it. Use session recording and user interviews instead, and build the traffic base first before committing to a split testing programme.
Businesses without a defined conversion goal. If your team cannot agree on a single primary success metric before the test starts, split testing will produce arguments, not decisions. "Conversion rate" and "revenue per visitor" often point in different directions.
Lock the metric before you touch the page, and make sure the person who will act on the result has signed off on that metric in advance — otherwise the result will be relitigated the moment it delivers an inconvenient answer.
Teams expecting results within a week. One week of data is almost never sufficient for a valid result, even on high-traffic sites. If your business cadence requires decisions faster than a two-week test window, you need a different decision-making framework — qualitative research, expert review, or benchmark-led design — not a split test. Rushing a website ab testing experiment to fit a reporting cycle is a reliable way to implement a losing change with confidence.
Brands running simultaneous major changes. If your development team is rolling out a new checkout flow, a new payment gateway, or a significant navigation change at the same time as a live experiment, the test is contaminated. Split testing requires a stable control environment. Pause experiments around major releases and resume once the new baseline has settled for at least two weeks — only then do the results mean anything you can act on.
Ready to build a testing roadmap that accounts for your actual traffic and South African payment behaviour?
Get a Free CRO Testing RoadmapWebsite AB Testing FAQ
What is website ab testing and how does it differ from multivariate testing?
It is a controlled experiment that splits live traffic between a control and a single variant, measuring which produces a better outcome on a defined metric. Multivariate testing tests multiple elements simultaneously and requires a multiple of the traffic that a standard A/B test needs. For most South African sites, a sequential programme of A/B tests produces faster, cleaner learnings than a multivariate approach.
How many sessions do I need before starting a website ab testing programme?
A practical minimum is around 8,000 monthly sessions, which allows a typical test to reach significance within four to six weeks at a realistic minimum detectable effect. Below that threshold, tests take so long that external factors contaminate the data. If you are below this level, focus on qualitative research and expert heuristic reviews first to build the traffic base.
Which elements should I prioritise for testing on a South African ecommerce site?
Checkout friction is the highest-priority area — payment method order, shipping cost presentation, and form field count consistently produce larger uplifts than design changes. For local context specifically, the prominence of EFT and instant-EFT options (Ozow, PayFast) relative to card is a high-value hypothesis. Baymard's research shows that extra costs and a complicated checkout are among the top drivers of cart abandonment globally, and that pattern holds locally.
How long should a website ab testing experiment run?
A minimum of two full business weeks, regardless of early results. South African shopping behaviour varies between salary-week and mid-month periods, between weekdays and weekends, and by traffic source. Two weeks captures at least one full commercial cycle. Call the test when you have reached your pre-calculated sample size AND two weeks have elapsed — whichever comes later.
Does POPIA affect how I implement website ab testing?
POPIA's relevance to split testing is primarily at the consent and data storage layer. If your testing tool records session data or links test exposure to an identified user profile, that processing likely requires a lawful basis under POPIA — typically legitimate interest or explicit consent, depending on how the data is used downstream.
The Information Regulator has not issued specific guidance on A/B testing tools to date, but practices commonly interpret POPIA as requiring that test data be treated with the same care as any other behavioural personal information.
What statistical significance threshold should I use for website ab testing?
95% confidence is the standard for decisions that will be permanently implemented. For exploratory tests where you are simply deprioritising a losing variant rather than making an irreversible change, 90% confidence is defensible. Never call a test at below 90% — the false positive rate at 85% confidence is high enough that you will implement losing changes more often than winning ones, which is worse than not testing at all.
Want a testing roadmap built around your actual SA traffic and conversion data?
We'll audit your current conversion funnel, calculate realistic sample sizes for your traffic volumes, and deliver a prioritised action plan for your first three website ab testing experiments. No obligation — we'll get back to you within 24 hours.
Get Your Free CRO Testing Roadmap

