JDKRUEGER&COAcademySign inDE
Module 1 of 8 · A/B Testing Mastery

A/B Testing for Non-Statisticians

⏱ 30 min · After completing this module, you'll be able to tell the difference between a real A/B test and a before-and-after comparison, explain why only parallel, randomized testing rules out external confounding factors, and write a falsifiable hypothesis that clearly names the variable, the expected effect, and the primary metric.
← A/B Testing Mastery A/B Testing for Non-Statisticians 1 / 28
continuing in 5
Start

A/B Testing for Non-Statisticians

Making decisions that hold up in reality, not gut feeling.

From Assumption to Proven Decision
Transcript of this slide

Welcome to the first module of the A/B Testing Mastery track. We start with a simple truth: most shops make daily decisions about layout, copy, and offers, and back them up with opinions. In this module, you'll learn how to turn those decisions into controlled experiments that tell you what actually works. The goal isn't to become a statistician. It's to become a confident, informed partner in data-driven growth.

Learning objective

Learning Objective

You'll recognize what makes a genuine A/B test.

  • You'll understand why parallel testing under identical conditions is more reliable than before-and-after comparisons.
  • You'll write your own falsifiable test hypothesis.
1
Formulate a hypothesis
2
Split randomly
3
Measure in parallel
4
Decide with data
Transcript of this slide

After this module, you'll consciously distinguish between an educated guess and a valid experiment. You'll know why parallel groups are the gold standard, and you'll be able to write a hypothesis that's actually testable. These three skills form the foundation for everything that comes in the next modules.

Self-check

Activate Prior Knowledge

How do you currently tell the difference in your shop between a successful change and a lucky coincidence?

  • When was the last time a new variant seemed "obviously better", but nobody actually checked the numbers?
Key Points

Activate Prior Knowledge

  • 1 How do you currently tell the difference in your shop between a successful change and a lucky coincidence?
  • 2 When was the last time a new variant seemed "obviously better", but nobody actually checked the numbers?
Transcript of this slide

Before we get into the details, two quick questions to ask yourself. Think about the last major change you made in your shop. Was it validated with a parallel test, or did it just go live because everyone was convinced it was the right call? That question matters, because it shows how far along your business already is on the path to a culture of experimentation.

Concept

The Business Problem: Seventy Percent Abandonment

The Baymard Institute analyzed over one hundred thousand checkout processes and found an average cart abandonment rate of sixty-nine point eight percent.

  • Every second shop owner underestimates this number, because it stays invisible in day-to-day operations.
  • The result: thousands of visitors leave the page just before making a purchase.
Drop-off 70 (70%) Purchase 30 (30%)
Average Cart Abandonment Rate in E-Commerce
Transcript of this slide

Imagine that in your physical store, seventy out of every hundred customers put their items back at the register and walk out. That's exactly what's happening online, except nobody tells you to your face. This number is the business case for why A/B testing is so valuable. Every abandoned cart doesn't just cost you revenue. It also costs you the money you already spent to get that visitor there.

Concept

Why More Traffic Isn't the Answer

With a conversion rate of three percent and a seventy percent abandonment rate, you need thirty clicks to generate one purchase.

  • Spending more on paid traffic at the same conversion rate only eats into your margins and drives up cost pressure.
  • The real opportunity is in your existing traffic, not in buying more.
10,000 visitors → 3,000 carts → 900 purchases
Transcript of this slide

A lot of shops try to buy their way out of abandonment. But if seventy percent are dropping off, you're paying for thirty clicks and getting one purchase. More traffic at the same conversion rate means you're pushing more water through the same pipe. The real opportunity sits in the visitors you already have, and that's exactly where A/B testing comes in.

Concept

What a Real A/B Test Does

Your existing traffic is split randomly into two equal groups.

  • Group A sees the current version; Group B sees the changed version.
  • Both groups browse at the same time, under the same conditions.
1
Random split
2
Variant A - Control
3
Variant B - Change
4
Simultaneous measurement
Transcript of this slide

A real test is a controlled experiment. Only one variable changes: the element you're testing. Everything else stays identical, including seasonality, day of the week, and active ad campaigns. That parallelism is the core of the method. It's what ensures that any difference in results can actually be traced back to the change you made.

Concept

What an A/B Test Is Not

It's not a design competition where the prettier variant wins.

  • It's not a platform for opinions like "nobody clicks green buttons in my experience".
  • It's also not a method for spinning up a new idea every week.
Data vs. Opinion
Transcript of this slide

Two persistent myths cost businesses money every year. First, that changing a button color will double revenue. Second, that we know our customers well enough to make data unnecessary. Both are dangerous. An A/B test isn't a beauty contest or a stage for gut instinct. It's a measuring instrument for real business impact.

Concept

Anatomy of an Experiment: The Hypothesis

A hypothesis states which change you're testing, why you're testing it, and what outcome you expect.

  • Weak: "We're testing a new button."
  • Strong: "If we make the CTA button in the cart larger, the conversion rate will increase by five percent, because the action becomes more prominent."
1
Observation
2
Assumption
3
Expected change
4
Metric
Transcript of this slide

Without a clear hypothesis, you're testing blind. A good hypothesis names the variable, the expected direction and magnitude of the effect, and the metric you'll use to read the result. You don't need to know whether the hypothesis is correct. In fact, that's exactly what you're running the test to find out. But you do need to be able to state upfront what you're measuring.

Concept

Control, Variant, and Randomization

The control is the current version. It's your status quo, measured under existing conditions.

  • The variant contains exactly one targeted change.
  • Randomization ensures that users don't choose which version they see. They're assigned at random.
Control vs. Variant
Transcript of this slide

Randomization is the most important step. If visitors can pick which version they see, you're comparing apples to oranges. Only random assignment makes the groups comparable. Imagine a drug trial where healthy participants get to decide whether they take the new medication. That's exactly what a test without randomization looks like.

Concept

Metric and Traffic Split

The primary metric is the single number that determines success or failure. That's usually conversion rate or revenue per visitor.

  • The traffic split determines what share of visitors sees each variant, typically fifty-fifty.
  • An uneven split either extends the test's runtime or weakens the statistical validity.
14 28 42 56 50 Variant A 50 Variant B
Typical Fifty-Fifty Traffic Split
Transcript of this slide

Choose your primary metric before the test starts. Switching to the metric that looks best afterward is cheating yourself. That's called cherry-picking. A fifty-fifty split is the standard because it gets you to a valid decision fastest. Any deviation from that needs a solid justification.

Example

Why Before-After Comparisons Lie: A Concrete Scenario

Monday: rain, three thousand visitors, conversion rate two point one percent.

  • Tuesday: sunshine, three thousand visitors, new header goes live, conversion rate two point four percent.
  • Conclusion: The header drives fifteen percent more revenue.
  • More likely: The weather and the day of the week skewed the result.
1 2 2 3 2.1 Monday 2.4 Tuesday
Before-After with an External Confounding Factor
Transcript of this slide

A classic example: the new header goes live on Tuesday, and the sun comes out at the same time. Suddenly more people buy. The header gets the credit, even though the weather shifted the mood. This is exactly why before-after comparisons are generally not a valid basis for decisions. Too many things change between the two points in time.

Concept

Campaigns, Holidays, and Promotions

On Black Friday, conversion rates rise almost everywhere, regardless of any changes you make to your shop.

  • A retargeting campaign can reach the control group differently than it reaches the variant group.
  • Anyone measuring before-after is ignoring these confounding factors instead of controlling for them.
Confounding Factors Hit Time-Based Comparisons Unevenly
Transcript of this slide

External factors aren't inherently bad, but they're powerful. Holidays, weather, newsletters, and ad spend all change how your visitors behave. A before-after comparison can't separate those influences. A simultaneously running A/B test can, because both variants experience the same external conditions at the same time. That's the decisive difference.

Concept

Parallel Testing Eliminates Confounding Factors

In an A/B test, both variants run at the same time.

  • Weather, day of the week, campaigns, and seasonal effects hit both groups equally.
  • The only systematic difference that remains is the change being tested.
1
Same time period
2
Same visitors
3
Same external factors
4
One single difference
Transcript of this slide

Because both groups are tested at the same time under the same external conditions, we can rule out weather, promotions, and day-of-week effects. That's the core advantage of parallel testing. It's why A/B tests are considered the gold standard in both research and practice. Not because they're complicated, but because they make a fair comparison.

Interim check

Quick Check: The Three Pillars

Real hypothesis: What changes, why, and which metric are you measuring?

  • Random assignment: The groups must be comparable.
  • Parallel measurement: That's the only way to neutralize external confounding factors.
1
Hypothesis
2
Randomization
3
Parallelism
Transcript of this slide

A quick check before we go deeper. The three pillars of a valid A/B test are: a clear hypothesis, random assignment of visitors, and parallel measurement over the same time period. If any one of those pillars is missing, the entire result is shaky. Keep this triangle in mind and you'll be able to evaluate any experimental setup quickly.

Concept

Primary Metric vs. Guardrail Metrics

The primary metric determines whether the variant wins, for example revenue per visitor.

  • Guardrail metrics protect against side effects: return rate, support requests, average order value.
  • A variant can boost conversion while simultaneously lowering order value.
Keeping primary metrics and guardrails in view
Transcript of this slide

Say your new discount button lifts conversion by eight percent, but average order value drops by twenty percent. Without a guardrail metric, you'd be celebrating a loss. That's why every test needs protective metrics alongside the primary metric. They prevent an apparent winner from quietly undermining your business model.

Example

Business metrics that actually matter

Conversion rate: the share of visitors who make a purchase.

  • Revenue per visitor: combines conversion rate and order value.
  • Customer lifetime value: shows long-term impact, especially for subscription models.
  • Example: With 50,000 visitors per month and an average order value of €50, a €0.50 increase in revenue per visitor already adds €25,000 in monthly revenue.
4 7 11 14 3.2 Conversionrate 4.8 Revenue pervisitor 12.5 Lifetimevalue
Example business metrics
Transcript of this slide

Click-through rates are a means to an end. Business-relevant metrics show whether your test ultimately drives more revenue or more profit, and that's what leadership actually cares about. So when you're presenting results, don't talk about click rates. Talk about revenue per visitor or lifetime value. A concrete example makes this tangible: with 50,000 visitors per month and an average order value of €50, a rise in revenue per visitor of just €0.50 already means €25,000 in additional monthly revenue. That wins any budget conversation.

Concept

What makes a hypothesis solid

Falsifiable: the opposite must be theoretically possible.

  • One variable: change only one element per variant.
  • Expected effect: state both the direction and the magnitude of the anticipated change.
1
Falsifiable
2
One variable
3
Expected effect
4
Measurable metric
Transcript of this slide

A hypothesis like "more trust increases revenue" isn't falsifiable. How would you measure trust directly? Better: "Adding a trust badge in the checkout increases conversion rate by four percent." The opposite is conceivable, the variable is clear, and the metric is measurable. Hypotheses like that can be tested, and they can be disproved.

Common misconception

Common mistakes that invalidate tests

Peeking: checking results daily and stopping the test the moment something looks significant.

  • Stopping too early: variants fluctuate randomly until enough data has accumulated.
  • Too many variants: with ten variants, one almost always wins by chance.
Avoiding peeking, early stopping, and variant inflation
Transcript of this slide

Peeking is the most expensive mistake. Check often enough, and you'll inevitably find an apparent winner that disappears in the next test. Plan your runtime in advance. The same logic applies to too many variants: the more variants you test, the higher the probability that one simply looks better by chance. Less is usually more.

Exercise

Your exercise: write a falsifiable hypothesis

Pick a page or an element in your shop.

  • Frame it as: if we change X, then Y will increase or decrease by Z, measured by metric M.
  • Check: is only one variable changing? Is the opposite outcome conceivable?
1
Choose an element
2
Write if-then statements
3
Check falsifiability
Transcript of this slide

Do this now, briefly. Every hunch you have in your head is a candidate for a test. The key question is: can you verify it with two parallel groups and rule out the opposite? Write your hypothesis down. That's already the hardest step done: you've turned an assumption into a testable statement.

Summary

Summary: the key takeaways

Seventy percent of all online shopping carts are abandoned without a purchase, and more traffic doesn't fix that.

  • A proper A/B test splits traffic randomly and runs both variants in parallel under identical conditions.
  • Before-and-after comparisons get distorted by weather, day of the week, and campaign activity.
1
Identify the problem
2
Formulate a hypothesis
3
Test in parallel
4
Decide with data
Transcript of this slide

The core idea in three sentences: it's not about more traffic, it's about better decisions within the traffic you already have. With random splitting and parallel testing, you get answers you can actually trust. And with a clear hypothesis, you already know before the test starts exactly what success looks like.

Summary

What you're taking away

Define a clear primary metric and appropriate guardrails before the test begins.

  • Write hypotheses so the opposite outcome is possible and only one variable is changed.
  • Avoid peeking, stopping early, and running too many variants.
From assumption to valid decision
Transcript of this slide

These three habits separate amateur tests from professional experiments: clear metrics, precise hypotheses, and discipline throughout the test runtime. Apply just these three principles in your next strategy meeting and you'll already have an edge. In the next module, we'll look at what it actually takes for a test result to be statistically reliable.

Intermediate step

The JDKRUEGER&CO promise

We run more A/B tests across the DACH region than most agencies ever sell, and we back every recommendation with data.

Measurable. Scalable. Proven.
Transcript of this slide

At JDKRUEGER&CO, A/B testing isn't a buzzword. It's the foundation of every growth program we run. We test what we recommend, and we only recommend what the data supports. In the next module, we'll look at what it actually takes for a test to be statistically sound.

Quiz

Quiz

Test your knowledge.

A shop has a conversion rate of three percent and a cart abandonment rate of around seventy percent. What's the most logical strategic conclusion?

A shop launches a new homepage on March 1st. In April, the team notices that the conversion rate has gone up. At the same time, a major spring campaign was running. Why is the conclusion "the homepage is working" questionable?

Which hypothesis is best suited for an A/B test?

A variant increases the conversion rate by eight percent but drops average order value by twenty percent. What does this tell you?

A team checks the live test data every morning and stops the test as soon as significance is reached. What problem does this create?

Exercise

Exercise

Apply what you have learned right away.

  • 1
    Your first falsifiable hypothesis
    mini-audit · approx. 20 min
    Pick one element in your shop that you want to change, for example a headline, a button, or a product page. Write your hypothesis using this format: If we change X, then Y will increase or decrease by Z percent, as measured by metric M. Check that only one variable is being changed and that the opposite outcome is, in principle, possible.
  • 2
    Reviewing parallelism in your business
    benchmark · approx. 15 min
    Look up the last three major changes made to your shop. For each one, decide whether it was rolled out as a straight live launch, a before-and-after comparison, or a parallel A/B test. Flag any decisions where you still have uncertainty today about whether the change actually made a difference.
Reflection

Reflection

A quick look back before you continue.

  • Which of the last three changes in your shop went live simply because everyone was convinced - and which ones should you have tested in parallel?
  • Write a hypothesis for your most important conversion lever, one that clearly names the variable, the expected direction, and the primary metric.
  • Where in your shop are you currently mistaking a weather or campaign effect for the effect of a real change?
Feedback

Feedback

Was this module helpful for your shop?

Sources

Sources & further reading

Here you will find links and materials to explore the topic in more depth. Take your time.

Overview & learning objective

This module is aimed at shop owners.

After completing this module, you'll be able to tell the difference between a real A/B test and a before-and-after comparison, explain why only parallel, randomized testing rules out external confounding factors, and write a falsifiable hypothesis that clearly names the variable, the expected effect, and the primary metric.

A/B Testing for Non-Statisticians