How an A/B test works
Half your visitors see A, half see B, and you measure which drives more of the action you care about. The split is random, so the only difference between the groups is the change you made.
The discipline is in deciding the metric before the test starts. A variant can lift clicks and still lose you revenue if the extra clicks come from people who were never going to buy.
| What you change | Typical effect | Feasible on B2B traffic |
|---|---|---|
| Button colour | Under 1% | No, never reaches significance |
| Headline wording | 1% to 3% | Rarely |
| Form length | 5% to 20% | Yes |
| The offer itself | 10% to 50% | Yes |
| Page structure and proof placement | 5% to 25% | Yes |
Why most tests prove nothing
Teams test assumptions instead of hypotheses. "Users prefer a cleaner page" is an assumption. "Removing the secondary nav from checkout will cut abandonment" is a hypothesis you can disprove.
Sample size is the other trap. Detecting a lift from 2% to 2.4% takes roughly 30,000 visitors per variant. Below that you are reading noise.
What to test on a B2B site
B2B traffic is too thin for button-colour tests to reach significance. Test the things with large effects:
- Form length, and which fields are genuinely required
- The offer itself, such as a demo versus an audit versus a guide
- Page structure, such as proof above the fold rather than below it
- Headline framing, where the change is a different claim and not a different adjective
Why it matters for B2B marketing teams
Most B2B sites do not have the traffic for testing, and pretending otherwise wastes quarters. Be honest about your volume, then choose the method that matches it. A confident decision from observed friction beats an underpowered test read as significant.

