Help CenterBroadcasts & A/B testing

A/B testing broadcasts

Last updated July 21, 2026

In short. Test two subjects or two templates on part of the list, measure for a window, and the winner goes to the rest automatically. No clear winner means variant A ships, never a coin flip.

A broadcast can run as a holdout test: part of the list gets the variants, SEMAOS measures for a window you choose, and the rest of the list gets whichever variant won. Turn it on in the broadcast composer with "Test on part of the list, then send the winner to the rest." The whole point of the design is that deciding is automatic and conservative; you set it up and the machinery does the honest thing.

Your listTest groupVariant A + Variant BRemainderWaits for the winnerMeasurement windowOpens or clicks compared;floor: 100 deliveredWinner selected
A holdout test: a slice of the list tries the variants, the measurement window runs, and the rest of the list gets the winner.

What can I test?

Two kinds of variant. Subject tests send the same body under different subject lines and pick the winner by open rate. Content tests send different templates and pick by click rate, because when the bodies differ, opens can't tell the bodies apart; only what recipients did inside the email can. The composer switches metric automatically based on what you varied.

How is the winner decided?

At the end of the measurement window, a winner must clear two bars:

  1. The floor: at least 100 delivered emails per variant. Below that, rates are noise; a 3-of-7 open "lead" means nothing. The composer warns you upfront when your test-group size can't reach the floor.
  2. The margin: a real lead over the runner-up. Subject tests need an open-rate lead of at least 8 percentage points; content tests need a click-rate lead of at least 2 points, since click rates run much lower than open rates.

If either bar isn't cleared, variant A (your control) ships to the remainder. Deliberately not "whichever was ahead": shipping a statistically meaningless leader would launder noise into a decision. Writing your safer, known-good version as variant A is therefore part of setting up the test.

What are the knobs?

The test-group percentage (how much of the list tries the variants), the measurement window (minimum one hour; longer windows suit lists that open slowly), and approve-first mode. With approve-first on, nothing goes to the remainder until you review the results and click send; the winner's numbers wait for you on the results page. With it off, the remainder goes out automatically the moment the window closes.

What does the results page show?

Per-variant delivered, opens, and clicks, with the winner badged once decided. The remainder's sends are attributed to the broadcast as a whole rather than to a variant, so variant stats stay a clean read of the test cohort only; the winner badge tells you what the remainder received.

When is a test worth running?

When the remainder is big enough to benefit. If your whole list is 300 contacts, a compliant test barely leaves a remainder to optimize for, and you're better off sending your best guess to everyone and applying what you learn next time. The floor-and-margin math starts paying for itself in the low thousands of recipients.