Observation, hypothesis, primary metric, sample size, and the decision rule — all written before the test starts.
We think the pricing question is killing our demo form. Write an A/B test brief.
Test is powered: 2,900 sessions per arm over roughly 16 days at current traffic. Detecting anything smaller than a 20% relative lift would take two months, so the brief targets the larger effect and pre-registers the three-way decision rule.
An A/B test template earns its keep by settling the argument in advance. This one converts a conversion problem into a testable hypothesis, computes the sample size your traffic can actually support, and records the decision rule before launch — so an ambiguous result cannot be reinterpreted after the fact.
A pre-registration document for one experiment. It states the observation that prompted it, the hypothesis in if/then form, the change being made, the primary metric, the sample size and duration required, and the rule that decides what each possible outcome means.
In if/then form with a mechanism: if we move the pricing objection above the form, then form starts will rise, because visitors currently leave to hunt for pricing. Without the mechanism you cannot learn anything transferable from a win.
Because most tests on B2B traffic are underpowered, and an underpowered test does not produce a weak answer — it produces a random one. Computing the minimum detectable effect up front tells you whether the test is worth running at all.
A record of the hypothesis and the result, including the losses. Teams that document only winners rerun the same failed idea every eighteen months, and the mechanism behind a loss is frequently more useful than the win.
The agent reads baseline conversion rates from GA4 or your analytics to size the test, and pulls the observation from your funnel or session data where available. The brief exports to Notion or Linear as the pre-registration record, and deployment stays in your testing tool where QA and rollback live.
What you noticed and where — a drop-off, a session recording pattern, a support complaint.
It writes the if/then with a mechanism and names what would falsify it.
Sample size, duration, and minimum detectable effect from your real baseline rate.
Pre-register the rule, run the test in your tool, then record the outcome against the brief.
Decision rule is written before launch, which prevents post-hoc reinterpretation
Sample size comes from your real baseline, not a rule of thumb
Declines to schedule tests your traffic cannot power
Archives losses with their mechanism, so ideas are not rerun blindly
The change, the expected effect, and the mechanism — in if/then/because form. The mechanism is what makes a win reusable on other pages instead of a one-off result.
Until it reaches the pre-computed sample size, and at least one full business cycle so weekday and weekend behaviour are both represented. Stopping when the result looks good is the most common way to produce a false positive.
Two, unless your traffic can power more. Every additional arm splits the sample and extends the test, and most B2B pages do not have the volume for a three-way split.
That is a valid outcome and the decision rule should already cover it — usually keep the control and move on. Inconclusive tests on small effects are the normal case, not a failure of process.