A/B testing that respects your traffic

Most store experiments are too small to conclude anything, and most "winners" are noise that regresses next month. Testing well means testing less.

A/B testing has a dirty secret in e-commerce: below substantial traffic volumes, most tests cannot detect the effect sizes that page changes realistically produce. The honest workflow is to fix obvious problems directly, reserve experiments for genuinely uncertain decisions on pages with enough volume, and size them so a conclusion is actually possible. The agents enforce exactly that discipline.

What deserves a test — and what does not

A broken mobile layout does not need an experiment; it needs a fix. Testing is for decisions where reasonable people disagree and the data could settle it.

  • Fix-directly: clear defects, speed problems, missing information, broken flows

  • Test: price presentation, offer framing, page structure choices with real trade-offs

  • Analyst checks in advance whether your traffic can detect the plausible effect size

  • Tests that cannot conclude are not started — that traffic is spent elsewhere

Hypotheses from the funnel, not the swipe file

Good tests come from your own drop-off data: a specific page, a specific segment, a specific hesitation. Analyst generates hypotheses from where the funnel actually leaks, which produces fewer, better experiments than borrowing test ideas from articles.

Honest conclusions, including "no difference"

A test that shows no effect is information — it means the decision does not matter and you should stop deliberating it. The agents report results with uncertainty intact, call peeking what it is, and record every result in shared memory so the same argument does not get re-litigated next quarter.

How it works

01

Screen the backlog

Obvious defects route to Builder as direct fixes; only genuine uncertainties become test candidates.

02

Size before starting

Analyst calculates whether your traffic can detect the plausible effect — underpowered tests are not run.

03

Run, conclude, record

Results are reported honestly, including null results, and stored so decisions stay settled.

What you get

  • Traffic spent on tests that can actually reach conclusions

  • Obvious problems fixed immediately instead of queued behind experiments

  • Hypotheses grounded in your own funnel data

  • A recorded history of results, including the null ones

Frequently asked questions

How much traffic do I need to A/B test?

It depends on the effect size you are trying to detect — small effects need very large samples. The agent calculates this per test before starting, which is precisely the step most teams skip.

What should a low-traffic store do instead?

Fix defects directly, make bigger and bolder changes where the effect would be large enough to see, and rely on before/after measurement with honest caveats. Pretending to run experiments is the worst option.

Why do winning tests stop winning?

Usually because the original result was noise, or was peeked at and stopped early. Proper sizing and fixed stopping rules prevent most of it — which is what the agent enforces.

Can it test email as well as pages?

Yes, and email often has better statistical conditions — larger samples, cleaner attribution. The same sizing discipline applies through the Retention agent.

Put the team to work on your store

Connect your storefront, analytics, and email stack, set a goal, and let the agents run the work end to end. Start on the free plan with your own model key.