Create a test
A/B Tests → Create New A/B Test. Four steps.1
Name it and pick the goal
The goal metric is what decides the winner:
Choosing Custom Metric makes you pick which one. Define it first under
Analytics → Custom metrics.
2
Configure the arms
- Clone: take one video and clone it into 2, 3 or 4 arms you then edit.
- Compare: pit existing, different videos against each other.
3
Pick the videos
One base video in clone mode; one video per arm in compare mode.
4
Review and launch
Check the split, then launch.
Embed the test, not the video
Once a test is live, the video’s Embed tab switches to a loader snippet pointed at the test. That snippet is what rotates viewers between arms.How the winner is decided
The statistical test
The statistical test
Rate metrics (conversion rate, completion rate, custom metrics) use a
two-proportion z-test. The denominator is unique plays, not raw plays, so a viewer
who resumes or restarts cannot inflate an arm.Watch time uses a two-sample t-test.
Why watch time sometimes says 'insufficient data'
Why watch time sometimes says 'insufficient data'
A t-test needs the real variance of each arm. When it is not available, TrackPlay
reports insufficient data and gives you the effect size only.It will not invent a p-value from an assumed variance. A fabricated p-value is worse
than no p-value, because you would act on it.
Testing 3 or 4 arms
Testing 3 or 4 arms
More arms means more comparisons at once, which makes a false positive more likely.
TrackPlay applies a Bonferroni correction: the significance threshold is divided by
the number of comparisons.
Significance alone is not enough
Significance alone is not enough
A winner must clear both bars: statistically significant, and a large enough
effect to be worth acting on. A 0.1% lift that is technically significant is not a
reason to rebuild your funnel.
Reading the results
A/B Tests → your test.- Traffic Flow shows where traffic actually landed, which is not always where you configured it to land.
- Statistical Analysis shows the p-value, the confidence interval, and the test
statistic (labelled
zortdepending on the metric). - Variant Performance compares the arms directly.
Declaring a winner
Click Declare Winner and pick an arm. All traffic moves to it. If the numbers are not ready, TrackPlay tells you why and makes you tick an explicit override before the button works.Overriding the guardrail is sometimes the right call: a launch ends, a deadline lands.
But you are choosing on a coin flip. Know that you are doing it.
Letting TrackPlay declare it for you
A test finishes itself. Ship the winner automatically is on for every new test, and you can switch it off on the test’s page at any time, even while it runs. It is the one setting you can change mid-run: the statistics themselves are locked once a test launches. Tests created before September 2026 keep whatever they were set to. With it on, TrackPlay checks the test once an hour and declares a winner only when every one of these is true:1
The result is statistically significant
At the Bonferroni-corrected threshold described above.
2
The effect is large enough to act on
The same practical-significance bar a manual declaration has to clear.
3
Each arm has enough traffic
The minimum per-variant sample is met.
4
The test has run long enough
Your minimum duration, 7 days unless you change it. This exists so a test cannot be
called on the first favourable hourly check.
5
The observed p-value clears your threshold
Your decision threshold, 0.95 by default, which means the observed p-value must be
0.05 or lower. Not the configured confidence level: the p-value the data actually
produced.