Skip to main content
Two versions of a video. Traffic split between them. The winner decided by a statistical test on the metric you actually care about.

Create a test

A/B Tests → Create New A/B Test. Four steps.
1

Name it and pick the goal

The goal metric is what decides the winner:Choosing Custom Metric makes you pick which one. Define it first under Analytics → Custom metrics.
2

Configure the arms

  • Clone: take one video and clone it into 2, 3 or 4 arms you then edit.
  • Compare: pit existing, different videos against each other.
Set the traffic split (it must total 100%) and, optionally, a start and end date.
3

Pick the videos

One base video in clone mode; one video per arm in compare mode.
4

Review and launch

Check the split, then launch.
In clone mode the arms start out identical. A test between two identical videos measures nothing. Edit each arm (the thumbnail, the hook, the CTA timing, the offer) before you send traffic. The dashboard flags any arm still identical to the control.

Embed the test, not the video

Once a test is live, the video’s Embed tab switches to a loader snippet pointed at the test. That snippet is what rotates viewers between arms.
A hand-copied embed pinned to one video code does not rotate. Every viewer sees that one arm, your test collects nothing, and the analytics show zero plays on the others. Re-copy the snippet from the Embed tab after you launch.

How the winner is decided

Rate metrics (conversion rate, completion rate, custom metrics) use a two-proportion z-test. The denominator is unique plays, not raw plays, so a viewer who resumes or restarts cannot inflate an arm.Watch time uses a two-sample t-test.
A t-test needs the real variance of each arm. When it is not available, TrackPlay reports insufficient data and gives you the effect size only.It will not invent a p-value from an assumed variance. A fabricated p-value is worse than no p-value, because you would act on it.
More arms means more comparisons at once, which makes a false positive more likely. TrackPlay applies a Bonferroni correction: the significance threshold is divided by the number of comparisons.
A winner must clear both bars: statistically significant, and a large enough effect to be worth acting on. A 0.1% lift that is technically significant is not a reason to rebuild your funnel.

Reading the results

A/B Tests → your test.
  • Traffic Flow shows where traffic actually landed, which is not always where you configured it to land.
  • Statistical Analysis shows the p-value, the confidence interval, and the test statistic (labelled z or t depending on the metric).
  • Variant Performance compares the arms directly.

Declaring a winner

Click Declare Winner and pick an arm. All traffic moves to it. If the numbers are not ready, TrackPlay tells you why and makes you tick an explicit override before the button works.
Overriding the guardrail is sometimes the right call: a launch ends, a deadline lands. But you are choosing on a coin flip. Know that you are doing it.

Letting TrackPlay declare it for you

A test finishes itself. Ship the winner automatically is on for every new test, and you can switch it off on the test’s page at any time, even while it runs. It is the one setting you can change mid-run: the statistics themselves are locked once a test launches. Tests created before September 2026 keep whatever they were set to. With it on, TrackPlay checks the test once an hour and declares a winner only when every one of these is true:
1

The result is statistically significant

At the Bonferroni-corrected threshold described above.
2

The effect is large enough to act on

The same practical-significance bar a manual declaration has to clear.
3

Each arm has enough traffic

The minimum per-variant sample is met.
4

The test has run long enough

Your minimum duration, 7 days unless you change it. This exists so a test cannot be called on the first favourable hourly check.
5

The observed p-value clears your threshold

Your decision threshold, 0.95 by default, which means the observed p-value must be 0.05 or lower. Not the configured confidence level: the p-value the data actually produced.
When it fires, the test is marked completed, the winning arm goes to 100% of traffic, and every other arm goes to 0%. The result is recorded as automated, so you can always tell which tests you called and which the math called. TrackPlay emails the workspace owner when it happens, or the notification addresses on the test if it has any. Settings experiments (a single change tested inside one video, started from Customize or from an Insights suggestion) work the same way: the winning version is applied for you once it wins with confidence, unless you switched that off. A win for your current settings changes nothing.
This moves live traffic without asking. The guardrails above are real and every one of them must pass, but the outcome is still a traffic change you did not click. Leave it off if you want to make that call yourself.