Experimentation Measurement Austin, Texas

A million people saw the ad. 385 bought. Do you know which step took the rest?

Everybody has a guess about that step, and the guess is usually the part of the site that is easiest to redesign. We measure which step is actually costing you, then run the tests that move it, and we tell you in writing when a test cannot answer the question.

Run 01  /  checkout copy test  /  scroll to advance the days

The same test says yes on day one, no on day five, and yes again on day thirteen.

Not significant
Day
1
Visitors per variant
5,800
Observed lift
+21.4%
95% interval
+1.2% to +41.6%

Simulated run, not a client result. Baseline 3.24%, true lift 6.8%, 5,800 visitors per variant per day, two sided 95% interval on the relative difference. The interval is recalculated from the sample collected so far, exactly as it would be in a live test.

02 / 06

Teams call winners because they stopped looking at the right moment.

Every number below comes from people who ran experiments at scale and wrote down what happened. None of them come from us, and every one has a link.

1 in 3
ideas tested on the Microsoft experimentation platform improved the metric they were designed to improve.
Kohavi & Longbotham, 2015, Encyclopedia of ML & DM
~10%
of Google controlled experiments led to business changes, in the figure Jim Manzi reported.
Manzi, 2012, quoted in Kohavi & Longbotham, 2015
80%
of the time we are wrong about what a customer wants, in Avinash Kaushik's testing primer.
Kaushik, 2006, quoted in Kohavi & Longbotham, 2015
5%
of true null results still come back significant at a 95% confidence level. That is the design, not a bug.
Kohavi, Longbotham, Sommerfield et al., 2009

The fourth number is the one that costs money. If you run twenty tests where nothing is really happening, one of them will still look like a win at the usual threshold. Peek at the dashboard every morning and that rate climbs well past one in twenty.

03 / 06

Which part of that is costing you the most?

Three ways we work, depending on whether the problem is what you are testing, whether the numbers can be trusted, or whether one particular result is about to be believed.

01 · Test programme

You get a decision, not an argument

A sized backlog, a stopping day agreed before the code exists, and one page per readout. Including the readouts that say no.

  • Sized before anyone builds it
  • One metric, agreed in advance
  • A written no when it does not hold
02 · Measurement build

The same question gets the same answer twice

Most disputed results are a tracking problem wearing a statistics costume. We rebuild the event layer so the arguing stops.

  • Tracking plan tied to real questions
  • Warehouse tables you can query without us
  • Dashboards with intervals, not just points
03 · Analysis review

Somebody checks it before you scale it

A 12% lift is about to become a roadmap. We check the sample, the segments and the stopping point first.

  • Re analysis from raw assignment data
  • Peeking and sample ratio checks
  • One page you can hand to your CFO
See what each one includes
04 / 06

Sample size and duration, before you build anything.

Move the four sliders. The tool gives you the sample you need per variant and how long your traffic takes to produce it. Nothing is sent anywhere until you ask for the plan.

Sample needed per variant
Time to collect it
Move a slider to start.

We reply from a person, usually within one working day.

Method behind the number

Two proportion comparison, two sided. Per variant sample is n = (zα/2·√(2·p̄·q̄) + zβ·√(p₁q₁ + p₂q₂))² / (p₂ − p₁)², where p₁ is your current rate, p₂ is that rate lifted by the effect you chose, and p̄ is their average. Traffic is split evenly across variants. Duration rounds up to whole days, and we round again to whole weeks in the advice because partial weeks bias a test toward whichever days it happens to catch. This is our reading of standard practice, not advice for your specific case, so check it with your own analyst before you commit a quarter to it.

05 / 06

Field notes

All notes →
06 / 06

Tell us what you are about to ship.

Send the change you are planning and the traffic you get. We will tell you whether the test can answer the question at all, and we will say so plainly when it cannot.

Small Lift Labs
Austin, Texas
Monday to Thursday, 9 to 6 Central
Sources on this page
  1. Kohavi, R. and Longbotham, R. Online Controlled Experiments and A/B Testing. Encyclopedia of Machine Learning and Data Mining, 2015. PDF. Checked 12 August 2026.
  2. Manzi, J. Uncontrolled, 2012, as quoted in the paper above.
  3. Kaushik, A. Experimentation and Testing primer, 2006, as quoted in the paper above.
  4. Moran, M. Do It Wrong Quickly, 2007, p. 240, as quoted in the paper above.
  5. Baymard Institute. Cart abandonment rate statistics, a meta analysis of 50 published studies, average 70.22%. Source. Checked 12 August 2026.