Traffic is split: half see version A, half see version B, and the only difference is the one thing being tested, a headline, a price presentation, the order of form fields. When enough visitors have passed through, the difference in conversion is either real or noise, and statistics can tell you which. The discipline is testing one meaningful change at a time and letting the test finish.
Two things. First, testing trivia: button colors on a page whose headline is the actual problem. Test the big levers first, the offer, the promise, the first screen. Second, sample size: a local business with 300 visitors a month cannot detect small differences, so it should test big swings, entirely different headlines or offers, where the winner is obvious sooner.
Honesty over hype: if a page gets a few hundred visits a month, session recordings and five real customer conversations beat a three-month test. Instrument first, understand second, test third. The point was never the test, it is compounding: each verified win becomes the new baseline, and baselines stack.
Until it reaches the sample size you calculated before starting, usually at least one or two full business cycles so weekday and weekend behavior both count. Stopping early because you like the trend is how teams fool themselves.
The measurement matters more than the tool. Clean event tracking on the outcome, a consistent split, and patience outperform an expensive platform used sloppily.
The promise above the fold on the money page, then the first step of the booking or contact flow. Those two levers dwarf everything else.
Everything in this glossary, we build and operate for real businesses. Thirty minutes maps it to yours.