You have six of something. Six products you could feature, six classes you could advertise, six links in a newsletter, six dishes you could put at the top of the menu. You want to promote the one people like best, and quietly retire the one nobody wants.
So you count the clicks. Or the orders, or the sign-ups. One option comes out ahead. You promote it.
Here is the uncomfortable part: you have measured your own layout at least as much as you have measured anyone's preferences. If those six are anywhere near each other in appeal, shuffling the list before you started would have crowned a different one by a similar margin. Counting harder does not fix this. Counting for longer does not fix this. The count was never the problem.
None of the fix needs software you don't have. It does need more patience than you were hoping for.
Five reasons people click the thing they click
When someone picks one option over another, at least five forces are pulling on them, and only one of them has anything to do with which option is better.
Position. This is the big one, and it is bigger than most people believe. The first item in a vertical list gets chosen more than the second, which gets chosen more than the third, and this holds whether the list is search results, a menu, a shelf, or a ballot paper. It is well enough established in elections that several US states rotate the order of candidates' names from district to district, so that nobody gets the top spot everywhere. Being listed first is worth a point or two in a normal race, and considerably more when voters know little about the candidates. Your six products are a race where the voters know very little about the candidates.
Layout, which decides where "first" is. In a vertical list, the top wins and the bottom picks up a smaller bump. In a horizontal row, the middle often wins instead. On a menu, the top of a section and the bottom of a section both outperform the squeezed middle. The useful conclusion is not "put it in the good spot". It is that there is no fixed correction you can apply, because the shape of the advantage changes with the shape of the layout. You cannot subtract position bias. You have to cancel it.
Effort. Any option that costs an extra tap, an extra scroll, or a click on "see more" is not competing on merit. It is competing on reachability, and losing. If two of your six live behind a "more options" link, you do not have six options in the running. You have four, and two footnotes.
Wording. Whether the label names a thing or a benefit, whether it is three words or nine, whether it happens to match the phrase already in the reader's head. Two labels for the identical underlying option can pull very different numbers, which means a click is partly a vote for the copywriting.
Familiarity. People pick what they recognise. Anything new starts behind anything established, regardless of quality, and this creates a loop that is worth naming: whatever you promoted last quarter is more familiar now, so it wins again, so you promote it again. Promotion manufactures its own evidence. If you have been running "featured items" for a while, some of your ranking is just an echo of last year's decisions.
Five forces. Rank your options by raw clicks and you have measured all five at once, mixed together, with no way to separate the one you care about.
Why your count is not an answer
You are counting the wrong thing. What you want is a preference, and a preference is a rate: how many people chose it, out of how many people had the chance to. Clicks are only the top half of that fraction. The bottom half is how many people saw the option at all, and almost nobody records it. If option one sits at eye level and option six is below the fold on a phone, comparing their click counts is comparing a busy shelf to a stockroom.
A zero has two meanings. This is the one that catches careful people. Most systems, and most people with a notebook, record something when a choice is made. Nothing gets recorded when a choice is not made. So an option that was put in front of a thousand people and picked by none of them leaves behind exactly as much evidence as an option nobody ever saw: none at all.
Which means your "least popular" list is quietly drawn only from the options that already worked at least once. The genuinely dead ones are not at the bottom of the list. They are absent from it. If you have ever looked at a ranking and thought "hang on, where is the blue one?", that is this.
How to measure it properly
Rotate the order
Give every option a turn in every position. With six options and six slots, that is a cycle: today's first is tomorrow's second, and after six rounds everything has been everywhere.
This is the ballot trick, and it works for the same reason. You are not removing the advantage of being first, which you cannot do. You are handing it out equally, so that it stops being an advantage for any particular option and becomes a constant that applies to all of them. The bias is still there. It just no longer has a favourite.
How often to rotate is a real decision with real consequences:
- Every visitor gets a random order. This cancels position fastest and most cleanly. The catch is bookkeeping: you now have to record what each visitor actually saw, or you cannot attribute anything to anything. If you shuffle the order but only write down the clicks, you have made your data worse than before, because now you do not even know which option was in the good spot.
- A fixed order per day, or per page, rotating on a cycle. Much easier to record, since one line in a notebook covers the whole day. The cost is that the rotation is now tangled up with everything else that varies by day. Tuesday's traffic is not Saturday's traffic, so make sure every option gets a turn on every weekday before you read anything into the result.
- Rotating once a month. Barely better than not rotating. The gap between changes is so long that each option is really being measured in a different season, against a different audience.
If you are doing this by hand, any randomiser will produce the order for you — a spin the wheel with your six options on it, spun once a day, is a perfectly respectable rotation schedule, and it stops you unconsciously putting your favourite near the top.
Rotation is not free. Once every option has been in every position, each option's clicks are spread across good spots and bad spots, so the average is fair but noisier than a single-position measurement would be. You are buying fairness with time. That is almost always the right trade, but it does mean the answer arrives later than you would like.
Count how many people saw each option
Write down the denominator. In a shop, that is how many people walked past the shelf, not how many picked something up. On a page, it is how many people had the option in front of them.
There is an important shortcut here that saves an enormous amount of traffic. If all six options are visible at once, then every single visitor counts as one chance for all six, and you only need to count visitors. Six hundred visitors gives you six hundred chances per option, all at the same time. If instead you show one option at a time and rotate the slot, each option only gets a sixth of your traffic, and you need six times as many people to reach the same confidence.
So: showing everything at once is dramatically cheaper to measure. It just makes position the dominant force, which is exactly why the rotation above matters.
Decide when you will stop, before you start
If you check the numbers every morning and stop as soon as one option is ahead, you will find a winner. You will find one even if all six options are identical, because six numbers wandering randomly will, sooner or later, produce a gap that looks impressive on the morning you happen to look. Continuous peeking with a stop-when-happy rule does not have a small error rate. Given enough mornings, it has a very large one.
Write down in advance how many visitors, or how many days, or how many clicks you are going to collect. Then collect them, then look. The stopping rule is part of the measurement, not an administrative detail around it.
How to know when you actually know
Here is the arithmetic, and it is genuinely one line.
A gap between two counts is only real if it is bigger than twice the square root of the two counts added together.
That is it. If one option got 40 clicks and another got 30, add them (70), take the square root (about 8.4), double it (about 17). The gap is 10. Ten is less than seventeen, so those two options are tied. Not "narrowly ahead". Tied. Run that experiment again next month and the 30 could easily come out on top.
The reason this works is that counts of independent events carry a wobble of roughly the square root of the count itself. Forty clicks is really "forty, give or take six or so", and once you allow both numbers to wobble, small leads evaporate.
A worked example
Six products on a shop's front page, all six visible at once, order rotated daily. After a couple of months, roughly a thousand visitors have seen each:
| Option | Clicks |
|---|---|
| A | 62 |
| B | 55 |
| C | 48 |
| D | 45 |
| E | 41 |
| F | 12 |
The instinct is to read this as a ranking, promote A and drop F. Half of that is right.
- A against B: gap of 7. The threshold is twice the square root of 117, which is about 22. Nowhere near. Tied.
- A against E: gap of 21, threshold about 20. This one just barely clears the bar, which in practice means "probably real, do not bet the shop on it".
- A against F: gap of 50, threshold about 17. Comfortably real.
So the honest reading of that table is not "A wins". It is: five options are indistinguishable, and one is clearly unpopular. That is a less satisfying sentence and a much more useful one, because it tells you the truth about where your evidence actually is.
Notice which end of the table carries the finding. Big gaps are easy to prove and small gaps are hard, so a well-run measurement usually gives you a confident answer about your worst option long before it gives you one about your best. That is not a flaw. Retiring the thing nobody wants is a real decision, and you can make it months earlier than the other one.
How much data do you need?
Work backwards from the same rule, and you get a table worth pinning up. This is how many clicks the leading option needs before a difference of a given size becomes visible:
| If the runner-up is truly... | The leader needs about |
|---|---|
| 50% as popular | 25 clicks |
| 80% as popular | 180 clicks |
| 90% as popular | 760 clicks |
| 95% as popular | 3,100 clicks |
Near-ties are extraordinarily expensive to resolve. Going from "is it double?" to "is it 5% better?" costs more than a hundred times the data. If your six options are genuinely similar, you will not separate them, and no amount of patience will change that at the traffic you have.
The other direction is more encouraging. Translate clicks into visitors: if about 5% of people who see an option click it, then 180 clicks needs roughly 3,600 people to have seen it. With all six visible at once, that is 3,600 visitors total, not 21,600. A small shop, a village noticeboard, a modest newsletter — that is months, not years. It is reachable.
And if it is not reachable? Then the honest answer is that you cannot tell, and you should stop pretending the numbers are deciding for you.
What to do with the answer
Promote the winner only if there is one. If your top five are tied, choosing between them is a coin flip wearing a lab coat. Pick on other grounds — which has the better margin, which you actually want to be known for, which you can supply — and be honest that the data did not choose it.
Act on the clear loser. This is where the confident finding usually lives. But check its exposure before you retire it: if F was below the fold on phones for half the test, F's twelve clicks are a fact about your layout, not about F.
Never delete an option that was never seen. Sounds obvious. It is the single most common way this goes wrong, because an unseen option and an unwanted one produce identical evidence, and only one of them deserves what happens next.
Re-measure after you promote. Once you move something to the top spot, its numbers will go up. That rise is not confirmation that you chose correctly. It is the position effect you spent all this effort cancelling, reappearing the moment you stopped cancelling it. If you want to keep learning, keep rotating.
Let "no winner" be a result. It is a genuinely valuable one. It means the choice is free, and you can spend it on something other than guessing.
The rule, with nothing attached
If you remember one sentence from this:
A count of choices is not a measure of preference until every choice had the same chance of being chosen.
Rotation, counting impressions, fixing your sample size in advance — all of it is bookkeeping in service of that one sentence. And the corollary, which is the part that costs people money: an option with no clicks and no measured exposure has told you nothing at all, and it is very easy to mistake that silence for a verdict.
What I would still be careful about
The square-root rule assumes each choice is independent and that exposure really is equal. Both can quietly fail. One customer clicking five times in a session is not five people. A rotation that gets interrupted for a week leaves one option with a fortnight in the best slot. Neither shows up in the arithmetic — they show up as a lead that does not reproduce.
And none of this tells you why an option won, only that it did. A label that happens to match the words in people's heads will beat a better product with a vaguer name, and the measurement will report that as a preference. If you want the why, you have to ask people, which is a different exercise with its own sharp edges.
The measurement will not make the decision for you. It will just stop you making a confident one for the wrong reason, which is most of the value.