researchtestingthumbnails

Ask an AI Which Thumbnail Is Better, Then Swap the Order

We showed an AI 161 pairs of real videos with a known winner, each pair in both orders. It picked whichever thumbnail came first 64.7% of the time, and changed its answer on almost a third of pairs.

September 14, 20268 min read
Two abstract thumbnails with a looping arrow showing their order being swapped

If you have ever pasted two thumbnails into an AI chat and asked which one will get more clicks, you ran an experiment with no control. We ran it with one: 161 pairs of real videos where we already knew which performed better, each pair shown to the AI twice, once in each order. Which thumbnail it saw first turned out to decide a large share of its answers.

Key takeaways

  • Shown two thumbnails from the same channel, the AI picked whichever came first 64.7% of the time. A judge with no preference for position would sit near 50%.
  • When the real winner happened to be shown first, the AI was right 77.3% of the time. When the same winner was shown second, 48.0%.
  • Averaged over both orders it was right 62.7% of the time, exactly matching a score that rates each thumbnail separately. Comparing side by side added no accuracy.
  • It gave opposite answers when nothing but the order changed on 30.7% of pairs.
  • It almost never admitted doubt: a typical answer claimed 85% confidence while the model was right about 63% of the time.

Can an AI tell you which of two thumbnails will perform better?

Sometimes, but not in a way you can trust from a single answer. Averaged over both orders, the model we tested picked the better-performing thumbnail 62.7% of the time — exactly as often as our score managed by rating each thumbnail separately. But its choice depended on which image it saw first: 77.3% right when the winner came first, 48.0% when it came second.

Put plainly, a single answer behaves like a coin that lands on the first option more often than not. The information is in there — when the model gives the same answer both ways round, it is right more often than chance — but you only find it by asking twice.

How we tested it: real pairs with a known answer, shown both ways round

The pairs came from our published test of the score: two videos from the same channel, published close together, one that clearly beat the channel's own recent average and one that clearly fell short. Using the same channel cancels subscriber count, audience and production budget, so what differs between the two is mostly the packaging.

We used the 161 pairs from half of the study's channels and kept the other half untouched. The AI — the same vision model that powers our score — was shown both thumbnails with their real titles, told that one had outperformed the channel's usual views and one had underperformed, and asked which was which. Eleven of the 322 answers came back unusable, which left 150 pairs answered in both orders.

Why every pair had to be shown twice

An AI comparing two options can prefer one position regardless of content, much as some people favour the first item on a list. If that preference exists and you only ever show each pair one way, you cannot tell it apart from skill: a judge that always picks the first image looks brilliant whenever the better thumbnail happens to be listed first.

Showing every pair twice, with the order swapped, separates the two. Skill should survive the swap. A preference for a position should flip with it.

The thumbnail shown first won most of the arguments

Across 300 answers, the AI picked whichever thumbnail it was shown first 64.7% of the time. That preference changes what any single answer is worth:

How the pair was shownHow often the AI picked the real winner
Winner shown first77.3%
Winner shown second48.0%
Average over both orders62.7%
Our score, rating each thumbnail separately62.7%
Same correct answer in both orders47.3% (random guessing: about 25%)

The average row is the fair comparison with a score that rates each thumbnail on its own, because both give one answer per pair. On that footing the two methods tied at 62.7%. Comparing side by side did not add accuracy; it added a dependence on order.

The last row is the stricter test: an answer counts only if the AI chose the same, correct thumbnail both times. A judge guessing at random would pass it about a quarter of the time, and one that always picked the first image would never pass it. The AI passed on 47.3% of pairs, so the signal is real — it is just tangled up with position.

It changed its mind on almost a third of pairs

On 46 of the 150 pairs, 30.7%, the AI gave opposite answers depending only on the order. Nothing about either thumbnail or either title had changed between the two questions.

On the other 104 pairs it was consistent, and on those it was right 68.3% of the time. Our score, on those same 104 pairs, was right 63.5% of the time. That gap is too small to be sure of with this many pairs, but it points somewhere useful: an answer that survives the swap is worth more than one that has never been tested.

Its confidence number is not a probability

We asked for a confidence figure with every answer and told the AI that 50 meant a coin flip. It barely used the lower half of the scale. Its lowest stated confidence across all 300 answers was 65, a typical answer claimed 85, and only 4 answers came in below 70.

Meanwhile it was right about 63% of the time. Answers it rated between 80 and 89 were right 61.6% of the time; answers it rated between 70 and 79 were right 63.5% of the time. Across the range where nearly all its answers fell, the number told you nothing about which answers to believe.

The 17 answers it rated 90 or above were right 82.4% of the time, which is suggestive, but 17 is too few to rely on. The practical reading is simple: a confident tone from an AI comparison is its default register, not evidence.

How to check any AI thumbnail judge in five minutes

We tested one model, and other tools may behave differently. You do not need to take our word for it either way, because the check is quick to run yourself:

  1. Pick two thumbnails where you already know the answer: two past videos from your own channel, published close together, one clearly above your usual views and one clearly below.
  2. Ask the AI which performed better, with the stronger one shown first. Note its answer and any confidence it gives.
  3. Start a fresh conversation, so it cannot remember its first answer, and ask again with the order reversed.
  4. Repeat with at least ten pairs. Count how often it picks whichever image comes first, and how often its answer survives the swap.
  5. Trust only the answers that survive the swap, and ignore the confidence figure entirely.

If you would rather have a number that does not depend on presentation order at all, you can score each thumbnail and title on its own. A separate score for each option, compared afterwards, has no first position to prefer.

What this experiment does not show

It tested one model with one set of instructions. A different model, or the same model asked differently, might prefer the first position more or less strongly. The point is not that every AI behaves like this one; it is that a single comparison cannot tell you whether the one you are using does.

The pairs were chosen because they differed sharply in performance, so these accuracy figures describe clear hits against clear misses, not two thumbnails that performed almost identically. Every channel was established and English-language.

And none of this is about real split testing. A test run on real viewers measures what those viewers actually did. This experiment concerns only an AI trying to predict that result in advance.

Frequently asked questions

Can an AI chatbot pick the best YouTube thumbnail?

Partly. The model we tested picked the better-performing thumbnail about 63% of the time on average, which beats chance but matches rather than beats scoring each thumbnail separately. Its answer depended heavily on which image it saw first, so a single reply is unreliable. Ask twice with the order reversed, and trust only answers that survive the swap.

What is position bias in AI comparisons?

Position bias is a tendency to prefer an option because of where it appears rather than what it contains. In our test, the AI chose whichever thumbnail it was shown first 64.7% of the time. The remedy is to present every comparison in both orders and treat an answer that flips as no answer at all.

Is it better to rate thumbnails separately or compare them side by side?

In our test the two were equally accurate on average, at 62.7%. The difference was stability. Rating each thumbnail on its own gives the same result whatever order you check them in, while comparing side by side produced answers that changed with the order on almost a third of pairs.

Why does the AI sound so sure of its answer?

Because confident phrasing is its default, not a measurement. Asked to rate its own certainty, the model we tested almost never went below 70 out of 100 and typically claimed 85, while being right about 63% of the time. Within that range, a higher stated confidence did not mean a more reliable answer.

Does this mean thumbnail A/B testing is pointless?

No. A real test shows your thumbnails to real viewers and measures what they do, which no prediction replaces. This experiment only concerns asking an AI to guess the winner beforehand. For a real test, the harder question is whether the difference you see is larger than your channel's normal variation from one upload to the next.