Home/Meta Ads/Testing Ad Creative

Testing ad creative: how to tell if humour actually works

Tone gets decided in meetings by whoever is most senior, and it is one of the few creative questions you can settle with evidence instead. A humour test is cheap, it takes about three weeks, and it fails in a very specific way: the funny version nearly always wins the metric that does not pay you. Here is how to run one that answers the question you actually asked.

Moeez Abbas, founder of BoltClicks

By Moeez Abbas

Published

START HERE

The argument nobody can win in the room

Somebody suggests a funnier ad. Somebody else says it will make the business look unserious. Both are certain, neither has evidence, and the person with the most seniority wins. That is how most tone decisions in advertising get made, and it is a poor way to spend money.

The mistake is treating tone as a matter of taste. It is not. It is a variable, it produces a measurable difference in cost per result, and you can find out which way it goes for your business in about three weeks for the price of a test.

This is how to run that test so the answer means something — and, just as importantly, how to avoid the versions of it that produce a confident conclusion from numbers that never supported one.

DEFINITION

A test changes one thing, and holds the rest still

Most things people call creative tests are not tests. They are two different ads, aimed at two different audiences, pointing at two different pages, launched in different weeks. When one performs better, nobody can say which of the four differences caused it.

A real test holds everything constant except the thing you are asking about. If the question is tone, then the offer stays the same, the audience stays the same, the landing page stays the same, the budget stays the same and the dates are the same. The only difference between the two ads is that one is funny and one is not.

That sounds obvious written down. In practice it is hard, because the moment you write a humorous version you will want to change the headline too, and shorten the copy, and pick a livelier image. Every one of those changes is another variable, and each one costs you the ability to answer the question you started with.

Write the straight version first. Then change only the tone.

A creative team laughing as they pin two ad versions, a cartoon dog and a pair of trainers, to a glass wall
Judged on engagement the funny version wins. Judged on what an enquiry costs, it does not. Which number you agreed on beforehand settles it.
THE WRONG METRIC

Funny ads win the metric that does not pay you

Here is the trap, and almost everybody falls into it once.

Humour reliably produces engagement. People react to it, comment on it, tag each other under it and share it. If you judge the test on likes, comments, shares or reach, the funny version will usually win, and it will win convincingly.

None of those things is what you are buying. You are buying enquiries, bookings or sales. The comparison that matters is cost per result — the actual result, defined before the test started — and on that measure the answer is genuinely uncertain in advance. Sometimes the funny version wins there too. Sometimes it collects a great deal of attention from people who were never going to buy anything, at a higher cost per enquiry than the plain version it replaced.

Decide the single number you will judge on before either ad goes live, and write it down. The discipline sounds excessive until you have watched somebody scale a campaign because it got a hundred comments.

PREREQUISITES

What has to be true before a test can tell you anything

A creative test is only as good as the measurement underneath it. Three things need to be in place first, and if any of them is missing the test will produce a number that looks like an answer and is not.

  • Conversions are tracked, and tracked as the right event. If the platform is counting page views or link clicks as your result, the test measures curiosity rather than intent — and curiosity is exactly what humour inflates. The event has to be the enquiry, the booking or the sale.
  • There is enough volume to separate the two. A handful of results split across two ads tells you nothing. If a week of spend produces single-digit conversions in total, you cannot run a tone test yet; you can only run one ad properly and learn from it slowly.
  • The landing page is the same and is not the bottleneck. If the page is slow or confusing, both versions fail and you learn nothing about tone. A test sitting on a page that loses people before the form measures the page, not the creative.

If those three are not true, the honest answer is that you are not ready to test tone, and the money is better spent fixing whichever one is missing. That is a less satisfying conclusion than a test result, and it is usually the correct one.

THE BUILD

Making the funny version a fair comparison

The humorous variant has to be a serious piece of work, or the test is rigged before it starts. A half-hearted joke losing to a well-made straight ad tells you nothing about humour.

Keep the claim identical. Whatever the plain ad promises, the funny one promises exactly the same thing. Humour goes into how it is said, never into what is offered. The moment the two ads make different promises you are testing offers, not tone.

Keep the call to action identical and in the same place. Keep the format identical — a static image against a static image, a video against a video of similar length. A funny video against a plain image is not a tone test.

And make the joke land for the buyer rather than for the industry. Humour that depends on knowing the trade is a joke for competitors, not for customers, and it is remarkably easy to write by accident when you know a subject well. If somebody outside the business does not get it in two seconds, it will not work in a feed.

READING IT

Reading the result without fooling yourself

Three things go wrong at this stage, and they go wrong in a predictable order.

Calling it too early. Campaigns need time to settle before the numbers mean anything, and the first days of any new ad are the least representative days it will ever have. Deciding on day two is not decisiveness, it is reading noise. Agree the run length in advance and leave it alone.

Treating a small gap as a result. If one version comes in at a slightly lower cost per enquiry on a modest number of conversions, the honest reading is that you did not find a difference. That is a legitimate outcome and a useful one: it means tone is not your constraint, and the money should go somewhere that is.

Editing mid-test. Every meaningful change restarts the learning, and a test you adjusted halfway through is not a test any more. If you cannot resist adjusting it, you are not ready to run it.

The outcome you want is one of three sentences: the funny version costs less per enquiry, the plain version costs less per enquiry, or there was no clear difference. All three are worth knowing. Only the third is disappointing, and it is still cheaper than arguing about it for a year.

CONTEXT

Where humour helps, and where it quietly costs you

The test answers the question for your business, which is the only answer that matters. But there are patterns worth knowing before you design it.

Humour tends to have more room where the purchase is low-risk, where the category is crowded and largely undifferentiated, and where being remembered later is part of the point. If everybody in your market says the same sensible thing in the same sensible way, being the one that is enjoyable to look at is a genuine advantage.

It tends to have less room where the buyer is frightened, where the decision carries professional or financial consequence, and where the searcher is acting urgently. Somebody choosing a solicitor, or trying to get a burst pipe fixed tonight, is not in a receptive mood. A joke there does not read as personality; it reads as not taking the problem seriously.

There is also a cost that never shows up in the ad account. A joke that misfires in public attaches to the business, not to the campaign, and it outlives the budget. That risk is real, it is asymmetric, and it is a reasonable argument for testing humour on a small share of spend rather than a large one.

AFTERWARDS

What to do once you have an answer

A single result is a finding, not a law. Treat it as the first entry in a record rather than a permanent decision about your brand.

Move budget to the winner, but keep the loser running at a small share. Creative fatigues, audiences shift, and the version that lost in March is occasionally the version that wins in September. Keeping a fraction of spend on the alternative is cheap insurance against believing an old answer for too long.

Then test the next variable. Tone was one question; the offer, the opening line, the format and the audience are all separate questions, and each needs its own clean comparison. This is the ordinary rhythm of running paid social properly — a small number of deliberate questions asked one at a time, rather than a redesign every quarter.

Keep a written log of what was tested, what the result was, and how many conversions it rested on. That last column is the one that stops a thin result being quoted as settled fact eighteen months later, which is how most agency folklore gets made.

And carry the finding across channels carefully. A tone that works in a social feed, where people are browsing, does not automatically work in a search campaign, where somebody has typed a problem and wants it solved. Different mood, different test.

COMMON QUESTIONS

Creative testing questions, answered plainly

How long should a creative test run?

Long enough for the campaign to settle and to accumulate enough conversions that the gap between the two versions is not noise. In practice that usually means weeks rather than days, and it depends far more on your conversion volume than on the calendar. A useful rule: if you would not bet your own money on the result, it has not run long enough.

How many conversions do I need before the result means anything?

There is no single number, and anybody who gives you one without knowing your conversion rate is guessing. The practical test is to ask what would happen if two or three conversions landed on the other side instead. If that would flip the winner, you do not have a result, you have a coincidence. If it would not, you probably do.

Can I test humour and a new offer at the same time?

You can run both, but you will not be able to attribute the outcome to either. If the combined version wins you will not know whether it was the joke or the discount, and you will carry the wrong lesson into the next campaign. Test the offer first, since it almost always moves results more than tone does, then test tone against the winning offer.

What if the funny ad gets far more engagement but fewer enquiries?

Then it lost, and it lost usefully. That specific pattern — high engagement, worse cost per enquiry — is the most common outcome of a humour test, and it is the reason the metric has to be agreed in advance. Attention that does not convert is not a partial success; you paid for it and it did not produce work.

Does humour hurt a business that wants to seem professional?

Not automatically, and the assumption that it does is exactly the untested belief worth checking. What genuinely damages credibility is a joke about the customer’s problem rather than about your own category or yourself. Self-deprecating rarely offends; making light of somebody’s flooded kitchen or their legal trouble does.

Should I copy a competitor’s funny campaign if it seems to be working?

You cannot see whether it is working. You can see that it is running and that people are commenting, and neither tells you what it costs them per enquiry. Campaigns run for months while performing badly, for reasons ranging from brand budgets to nobody checking. Treat a competitor’s creative as a hypothesis worth testing on your own account, never as evidence.

Is a small budget enough to test creative at all?

Often not, and this is worth being honest about. Splitting a small daily budget across two ads can mean neither accumulates enough data to exit the learning phase, so you get two unstable results instead of one reliable one. With limited spend, run one well-built ad, learn from it, then replace it with a deliberate variation and compare the periods. It is slower and less clean, but it is better than a split test that cannot resolve.

Who should write the humorous version?

Somebody who knows the buyer rather than somebody who knows the industry. The most common failure is an in-joke that everybody inside the business finds funny and nobody outside it understands. Show the draft to a person who does not work in your trade, give them two seconds, and ask what it means. If they hesitate, rewrite it.

NEXT STEP

Send the account before you argue about tone

Give me access to the ad account and I will tell you whether you can run a tone test at all — whether the conversion event is the right one, whether there is enough volume to separate two versions, and whether the landing page would sink both of them. It costs nothing and there is nothing to sign. If the honest answer is that your budget cannot resolve a split test yet, that is what the answer will say.


Scroll to Top