Skip to content
February 2025 ARVIC

Creative Testing for Ad Effectiveness: Super Bowl Edition

Bill McDow, President of Accelerant Research, walks through the firm's framework for creative asset testing using Super Bowl ads as a live case study. Drawing on surveys of roughly 45 to 50 Super Bowl spots, with over 12,000 survey completes collected within days of the game, he covers how to structure an ad testing program, which metrics matter most at different funnel stages, and how to interpret results in context rather than in isolation.

Key Takeaways

  • Brand recall is table stakes. A high-scoring ad that fails to create a clear brand association is a problem, regardless of how well it performs on other measures.
  • Top-of-funnel metrics (recall, likeability, engagement, viral potential) are the ones creative actually moves. Bottom-of-funnel metrics (consideration, intent, recommendation) are driven more by category, brand equity, and purchase cycle than by any single ad.
  • Always compare within context. B2B ads, niche categories, and cause-driven spots will score differently than general consumer ads. Benchmarking them against snack food or beer norms will mislead rather than inform.
  • Dial testing is a practical diagnostic tool. Tracking moment-by-moment audience response helps identify which content to cut for shorter edits, flag potentially offensive scenes, and pinpoint when talent or characters help or hurt.
  • Build your own norms over time. Third-party benchmarks are a useful starting point, but teams running repeated ad tests should standardize their approach and develop internal benchmarks that reflect their own category and audience.
  • Consumers are irrational, and that is a feature, not a bug. Pre/post Super Bowl survey data showing a 20-point swing in admitted Chiefs support after the Eagles won is a useful reminder to always probe the 'why' behind scores, not just the scores themselves.

Questions & Answers

Did ads that elicited humor and nostalgia perform better than those that elicited disgust or discomfort?
Bill noted that this year's results did not show a clean pattern in that direction, but acknowledged it is exactly the right type of question to pursue. As an organization runs more ad tests, it becomes possible to assess whether certain emotional triggers, such as humor, nostalgia, or heartstring moments, consistently outperform others. He encouraged teams to build that analysis into their programs over time rather than relying on individual tests.
How did results vary by demographics or psychographics?
All results presented were general population, census-balanced. Bill emphasized that every survey should capture demographic, brand usage, category usage, and psychographic data so results can be sliced by subgroup. An ad that underperforms overall may perform significantly better among a specific target segment, which can inform media placement decisions even if the ad is not a top overall performer.
What were the OpenAI/ChatGPT ad results?
The ad was among the lowest performers on several measures, including brand recall. Bill speculated that ambiguity between the OpenAI and ChatGPT brands may have contributed to the weak recall score.

Session Notes

Context and Setup

This session was presented at the February 2025 Accelerant Research Virtual Insights Conference (ARVIC). Accelerant Research tested approximately 45 to 50 Super Bowl ads that aired that year, running each through a roughly 5 to 8 minute survey with approximately 250 respondents per ad. In total, the team collected over 12,000 survey completes within days of the game.

Pre/Post Super Bowl Consumer Survey Findings

Accelerant runs a pre-game and post-game survey each year to track what consumers were excited about going in versus what they felt delivered afterward. Key observations from this year:

  • The game itself was considered a 'bit of a yawner.' Excitement going in was higher than satisfaction coming out.
  • The halftime show was fairly anticipated and generally delivered, despite some social media discussion afterward.
  • Ad excitement was roughly flat pre- to post-game, an improvement over the prior year when ads significantly disappointed.
  • A notable data point: before the game, Chiefs and Eagles fan support was roughly even. After the Eagles won, 20 percentage points fewer respondents were willing to admit they had been rooting for the Chiefs. This illustrates consumer irrationality and the importance of understanding the 'why' behind any survey result.

The Creative Testing Framework

Accelerant's approach to communications research is organized around five core questions:

  1. What should we say? (early-stage strategic research)
  2. How should we say it? (creative development)
  3. Where should we say it? (channel and media mix)
  4. How well did it work? (pre- and post-release evaluation)
  5. What does it mean? (synthesis and action)

The session focused primarily on the evaluation stage, specifically testing finished or near-finished creative assets before or after they go to market.

Survey Metrics in the Scorecard

The standard scorecard captures the following measures, each rated on a five-point scale (except brand recall):

  • Brand recall: Did respondents correctly associate the ad with the right brand? This is the baseline. An ad that scores well on everything else but fails here has a fundamental problem.
  • Likeability: Overall positive or negative reaction to the ad.
  • Engagement: Likelihood to actually watch the ad rather than skip it.
  • Viral potential: Likelihood to share the ad on social media, a proxy for earned impressions and ongoing reach.
  • Consideration: Likelihood to consider the brand or product.
  • Intent: Likelihood to purchase or use.
  • Likely to recommend: Net Promoter-style forward intent.
  • Relevance: How personally relevant the ad felt.
  • Comparison to other brands: Relative competitive perception.

Top-of-Funnel vs. Bottom-of-Funnel Metrics

Recall, likeability, engagement, and viral potential are directly influenced by the creative itself and are the primary measures to optimize. Consideration, intent, recommendation, and relevance are much harder for a single piece of creative to move. These are shaped by factors outside the ad, including category dynamics, brand equity, purchase cycle length, and product complexity. Teams should not judge an ad as a failure because it did not spike consideration scores in isolation.

Dial Testing as a Diagnostic Tool

Dial testing tracks audience sentiment on a moment-by-moment basis as respondents watch an ad. Participants move a slider left (dislike) or right (like), and the result is a line chart overlaid on the video playback. Practical applications include:

  • Identifying which content to drop when cutting a 60-second spot down to 30 or 15 seconds. Segments where the dial dips are natural cut candidates.
  • Flagging scenes or characters that may generate unexpectedly negative reactions.
  • Pinpointing high-impact moments, such as a celebrity cameo or a punchline, that should be preserved or featured earlier.
  • Serving as a digital replacement for in-person dial testing sessions that were once conducted with physical paddles in a room.

Using Norms and Benchmarks

A raw score of, say, 81% likeability is meaningless without a point of comparison. Norms provide that context. Accelerant maintains a database of benchmarks from hundreds of ad tests. However, Bill emphasized that the most valuable long-term approach is for organizations to build their own internal norms by standardizing their testing methodology over time. Third-party norms are a useful starting point but should not be applied blindly, especially when category, audience, or business model differ from the norm base.

Super Bowl Ad Results: Category Highlights

Brand Recall: Who Landed the Association

  • Strongest: Doritos, Ritz, and Pringles. Salty snack brands led on brand recall this year.
  • Weakest: OpenAI/ChatGPT, the Snoop Dogg and Tom Brady prosocial spot ('No Reason to Hate' or similar), and the Meta/Ray-Ban collaboration. In several of these cases, context explains the result. A public service-style ad with no clear brand to purchase will naturally underperform on recall.

Likeability: What Audiences Enjoyed

  • High performers: Lay's (the spot following a little girl planting a potato and watching it grow into a bag of chips), Budweiser Clydesdales, and Booking.com featuring the Muppets.
  • Lower performers: CoffeeMate ('traveling tongue' concept), a 'head shaped like a cowboy hat' B2B-style ad, and Dunkin', which had a strong showing the prior year but did not repeat that performance.

Engagement: Would Viewers Actually Watch?

  • High performers: Budweiser Clydesdales, Lay's, and Michelob Ultra (Pickleball Hustlers).
  • Lower performers: CoffeeMate, DoorDash ('Got Trouble With Your Finances'), and a soccer-themed coffee spot featuring Channing Tatum.

Viral Potential: Earned Reach After the Game

  • Highest: The Angel Soft spot, which was structured more as a public service message than a traditional ad, scored extremely high on viral potential despite underperforming on most other measures.
  • Also strong: The Snoop Dogg/Tom Brady prosocial spot and Budweiser Clydesdales.
  • Weakest: CoffeeMate, Hims & Hers, and DoorDash.

Notable Individual Ad Observations

  • Helman's (When Harry Met Sally homage): A 60-second spot that was generally well received. Dial testing showed a dip in the second 30 seconds, suggesting a clean 30-second cut was possible without losing the key content. A spike at the end when Sydney Sweeney delivers the iconic line was a clear retention anchor.
  • Little Caesars vs. Pringles (facial hair battle): Both ads around flying facial hair, mustaches for Pringles and eyebrows for Little Caesars, were strongly received. Little Caesars had a slight edge. CoffeeMate's 'traveling tongue' concept used similar visual logic but failed to land with audiences.
  • Meta/Ray-Ban collaboration: Low brand recall is a common outcome for co-branded or 'collab' ads where it is unclear which brand owns the message. Instacart's multi-brand mashup ad was received more favorably and avoided that pitfall.
  • Michelob Ultra (Pickleball Hustlers): Performed well across recall, likeability, and engagement.
  • Hims & Hers (weight loss): Performed roughly at norm on likeability due to strong imagery but struggled on bottom-of-funnel measures. This is expected for a product category involving injections or significant side effects, where likelihood to purchase is inherently lower and harder to move with advertising alone.
  • DoorDash ('Got Trouble With Your Finances'): Missed the mark across multiple measures.
  • OpenAI/ChatGPT: Among the lowest performers on several measures including brand recall, partly because the brand distinction between OpenAI and ChatGPT may have confused respondents.
  • B2B ads (GoDaddy, Squarespace): Consistent with prior patterns, B2B-oriented ads tend to underperform against general consumer norms in a Super Bowl context. Small businesses represent roughly 15 to 20 percent of the general population, making Super Bowl targeting feasible but inherently challenging when measured against snack or beverage category benchmarks.

Recommendations for Building an Ad Testing Program

  1. Test earlier in the pipeline. Animatic or storyboard-level testing can surface major issues before production costs escalate.
  2. Standardize your methodology so results are comparable over time.
  3. Build internal norms. The more tests you run, the more meaningful your benchmarks become.
  4. Always interpret results in context. Category, target audience, campaign objective, and brand equity all shape what a score means.
  5. Use dial testing diagnostically. It is particularly valuable for editing decisions, talent evaluation, and identifying scenes that may generate unintended negative reactions.
  6. Segment your data. Gen pop results tell you one thing. Slicing by demographics, psychographics, and category usage often reveals where an ad performs best and where it should be placed.

Transcript

Read the full transcript

hello everyone welcome to uh the February excuse me 11th accelerant research virtual insights conference uh my name is Bill McDow I am the president at accelerant research and I'm going to be your host for today's uh festivities as it were I'm going to apologize ahead of time because uh my voice may be a little in and out I've been under the weather weather last week um but you all get the benefit from my cool smooth R&B radio d a voice that may kind of go in and out and Fade Into You Know 13-year-old pre prepubescent cracking um there's no telling what may happen um so welcome everybody if you've never attended one of our uh webinars or virtual conferences in the past uh welcome if you have attended before welcome back appreciate it um if you have attended before you probably know kind of the Spiel that we tend to give at the beginning of these uh these events these webinars um and this is our ground rules that we operate by we started doing uh these webinars uh back in Co times actually during quarantine um and we were really big on these uh these rules because nobody really knew what we were doing this whole Zoom thing was brand new uh these are typically not things that we have to to deal with too much but you never know right so we are at the mercy of tech uh if something did crap out if our stream went out uh we're going to apologize uh hopefully that doesn't happen uh be respectful is our other rule to live by and that is really just uh to one another uh to me as a presenter you can say what you will I'm okay with that uh but no I want you all to uh Network engage interact with one another um to the extent possible this is not a an in-person conference obviously this is a digital virtual one but if you can get on the socials start interacting with one another use the hashtag a RV I um or if you are in the zoom or YouTube um well you take advantage of the the chat function the Q&A uh drop questions drop comments um that's kind of what we're here for I would love nothing more than um to see a lot of chatter and offshoot conversations even um if you're at an in-person conference it's kind of rude to to have sidebar conversations for something like this yeah go for it why not um it's totally fine I think we're what we're going to be talking about is you off the press's Super Bowl ads and some of the the information that that we've been seeing so if yall want to debate with one another talk with one another one another please feel free uh before we start getting into our content what I would like to do is just kind of give the heads up I mentioned we we do a lot of these virtual conferences webinars uh got a couple of events coming up one of which is going to be later this month on the 27th of February uh Brett from my team team as well as Pavo from Amazon are going to be getting together having one of these mini conferences uh just similar format to today uh they're going to be talking about transforming insights into results uh advancing influence in corporate Landscapes from the bottom up um which is a really good and kind of you know the intent of you know what we do is wanting to basically rise the tide as it were for for researchers and insights Prof professionals uh we want to be basically learning with one another so let's use these opportunities we want to be sharing and knowledge sharing as it were right um coming up in excuse me in March uh March 27th is going to be one of our more we'll call it traditional uh but original format uh all day virtual insights conferences and what those tend to be is just kind of just back-to-back webinars over the course of the day we bring in some really compe in speakers who're going to talk about lots of different topics uh within research um and impacting of and by research um and then next one is going to be March 27th like I say and if you want to sign up for any of these upcoming webinars upcoming events just go to our website Excell research.com and you'll click on the resources dropdown conferences and you'll have all the info about upcoming agendas uh rosters and you know you can go ahead and register and sign up for the conferences all right so today's conference uh myself I will be presenting on Creative testing for at Effectiveness Super Bowl Edition um and as we all know Super Bowl was a couple couple of days ago uh so we got a lot of hot off the presses uh information results um the idea is we want to just basically share a framework of creative asset testing right so if you're going to be doing ad testing over time this is a framework that you can use um you know this is best practices that we have seen at excelerant research um I think I want to definitely in the spirit of of our conversation uh not always the right recipe for all particular corporations for particular researchers um but that's the whole idea of this float some ideas share some interesting content and maybe we can all learn a little something right uh before we get going I would definitely uh like I said just kind of encourage everyone so we're talking about Super Bowl ads right so if you have your favorites if you you know have particular ones that you were interested in like I say jump in the chat jump in the Q&A uh share tell me about what you thought um any opinions about who did it well who didn't do it well this past Super Bowl um as we know almost everybody watched so this is obviously you know compelling and fresh content um Okay so and kind of also in in that Spirit or in that vein um we're we tested a bunch of Super Bowl ads right the Super Bowl airon on Sunday we at excellent research have tested uh I believe 45 50 of the ads that did air um we run them through uh 250 person um just about 5 to 8 minutes survey um for each of the ads that were tested so I think that translates to you know 56 the ads uh over 12,000 survey completes that we conducted so huge shout out to my team at accelerant because you know within the couple of days that's that's a pretty Herculean effort to be able to fi a bunch of surveys uh analyze a bunch of data and put together some results so I can talk about them today um so what I'm showing right now is just conceptual framework for uh conducting ad testing or or Communications research um and I like to really Bo boil that down to just questions all right so you know what are you looking to to glean from this research or This research program uh what do you want an ad to say how should you say it um where to say it how well did it work um and typically internal marketing agency business partners these are folks that can kind of glom on to pretty simple questions like these right um depending on what your objectives are depending on which of these questions you you need answered uh when you are looking to do Communications research uh the type of research that's going to align to that can vary quite a bit um so you can have qualitative or quantitative solutions that can basically be used to to answer each of these questions right so if we are at you know initial creative brief stage of research just trying to understand you know what should we build or what should we you know include in our marketing campaign well that's going to be an early on kind of phase one of you just understanding what to say how you know what are the uh emotional functional you know communication levers that we should be even be pulling in this advertising that is just beginning to take shape so very front of pipeline trying to understand you know what what to test um focus of what we're going to be looking at today is going to be primarily down here at the bottom and that's just going to be evaluation Reserve so a thing an ad has been created and now you want to test that thing and you want to understand how well that thing performs um either before or after having actually released it to market right um before we jump into results I just want to give you know the very brief sort of uh time share pitch of of the webinar um no Agro USA so accelerant we have our own panel um our panel is actually among the larg lest in the business when it comes to um proprietary panels um we are very proud of the level of quality that that that we have in in our panel uh our panel is basically the lifeblood and the raw material for any research that we conduct either qualitative or quantitative um and it's the source for which we um gathered these insights that we're going to be talking about today um Second To None I recommend it it's one of the best in the big is uh quality as you know anybody who's been in research especially for the past several years knows um has become a huge issue in our business and one that we we take very seriously and ultimately looking to take that raw material of just the people taking part in research and deliver it and produce great insights based on it right okay uh before we jump in um we actually conducted a little bit of you know having nothing to do with the ads themselves particularly uh we like to do every year a pre and a post uh Super Bowl pregame and and postgame survey and you know some of the interesting findings again these are kind of hot off the presses uh but this year versus years prior um so we look at pregame postgame so what were people consumers uh excited about for for the big game um on from the pregame standpoint and what did they like best about the game afterward right so the game itself no real surprises here it was a bit of a yawner um so you know folks were more excited about the game going into it uh less so coming out uh halftime show even in spite of you know some of the uh kind of social murmurings that that have been going on on afterward uh halftime show you know was fairly anticipated and kind of delivered from that standpoint uh the ads themselves a little bit of a dip but if I look at actually um last year's results and that's what I'm flashing right now um the game much more compelling the ads themselves kind of disappointed last year uh if we go back to to this year you know fairly flat and even um and then we've got this other response which better Pro post game um don't actually have an open end associated with that but I can assume that can only be like Tor Swift right it has to be post game results uh another interesting finding and this is kind of to underscore a lot of what we uh what we need to always keep in mind as researchers um and especially kind of the subject matter that we're delivering uh is that consumers at the end of the day um and those are the folks that we're trying to measure and understand and find out what makes them tick so that we can deliver better ads in this case better products better experiences um they're irrational and and kind of crazy so you know it's important that as researchers we understand that and bring that context into whatever it is that we're doing and this is kind of from our prepost game survey little fun fact that's always interesting um to kind of underscore this right people are crazy so if we ask people before the Super Bowl who they're pulling for got a pretty even split between Chiefs fans and Eagles fans if we ask there a demographically ident identical group of people after the game who were pulling for now you would expect for these to to be even Steven right but everybody loves a winner and more far fewer people in fact 20 points fewer people are willing to admit that they were pulling for the Chiefs going into the game um again fun fact just kind of underscores the fact that that you know consumers are irrational they do crazy things but you know it keeps us employed as researchers because we're always chasing that consumer mindset as a were right all right so I got a little poll here if you are in the zoom please answer if you are in uh YouTube you can definitely um feel free to just drop in the chat but all right so we all watched the Super Bowl uh what do we think among these who do we think was the uh the winning Super Bowl ad right and this is in the minds of consumers what did they think all right just loost launched a poll please feel free to answer um got some answers coming in all right and we'll re revisit that as as we go through um so evaluating creative right so you know we're going to be sharing a lot of examples throughout this uh presentation of just you know different ads from the Super Bowl um so we're going to be sharing actual data from from surveys that that we conducted um but this is a sort of recipe and framework that we've developed for for assessing or or testing ads um and it boils down to a handful of survey metrics um so we start you know top of funnel by you know just fing a survey asking folks all right so exposing them to an ad uh in this case a a Super Bowl TV spot uh but you can this is very interchangeable for basically any type of media this could be done for for digital this can be done for print social you name it um but in this case the examples that we're going to be drawing from is going to be your your Super Bowl TV spots um but we start with you know top of funnel brand recall right does the ad you know create that association with the brand that that was intended um likeability to what extent did folks like or dislike and these are survey ratings uh with the exception of brand recall fivepoint scale you know pretty simple survey ratings uh engagement to what extent would folks actually you know sit through and watch this ad uh viral poent potential and that's the degree to which they might or or would expect to share it on on their socials um consideration How likely would you be to consider this brand or or thing uh if it were available to you uh intent your know likelihood to actually take action and either purchase or use the thing uh and then a few more valued measures here we've got likely recommend uh finding the content relevant and comparison to other brands um we like to show in this sort of one pager scorecard um you know the comparison right so I'm going to kind of break down you know the component parts of of this scorecard or these uh Norms as it were so all right this is what the scorecard typically typically looks like um we've got scores we've got some thumbs ups and downs which I'll talk through and we got a little bit of a video playback uh call it a dial testing measure right all right so top of the funnel scores that's you know the these top measures right here so these are your evaluative ratings that are most influenced by the given bit of creative that that we're assessing right so again the brand recall the likability engagement viral potential these are the things that are highly impacted by the the creative asset that is being tested right then and there when you get down to the bottom of the funnel and I will say uh before that even so recall you know does the brand associate with the ad that's your table Stakes right if you are an ad that is you know through the roof in terms of How It's scoring if folks are not making that assciation then that is potentially a problem the bottom of funnel metrics right so this is the your consideration your intent likely to recommend uh relevance and comparison to other brands these are much more difficult for a given piece of creative to move the needle on right and these are going to be more subject to you know an overall Brand's Impressions or likelihood to use or EAS or difficult of of signing up or long purchase Cycles um these are the things that you know for a given creative test you may not see a whole lot of movement with right I mean it is certainly the desire that you put an hat out there and it instantly you know elicits a response to go out and purchase the thing that you're advertising but you know that that's a tall order right I will say that so we're showing finished ads in this case um I I am a big fan of doing this type of research for um even under construction ads uh so for example you know agency is producing a handful of spots you know they're trying to decide among a few different potential creative Avenues it is a really great and it's gotten a lot easier with with some of our AI tools out there to create a bordom Matic animatic type of um you know not quite finished deliverable but something that that consumers can easy easily wrap their heads around I'm a huge fan of you bringing in research like this earlier in the pipeline right not necessarily testing a finished ad but actually getting some you know quick consumer input prior to you know spending time in effort on hiring talent in you know for some ads getting into a lot of post- production and CGI and things that can you drive up those budgets right and the last component of this particular scorecard is going to be video dial testing and that's basically just a playback of the video that has has just been uh evaluated and I don't know I'm a child of the 80s but I used to I grew up with the Atari and I had the the the controller that had the paddle which would I could dial back and forth that's basically dial testing if you've never experienced it or done it um in person it's actually something that used to be done and is still done a fair amount uh with in-person research so you bring in you know a group of people sequester them show them a video product a movie trailer a full you know 30 minutes of a of a a TV pilot and you give them a handheld dial which they dial to the left and they see something they dislike dial to the right when they see something that they like um in this case this is a digital variant of that uh where folks you know are given a video they have a a slider that they can move to the right and left when they see something they like or dislike within the ad um so this is the balance of you've got the evaluative measures the survey based and then you've got the diagnostic which is going to be your dial testing right so if you're doing print ads for example um obviously you can't do a dial test but I like to try and work in some kind of diagnostic measure for for those as well uh it could be a heat mapping it could be a you know highlighter tool for for a print or a static image um in some cases you know replace the D test Al together and you know a lot of Brands like to do uh you know something more neuro in nature maybe some eye tracking or some facial cating uh just something to get at you know more of the frame by frame right uh and how should you use or how can you use something like a dial testing exercise it is a great tool for if you are trying to identify maybe and the idea is you know and I'll show you some examples of this but you you have a video playback and you can watch the results of that that dial testing exercise you can watch the the Peaks and values of of of a line chart that as you know certain favorable things happen on screen um you'll see you know that line go up when some certain unfavorable things happen you'll see that line go down uh great applications for that would be you know if you are trying to I don't know you got a 60-second spot that you are tasked with trying to figure out what content is important for the 30 30 second variant or the 15 is a great way to do it it right when you see those you know those dips that's potentially area where you don't have to necessarily include in in that you know in that next version as it were um in certain cases you'll see you know I've seen dial tests where you know a certain character or or Talent appears on screen and the moment they do you you see that dial dip that could present a problem down the road D test is a great gut check for you know are there any offensive scenes in in your uh video are there any things that you maybe didn't foresee when you were you know creating the app right uh and then you'll see these thumbs ups and downs um and those are going to be indicators of uh normative uh improvements or or or declines so so we at Excel research we've done this recipe for for testing ads hundreds and hundreds of times so we have a database of norms for and benchmarks for ads that we've tested in the past uh whenever you do an ad test so you go out and test a singular ad and you get a likability score of 81% you know in a vacuum without anything to compare to that 81% means next to nothing right so it is always about being able to have some point of comparison so that you can tell what what is good what is bad um in this case norms um in many cases and I'm actually a big fan of actually not using a thirdparty norm like like we might provide uh but instead as a company if you're doing this type of research over and over again you know to whatever extent you can start to standardize it and create your own Norms that's fantastic right because you know even these measures that that we include in our sort of standard for doing this type of research they're not applic able to every brand you know there are some measures which may be a lot more appropriate to include when you are doing this yourself I often recommend you know customizing right don't necessarily be beholden to a set of norms just because either you've tested it that way for for years or because you've bought into you know a company like ours that that has their own recipe for Norms you know again if you're doing ad testing over and over over again it it's best to create your own honestly um okay so let's pretend that you you know we you're a Super Bowl Advertiser you have four brands that you have uh you know that you're running ads for and you want to test all of them so you know how can we you as a standalone we show just the individual scorecard how well did it perform plus or minus our Norm um but you know again maybe you've got four ads delivered by the agency you want to test them all and decide which should be your winners which should be your losers which ones you know are warrant further development which ones look like they may be you know back to the drawing board as it were right one pager simple scorecard reporting can work really well and so instead of you know just like I said showing the individual ad you can show and compare a few um simultaneously um in this case you know you've got nerds budlight Helman and Mountain Dew compared to one another you know which ones scored better or worse than the others um so in many cases you set out to do a bit of research uh maybe your goal is not really caring so much about you know what's happening within you know the rest of the universe or comparison to other ads that are being tested but it's literally trying to make a decision among you got four ads to test you're only going to you know take to production two or three of them how do you prioritize this is a great way to do that just quick comparison on those value evaluative measures uh and you're looking for significant differences right so which ones perform Better or Worse um and we just like I said these are actual results from this year's Super Bowl ads uh interesting if if we take a look at this and this is something I was talking about with regard to some of those those bottom of metrics but uh intent you know much higher for a helman's than for nerds Bud Light Mountain Dew that's not necessarily because the ad the When Harry Met Sally you know offshoot ad was that amazing no a lot of that is going to be by virtue of the the product category that they exist within right um you know Candy Beer soda a little bit different from you know uh a competitive context right so whenever you're evaluating this type of research you never just want to copy paste and push forward your your results you always want to take into consideration why things are happening what they're all about what it means right um okay so of the ads that were tested this year uh just if we take a look at our our valuative measures um you know who scored best who scored worst uh and what you know the the balance of this presentation is really going to be is just kind of talking through a bunch of the examples and showing those of the ads right all right so if we take brand recall and that's again you know did the ad you know did folks actually take away which brand it was associated with um who did it best who did it worst among the Super Bowl ads this year Doritos Ritz Pringles uh best right so they are among the highest brand Association and that's those the salty snacks did it this year right um the brand recall that brands that did so less goodly um would be open AI chat GPT uh the stop haate I believe it was um this was the Snoop Dogg Tom Brady ad um and then the rayan and meta uh these were among the lowest scorers on that brand recall but remember we talked about context right uh for the the don't hate stop haate I cannot remember the name of it exactly um there was no real brand associated with it so you know if if the ad didn't score well on that well there may be you know reasons why right so that's why I say you always want to take take an consideration you know the context of what it is that you're evaluating don't just blindly you know report results but instead you need to know if you're doing well if you're doing poorly why that is right all right so likeability um who did well who did poorly likeability uh lays the little girl planting her potato and watching it grow into a bag of potato chips really performed quite well in terms of just overall liking um also Budweiser in Clydesdales as well as booking.com and the Muppets uh not so great uh there was the coffee mate uh I don't know if yall recall this one and we'll we'll show it in a few but this was the uh the traveling tongue uh ad I don't know if you remember um there was the 2B uh head shaped like a cowboy hat not so great um and then there was Duncan actually really good and very well received ad last year um I guess we credit JLo for that but this is you know Duncan didn't do his great this year uh engagement so this is you know how likely would folks actually you know tune in and view this thing uh the performers again you're probably starting to see some patterns emge but you got your Budweiser you got LS again and you've got Michelob Ultra um pickle ball Hustlers um not so great or again the the coffee mate uh there was the door Dash um got trouble with your finances at I don't know if you remember this one uh that one didn't do great when it came to engagement uh nor did the stock uh cob Brew Coffee in spite of dreamy C chanting Channing Tatum um being part of it they just just couldn't overcome possibly the fact that it's you know European football being shown during American football uh I don't know but not not so great from that standpoint uh viral potential this is the one that is you know your tendency to share these ad results with with others post on your social as it were right so this is you know getting a sense for you know what kind of afterlife and and you know and potential you know sharing and ongoing exposure and Impressions might your ad have uh who did it best uh actually the ad that wasn't an ad Angel Soft and this was through the roof with folks that you know were were you know saying that they would share this on their socials um the ad itself didn't perform well on much else but again it was kind of an unconventional um got Snoop Dog and Tom Brady again High viral potential again that one didn't have great uh awareness but that that wasn't you know there wasn't the brand Association right uh again again with the bud wiser Clyde sales uh and again on the on the negative side you've got the coffee mate not doing well uh there was a hims and hers ad uh that wasn't as as well received or or folks weren't as willing to to share it virally um and again with the door Dash so a bunch of examples here A bunch that I'm going to kind of continue to to share out um and this is kind of you know the illustration of you know that dial test in action how it you know how it performs right uh this was the helman's When Harry Met Sally uh this was a really well-received one like I said this was a 60-minute spot and I wanted to kind of highlight the uh the dial testing because you see you know a bit of a dip in this as you get into the second 30 seconds of the ad um prob could cut that down to your your your 30 second or your 15 second spot and you wouldn't lose a whole lot of the content right uh you do see a big jump at the end of the ad when uh Sydney Sweeney comes on and says you know gives the iconic line of I'll have what she's having um but this was you know generally a very well-received ad right no reason to hate was was the name of this one this was Snoop Dogg and Tom Brady uh again you know the recall number is not so great but you know it was impactful right it had strong like ability engagement now when you get into some of these again you take a a concept like this which is just be nice to one another um there's not necessarily A A brand to purchase or use you know there's not likelihood to recommend sure I can do that but you know you always want to think about all right what's your product category who is your target audience you know are these measures appropriate or or applicable if they're not you want to maybe really start thinking about doing something with custom in nature right uh this was the meta and uh rayan uh I think you had uh yeah the evaluation of the the metag glasses right this one kind of didn't hit the mark quite as much um and the brand recall for this one was low and that's not uncommon when you see these um kind of paired up or collab type of type of brand assessments you know it's hard to know you know is this a meta product is this a Rayband product you know who who is the appropriate sponsor of this thing as it were right um I think we have another example here of that actually where you know we had a bunch of Brands coming together this was the uh instacart I believe um where it was actually done quite well and very well received um so you know not as as powerful connotation with regard to meta definitely so with with regard to the instacart so there was the battle this year for the uh traveling facial hair I don't know if youall recall this but we had we had Pringles and we also had uh Little Caesars where in this case Pringles you got the mustaches being ripped off pe people's faces and and flying through the air um Little Caesars I can go to that one uh you got kind of the same same deal but in this case it's it's the eyebrows uh both really strongly received ads um I think Little Caesars has the edge in terms of which one one um and then if you get into the following ad and that would be your coffee mate um so appropriate for eyebrows appropriate for mustaches when it gets into the tongue uh coming out of the mouth people are not so so keen on that right so you know pretty poor ratings with regard to some of your likability engagement um kind of across the board as it were right and you know you're you're spending this much on an ad you're definitely wanting something on the more favorable side I will say however you know that that's not all always the the case or the intent I remember GoDaddy a few years ago did a Super Bowl spot where it was uh supermodels and and Average Joe's uh making out and people were very kind of disgusted by the ad um there are other ads that you know negative impression but in some cases that's kind of the intent of the ad right um now I recall there's there's a toenail fungus ad where the character was just was very creepy and you know it always made your skin crawl but it's a memorable ad so you know is it doing the job right uh B2B ads um Super Bowl Mass audience um most of your B2B you you'll see for example I think we're looking at GoDaddy here um tendency and and the trend is for these to not perform quite as well uh versus consumer products right and this is again where if you are a B2B brand and you are you know and you could argue that you know like a GoDaddy or uh I believe Squarespace was another one that we tested didn't do so well honestly Squarespace um but if you're you know targeting small businesses for example well you know if you're doing a genp kind of General consumer ad like a Super Bowl uh you know small business make make up you know depending on your definition as much as you know 15 20% of the the general population so you know it's not not a crazy targeting but when you're comparing a a an ad targeted to small businesses to you know cookies or chips well that's going to be very difficult for those to measure up so when you're doing nor normative comparisons you always want to kind of keep that context in mind keep in mind what you're comparing against who you're up against and you know not hang your hat on or you know feel too discouraged by you know if you are a a niche type product and you're comparing to these these very gener General market uh this was another one of the really you know really big Winners and that was your your Budweiser Clydesdale horse walks into a bar um you know this is they've been doing variants of this spot for years and years and it it continues to have that impact uh another one was really well received that was your your M mob Ultra U this was the pickle ball Hustlers um really well received ad very favorable in terms of recall liability engagement viral potential mentioned this one this was the hims and hers uh weight loss uh ad this one's a good to highlight just because of the bottom of funnel right again you're you know this one didn't do Orly per se except on likability um and a lot of that comes from you know it's some pretty strong IM imagery that's used especially in the the first end of this this ad um but you know performing at Norm is not the worst thing in the world for for for a given and AD right um performing poorly is something that you certainly want to to avoid uh but again for a brand like him and hers you know again these Norms are going to be compared against batteries or or snack food you know a product that requires you know potential injections or or you know heavy side effects obviously that's going to be really difficult to to measure up when it comes to you know some of these more uh you know call to action measures like like like uh likelihood to use or purchase uh like recommendation all those good things another one that really missed the mark this year and this was the uh the door Dash uh let's talk about your finances just just didn't resonate um I think that was all the ads um questions let's let's get into them I can see that some folks have been chatting have been talking throughout uh would love to see if you guys have any questions um again comments definitely bring them on I'm more than happy to to talk through some of these debate about some of them um so let's go for poll results first though um so pretty even split um looks like you know about 25% of folks thought Budweiser mob Ultra Helman were the winners guess what uh depends right depends on which measure you're you're looking at or talking about um and in fact if you're a young researcher this is your answer to every question that you're ever asked about research and that is it depends um you know what was the intent of your campaign you are you awareness building are you trying to really move the needle on something like consideration um you know the the answers to some of those questions is going to vary and is going to you know Drive uh to what exent your ad is successful or or a failure do you produce High performing ads over and over and over again um and do you does one variant come out and you know it's it's you know performing below your standard right uh Colleen is asking a question and I'll go ahead and address that and that's uh you look at the why Behind These numbers um did those that elicit elicited humor and Nostalgia perform better than those that elicited disgust or discomfort um in this case for this year not so much um but that is and I will say that you know what we're sharing is very scor Cardy in nature right um meant to just prop up and build those Norms but that you know getting behind the what that we're you know just pushing out in the scorecard is absolutely what what you should be and what we should all be exploring as researchers so if you're standing up a program like this um in your organization that is the thing you want to understand to whatever extent possible right um so are there certain things that you you see over time you know the more and more of this research you do the more you can start to assess these kinds of things so humor Nostalgia um tugging at the heartstrings you know are these things that tend to resonate versus others um similar similarly again for from Colleen this is um how did your results vary by demographics or psychographics um something not mentioned but I think it's very very important about this is to do exactly that so everything that we've reported on is just genp you know census balanced uh demographics right but you know in every survey you should be and we are capturing demographic information you know capturing some you know brand usage some some category usage some psychographics this is information that you can then use to compare and contrast slice and dice your data maybe you have have four ads that you're looking to roll out you've got one among them that isn't a very strong performer among gen pop but if you dig into you know some of your subgroups um maybe you notice that um I don't know a particular ad is really well received among folks who have you know who lose subscriptions for example right um that could be a really compelling reason why especially if you're doing your media mix um trying to basically place that ad that maybe wasn't as high a per former as the others from an over overall standpoint um you know how much is it doing better and is it something that we can place um you know in a more targeted fashion right uh Caleb asked about the open AI ad results uh this was the the chat GPT um I don't think we have a scorecard in this particular presentation but didn't perform great um it was among the lowest in uh several of of the measures that we look at uh in fact the brand Association wasn't wasn't there uh and that could be by virtue of the is it open AI is it Chad GPT uh what is it right uh Matt Roberts is uh posting that the Philadelphia chapter of the AMA holds an annual Smackdown of advertisers uh use of the uh Super Bowl event as platforms to leverage uh reach and engage with Target audiences that'ss on the board um and we'll want to pull some of these metrics into their convention ah sweet yeah go for it please do uh everybody yeah I mean we WE Post these results on on YouTube um they're open and available for anybody to go in and access uh the ads that we test I mean these are Super Bowl ads these are you know published by the advertisers um on YouTube so all public domain we have any other comments questions coming in about wrapping this up kind of time we got got a few minutes um if there are any final questions definitely feel free to drop them uh otherwise we're going to be posting this uh this archived video for for folks to to access as they wish um and we're going to be uh posting again the slides uh for folks to access and download as they wish um otherwise appreciate y'all attending uh definitely answer or connect with one another feel free to connect with me on LinkedIn um definitely register for for our next uh our next webinar is coming up and uh we appreciate it y'all

Get Involved

Present at the next ARVIC

Share a method, a study, or a hard-won lesson with a room of senior research practitioners.