Skip to content
ARVIC · Virtual Insights Conference

February 2025 ARVIC

February 11, 20251 Session

The February 2025 ARVIC conference focused on building rigorous, context-aware ad testing programs that translate creative measurement into clear, actionable decisions for insights and product teams.

Overview

The February 2025 ARVIC session centered on creative effectiveness measurement, using one of the highest-stakes advertising events of the year as a practical teaching tool. The talk moved from foundational framework to applied technique, giving attendees a concrete model for running ad tests that inform real decisions rather than just produce scores.

Framework First: Structure Your Ad Testing Program

Before interpreting any result, a well-structured testing program needs clear objectives, consistent methodology, and a plan for building comparable data over time. The session walked through how to design that foundation using surveys across roughly 45 to 50 ads and more than 12,000 completes as a working example.

Match Metrics to Funnel Stage

One of the sharpest distinctions in the session was between what creative can actually move and what it cannot. Top-of-funnel metrics are where creative earns its keep; bottom-of-funnel metrics are shaped more by category dynamics, brand equity, and purchase cycles.

  • Top of funnel (recall, likeability, engagement, viral potential): responsive to creative execution.
  • Bottom of funnel (consideration, intent, recommendation): driven by factors outside any single ad's control.
  • Brand linkage is non-negotiable at every stage. A high-performing ad that fails to connect to the brand is a problem, not a win.

Context Is Not Optional

Benchmarking without context misleads. B2B ads, niche categories, and cause-driven creative will naturally score differently than mass consumer spots. Applying snack food or beer norms to those categories produces false negatives. The practical fix is to build internal norms over time using a standardized approach, while treating third-party benchmarks as a starting point rather than a standard.

Dial Testing as a Diagnostic Tool

Moment-by-moment response data has specific, practical uses: identifying what to cut when editing for shorter formats, flagging content that may land badly with a subset of viewers, and pinpointing when talent or characters help or hurt audience engagement. It is a diagnostic instrument, not a replacement for overall scoring.

Probe the Why, Not Just the Score

Consumer responses are not always rational, and that is useful information. A 20-point swing in admitted fan allegiance following a game outcome illustrates why scores alone are insufficient. Understanding the reasoning behind a response is as important as the response itself, and any testing program should build in mechanisms to surface that reasoning.

Get Involved

Present at the next ARVIC

Share a method, a study, or a hard-won lesson with a room of senior research practitioners.