Skip to content
May 2021 ARVIC

Preference Testing in UX Research: Should You Do It?

Molly Malson of Charles Schwab makes the case that preference questions, as commonly used in usability testing, do not reliably correlate with actual usability outcomes. Drawing on psychology research and two internal case studies, she argues that cognitive biases make stated preferences a poor proxy for ease of use, effectiveness, or satisfaction. She closes with practical guidelines for teams that still need to include preference questions.

Key Takeaways

  • In both case studies presented, participants preferred a design option that scored lower on independent measures of ease of use, understanding, and satisfaction, suggesting stated preference does not predict usability.
  • Cognitive biases including the framing effect, choice blindness, the substitution effect, the aesthetic usability effect, and the decoy effect all distort design preferences in ways that are difficult to detect or control for.
  • If preference questions cannot be removed from a study, always collect independent usability measures (time on task, task success rate, SUS, CSAT) after each design option individually, and let those measures outweigh comparative preference results.
  • Asking participants to explain their preference choices can help analysts separate aesthetic or irrational drivers from usability-relevant reasoning, though participants often cannot fully articulate why they chose what they chose.
  • The Pepsi Challenge is a real-world cautionary tale: sip-preference tests produced results that did not hold up to actual consumption behavior, illustrating that preference data can mislead even outside the digital UX context.

Questions & Answers

The most common questions I receive from partners are 'Do they like it?' and 'Would they use it?' How do you respond when there isn't time to walk them through the full rationale?
Malson suggests asking the stakeholder whether they believe liking a design means it will be easier to use, and then prompting them to show data supporting that assumption. A short reframe: a lot of research shows the most reliable way to measure usability is to have people actually do things. She also notes there is not enormous harm in adding preference questions if you follow the guidelines and use independent measures, since the disparity between preference and usability outcomes can itself become a persuasive internal demonstration.
Should usability and aesthetic preference both still be tested, but kept completely separate? And should aesthetics be stripped from a usability study to remove noise, or does that reduce realism?
Malson's view is that established design patterns and accessibility guidelines already encode a lot of what is known about usable aesthetics. If an aesthetic issue is getting in the way of task completion it will surface in a task-based test. For conversion-focused questions about element placement, A/B testing is the appropriate tool. On emotional or brand-aesthetic impact she is less certain of the value, questioning whether measured differences translate into meaningful real-world outcomes. On fidelity, she recommends lower-fidelity testing early in the process to keep participants focused on interaction and flow, and higher-fidelity or coded-environment testing later to catch issues introduced during implementation.

Session Notes

About the Speaker

Molly Malson spent 12 years in technical writing before moving into user experience. She has worked across interaction design, information architecture, taxonomy, and research. For the past four years she has focused exclusively on UX research at Charles Schwab, primarily on digital and online experiences.

What Is Preference Testing?

In usability research, preference testing is most commonly defined as showing participants two or more design options and asking which one they like better. Common questions include which design a user prefers, which looks most trustworthy, and which appears easiest to use. Many online guides focus on asking users what seems more usable rather than observing actual use.

Malson's talk is specifically about traditional usability testing, defined as evaluating a product by having participants complete realistic tasks, and then using that format to ask preference questions when more than one design option is shown. The focus is on digital interfaces.

Where Preference Testing Comes From

  • Early use appears in the 1970s in animal behavior and welfare research.
  • Consumer sciences adopted taste tests and similar preference methods over subsequent decades.
  • When and how it entered UX research specifically is unclear; instructional articles on how to conduct preference tests simply appeared over time.

Malson's hypothesis for why it was adopted in UX is that preference questions have face validity: it seems logical that what users prefer will align with what is easier, more efficient, or more satisfying. She argues the evidence does not support that assumption.

Why Stated Preferences Are Unreliable: Key Cognitive Effects

Framing Effect

First documented by Tversky and Kahneman in 1981. People's responses are shaped by how information is presented, not just by the information itself.

  • Moxie, O'Connell, Mageddon, and Henry (2003): twice as many patients chose surgery when told they had a 90% chance of surviving versus a 10% chance of dying. A meta-analysis found patients were on average 1.5 times more likely to choose surgery when outcomes were framed positively.
  • Irwin and Gaith (1998): ground meat labeled '85% lean' was rated significantly more favorably than the identical product labeled '15% fat.' Participants also rated the lean-labeled sample as tasting better.
  • Pricing displays: showing an original price alongside a discounted price (with an anchor) drives higher conversion than a simple price, because buyers perceive they are getting a deal.

Choice Blindness

Discovered by Swedish researchers in the mid-2000s. When people's choices are secretly swapped for the option they did not pick, most do not notice and will confidently defend the substituted choice.

  • Faces study: only one quarter of participants noticed they were shown the face they had not selected; the rest defended the swapped choice, sometimes citing features that did not appear in the photo they had actually chosen.
  • Jam tasting study: fewer than 20% noticed the jams were switched; the rest explained why they preferred the one they had not picked.
  • Political opinions study: more than 90% accepted and endorsed at least one altered response. Participants who argued more strenuously for the opposite position were more likely to restate that view a week later.

Substitution Effect

Defined by Daniel Kahneman in Thinking, Fast and Slow (2011). When a question is hard to answer quickly, people unconsciously answer an easier related question instead. In design testing, 'which design do you like better?' may be answered as 'which has my favorite color?' or 'which looks shorter?'

Aesthetic Usability Effect

First observed by researchers at Hitachi in 1995. People tolerate minor usability problems when they find an interface visually appealing, and they rate perceived usability much more closely to perceived aesthetic appeal than to actual, observed usability problems. Nielsen Norman Group documented a case where a user experienced navigation problems but rated ease of use very highly, citing the calming colors and photographs.

Decoy Effect

Described by Huber, Payne, and Puto in 1981. Adding an asymmetrically dominated third option shifts people's preference between two otherwise comparable options. A National Geographic study (2013) on popcorn sizes illustrated this: adding a medium option priced at $6.50 shifted most buyers to the large bucket at $7.00, whereas without the decoy most chose the small at $3.00.

Case Studies: Preference vs. Independent Usability Measures

Study 1: Product Guidance Navigation (n = 129, within-subjects, 3 options)

  • Preference question: Option 1 was the clear winner.
  • Independent measures (ease of use, understanding, satisfaction rated after each option): Option 2 scored higher on average than Option 1.
  • Reasons given by those who preferred Option 1 but rated Option 2 higher were largely visual: icons, layout, an interactive behavior described as 'cooler.' These map onto the aesthetic usability effect and potentially the decoy effect (Option 3 was an existing design with known issues).

Study 2: Investment Projection Display, Line Chart vs. Alternatives (n = 160, within-subjects, 3 options)

  • Preference question: Option 1 (a line chart) was the significant winner.
  • Independent measures: Options 2 and 3 were rated significantly easier to understand. Option 3 inspired significantly higher confidence. Option 2 also showed higher (though non-significant) confidence and satisfaction scores.
  • On virtually every independent measure, the non-preferred options outperformed the preferred one.
  • Participants chose the chart format because of a pre-existing bias toward charts, not because they understood the content better when it was presented that way.
In both cases one design was rated higher on independent variables related to usability than what users preferred, providing support that preference doesn't correlate with usability.

Guidelines If You Must Include Preference Questions

  1. Ask participants to explain their choice. Carefully analyze those reasons to separate aesthetic or irrational drivers from usability-relevant feedback. Acknowledge that people often cannot fully articulate their reasoning.
  2. Measure additional dependent variables. Have participants complete tasks with each design and collect standard usability metrics: time on task (efficiency), task success rate and error rate (effectiveness), SUS or SEQ scores (overall usability and ease), and CSAT scores (satisfaction).
  3. Measure each option independently. Ask ease-of-use or satisfaction questions after each option, then compare results statistically. This typically requires unmoderated testing to achieve the sample sizes needed to detect significant differences.
  4. Let independent measures overrule comparative preference. Agree with your team in advance that if independent measures point in a different direction than the preference result, the independent measures take precedence. Define this in your research plan before fielding.

A Note on Preference Testing Beyond UX

Malson acknowledges that preference testing may have value in other applications such as market research, but raises a caution even there. The Pepsi Challenge taste test found people preferred Pepsi in sip tests because it is slightly sweeter, leading to a major ad campaign. The preference did not translate into sustained market share. Coke's response, a sweeter reformulation called New Coke, failed and was quickly withdrawn. The lesson: first-sip preference did not predict real consumption behavior or market outcomes.

Recommended Reading

  • Thinking, Fast and Slow by Daniel Kahneman
  • Blink by Malcolm Gladwell
  • Nudge by Richard Thaler and Cass Sunstein
  • Predictably Irrational by Dan Ariely

Transcript

Read the full transcript

okay we are at the top of the hour so i think we're gonna go ahead and get started with our next presentation um welcome back everybody it hasn't been long um so since we just wrapped a presentation i'm not going to go through our ground rules again other than the reminder bullet number one that we are at the mercy of tech because our next presenter molly malson from charles schwab was it was scheduled to be a presenter last month um the the internet internet connectivity gods decided not to smile on her that day and we ended up actually having to to cancel her presentation for last month but you know the silver lining is that that she has agreed to to come back and it look i can both see and hear her so i think we're looking pretty good in terms of uh being able to to kick off for a successful presentation this month um so molly welcome and thank you and i'm going to get out of the way and let you get started and do your thing well thank you bill but one we do have one issue i had to send you an email i just found out that our company apparently locks down sharing the screen and zoom if you can stop sharing i'll try but i did a test right before this so i might have to have you share my presentation while i speak not a problem i'll stop sharing and you just let me know yeah okay we are ready with plan b if need be well let's see i actually think it's working can you guys see a slide oh wonderful awesome everything's working thanks so much bill and i'm so glad that this all worked out today and thank you to everybody who's attending so today i'm going to talk to you about preference testing in the field of user experience research and whether or not you should do it but before that i'd just like to tell you a little bit about myself so i went to graduate school intending to work in the field of political communication but after my first job in dc uh thankfully in my opinion i ran away screaming from that field and i randomly landed in the field of technical writing after doing a job for a former teacher that turned out to be technical writing and so i did that for 12 years and then after a while i got sick of writing instructions for poorly designed interfaces and i wanted to fix them instead so i transitioned into the user experience field which i've been in for the past 16 years and during this time i've been somewhat of a jill of all trades i've done interaction design information architecture taxonomy research project management business analysis but about four years ago i really decided to solely focus on research because that's the thing that i i really like the most and feel the most passionate about so i've been at charles schwab for just about four years and in my experience i've worked in a lot of different industries and companies of many sizes but primarily i've been focused on digital online experiences so that's the focus of my talk today so what we'll be going over today i want to start out by talking about the definition of preference testing and then i want to talk a little bit about how valuable preference testing is as part of user experience research next we'll look at a couple of case studies and then last we'll go over some guidelines for this kind of testing so let's start with the definitions what is preference testing and so the way i most often hear people define a preference test when speaking of usability research is showing people some design options to participants and then asking them to choose their favorite and a usability hub post that i've referenced below lists these commonly used questions that they suggest in preference tests so questions like which design does a user prefer which design looks the most trustworthy which design looks the easiest to use and other online instructional guides focus on focus on the idea of asking users what seems more usable rather than actually doing some testing on usability another thing about these kinds of tests is they often focus on the visual or aesthetic appeal of a design now that things like fonts and colors and images are increasingly expected to express really complex brand traits like friend friendliness or trustworthiness and the norse nielsen norman group which is a prominent company in the field of user experience they have a an instructional article referencing how to do preference testing on these kinds of aesthetic impressions and the source for that is listed below but i wanted to clarify that that kind of testing is not going to be the focus of my talk today some people may debate whether that kind of research on aesthetics is truly usability testing i want to focus on what we define as traditional usability testing which is the evaluation and measurement of a product or service by having participants complete typical tasks with it and then using that format to ask questions about what users prefer or choose when showing more than one design option so in other words if you are a user experience researcher and you've been asked to or you want to evaluate user preference in some way as part of your research that's what i'm going to be talking about today and then a reminder the focus here is on digital projects and products and interfaces so this information may or may not apply to other product types so just want to be clear on what i focused on in my career so next let's get into how valuable preference testing is as a part of this kind of usability testing and research that i've been talking about and sometimes we do things because other people do them without really knowing why so i really wanted to just examine the value of doing this kind of work so before i go any further i just want to stop and take a pulse check on your thoughts about this topic so we're going to try doing a poll and we'll oh great so fast thank you bill um so this is the poll and here's a copy of the the actual question and if you can look at this question and just provide an answer so it's when sharing two or more designs do you think it's a good idea to ask usability test participants which design they like better so if you can just provide a quick answer i'll give about a minute or so so everybody can answer and bill if you see it all done before i interrupt let me know how's it looking for answers bill does it look like most people have finished the poll yeah it looks like we have a 70 response rate from our our attendees which i really want to figure out the secret sauce here because i want to apply that to my surveys that i conduct um yeah that's pretty good but we do have live attention so uh this is true this is true yeah i'll just wait like third like five more seconds if you haven't had a chance go for it otherwise we'll continue okay so you can stop the poll now bill and then just let me know what the answers are in the percentages i don't know if you can show it or if you just want to say it yeah i'll just read it off to you um we've got yes at 43 no 48 percent and i don't know at 10 percent okay awesome all right so thanks for that and then we'll get back to that later on in the presentation so i want to start with a little background about where preference testing came from and so i wasn't able to find a really good detailed history on the subject so i've pulled some information from a few places you might know more about this if so i'd love to hear from you after the session but here's some summary that i've come up with so in the early 70s apparently scientists use preference tests as a means of answering questions about animal welfare and then wikipedia also focuses on its earliest usage in annual animal behavior and motivation and then in consumer sciences and marketing preference tests such as taste tests have been around for decades and then i found it to be really unclear when it started being used in the user experience field i just know that i found articles talking about how to do it in usability testing so the next question is why do we do this so one reason why i think we have added this to our usability testing scripts is that asking people's preferences has face validity face validity means that it appears to measure what it intends to measure it's also called logical ability and so we assume that what people prefer will also align with other positive things we want to impact like an easier more efficient or more satisfying user experience which is the classic definition of usability that it's converting in some way improving sales or forming a favorable opinion of the brand or perhaps enhancing trust and i know my myself i love to ask people what their favorite this or that is and i like giving people choices so that they can pick the things that they like the best so it really seems logical that asking people what they like and what they want is is good and that if they pick one design over another that it would be a good idea to move forward with that one but how do we know that's true and so i started investigating this question over the past couple years and i look to the field of psychology and as it turns out research in the psychology field shows that we don't have a clear understanding of why people choose what they choose much less what that choice may or may not predict so as a spoiler alert and the rest of this presentation i want to argue that the existing evidence indicates that preference questions do not provide value to the digital product user experience evaluation process and actually may result in less usable experiences if we rely on that data so numerous books on the bestseller list you might recognize some of these here or maybe you have read some of them they talk about how people make choices and how typically irrational reflexive and fragile those choices tend to be so i want to walk through just a few examples of how fragile our choices really are and i want to start out with one that's probably the most commonly known effect on our choices called the framing effect and so framing is a cognitive bias where people's responses are influenced by the way the piece of information is presented to them and this is one of a general class uh several in a general class of judgment and decision making fallacies that have identified by researchers initially framing effects was demonstrated by taversky and kahneman in 1981. so i wanted to show you a couple examples of framing effects and how they work so one example is whether a piece of information is presented in a positive or a negative light and there was a study done by moxie o'connell mageddon and henry in 2003 where people were asked whether they would have a particular surgery operation and they were told one of these two things so one group was told you have a 90 chance of surviving the operation and another group was told you have a 10 chance of dying during the operation so obviously we see it's the same the same information just presented in a different way and these researchers found that twice as many patients would opt for surgery when shown the first one versus the second so the one that was positively framed initiated a different response than the other one and that same group of researchers did a meta-analysis of a bunch of studies in the same vein there are several that did things like this and they found that test participants were on average one and a half times more likely to choose surgery over other treatments when treatment efficacy was framed in positive terms so let's look at another example of the positive and negative framing because i just think this one is really fun this was a study by irwin and gaith in 1998 where people were shown two versions of this this ground meet and then asked to evaluate it and each group evaluated one of the two options just like before and you can see that they're exactly the same but the only difference is one has a stamp that says 85 percent lean and the other one has a stamp that says 15 fat so in this study they found that the 85 percent lean product was evaluated significantly more favorably than the one on the right and then not only that but the test participants then proceeded to sample the meat i hopefully not raw the way it's here and they chose the one that tasted better and guess what they chose you got it the 85 lean one so now let's look at another kind of frame and this is on value so what you'll see here is a classic example of ways of presenting cost of a product you can see the one on the left has a really simple price whereas the one on the right has an original price with a discount and you'll see too this option on the right i won't get into the another effect on our decision making but there's also an anchor here which you might have heard of and that's a cognitive bias where people's decisions are influenced by a particular reference point or an anchor that sets the value of something so in this case they've already established that this t-shirt costs a lot more than what you're going to pay for it so as you can guess the frame on the right leads to higher conversion purchases rates since people think that they are getting a deal and they see that they have a value here so that's the framing effect and now let's move on to another phenomenon that impacts people's choices and this is called choice blindness and this is a phenomenon where even when someone doesn't get what they want there's a strong chance that they won't even notice and they might even defend a choice that they think they made but they didn't and this phenomenon was uncovered by a group of swedish researchers in the mid 2000s so let's show some examples of that so one of the studies they did is that they showed people two faces not these ones they were this is just an example and they showed them two different faces and then asked people which to say which face they found more attractive and then a few seconds later they're shown the opposite one the one they didn't pick and then asked to explain why they picked that one so in other words they swapped out their choice and asked them to defend that choice and in the study only one quarter noticed that they had a different face than what they had selected and then the rest of them proceeded to defend their choices even sometimes with reasons that didn't appear in the picture they really selected like saying oh i liked her earrings when the other one didn't have earrings so that's one example let's look at another example of choice blindness and this is in sort of the taste testing realm where people were asked to taste two kinds of jam and then again they were a few seconds later they were given the other jam that they didn't pick and asked them to taste it again and say why they liked it so in this case less than 20 percent realized the jams were switched and then they proceeded to explain why they chose the other one so next let's look at an example that's a little bit heavier than jam or pictures and this was more in the field of political thought so that same group of swedish swedish researchers that's hard to say ask swedes during a general election about who they wanted to vote for and their opinion on a number of political issues and then just like in the other studies the researchers altered their answers so that they were actually picking the other point of view on these opinions and then asked people to justify their responses and in this case they found that more than 90 of people accepted and then endorsed at least one of the altered responses and then what was more surprising about this study is that they came back a week later to people to ask them about their positions again and they found that the more people argued for the opposite position so the more strenuously they supported the one they didn't choose the more likely they were to remember and then restate that view in a week okay so let's go on to another um effect that impacts choice and this is called the substitution effect this was another known issue with choice making that was observed and defined by daniel kahneman in thinking fast and slow in 2011 and the effect says that if we can't come up with a choice quickly we find a related question that's easier to answer and we answer that instead the receipt the researchers defined this effect by comparing answers to multiple substitute questions that they hypothesized for questions that weren't easy to answer and then they found ones that strongly correlated on various psychological measures so that they knew they were getting the right substitute question so here's some examples of that one example is when they ask someone how happy are you with your life these days they were actually answering what is my mood right now so you might be wondering what does this have to do with preference questions well it's easy to see how this might happen when a person is asked to choose between designs so if we ask them which design they like better and it's not a fast or easy choice could they be answering questions like which one has my favorite color or which one has an image i react to more favorably or which one looks shorter and this will kind of come into play in the next choice or the next effect effect on choice where we might be seeing some evidence of this substitution effect and that's called the aesthetic usability effect and this is that people are more tolerant of minor usability issues when they find an interface visually appealing and so that means that a user could choose a design that's less effective in terms of task completion just based on something that's that's related to the aesthetics of the design and this effect was observed by researchers at hitachi in 1995 and how they did that is they conducted experiments to see what the relationship was between apparent usability which is what people like say is usable versus inherent usability which is the observed number of problems with the experience and they found during this analysis that the apparent usability is is far more correlated with the apparent beauty in other words what people think is aesthetically pleasing than it is with the inherent usability an example of this was a test done by the same nielsen norman group that i mentioned before and a user experience they watched a user experience several usability issues from minor annoyances to serious navigation problems but in a post-task questionnaire she rated her experience in terms of ease of use very highly commenting that it's the colors they use looks like the ocean it's very calm very good photographs so next is the last example i'll give you of how our choices can be fragile although there are a lot of others that we could go through these are the ones i felt might be the most relevant to choosing designs and this is called the decoy effect this was first described by huber payne and puto in 1981 and this is a phenomenon in which people change their preference between two options when presented with a third option and that third option is considered to be the decoy or it's much less attractive than the other two options so that means that it is asymmetrically dominated and i'll show you an example of this that you've probably seen before so national geographic did a study on popcorn sizes in 2013 don't ask me why national geographic was studying popcorn but apparently they were and uh they started out with two groups and they gave them two offers either the small popcorn for three dollars or the large popcorn for seven dollars and in this case most people choose a small option saying that's all they needed and that was a good price but then a decoy was added and i can see where this has led to pricing in theaters and you see the medium is not five dollars as you think it would be but it's 650. and so in this case most people chose the large bucket since they saw more value and having that much more popcorn for only 50 cents and you can also see some of the other effects that we've discussed coming into play here like anchoring you know showing what the cost or value is of something and framing the way you present the information so in conclusion from all of that research given all the ways that our choices can be easily influenced or manipulated or really not just not very rational i ask why do we want to measure those in our usability tests because as i mentioned before usability researchers researchers support user efficiency effectiveness and satisfaction and we already have valid and reliable ways of measuring usability that just don't have anything to do with people's choices so we don't need this extra type of questioning if it doesn't help and so i started focusing on this topic as i mentioned before first because i've gotten requests for this type of research often in my career and because there's all these articles on how to do it online for usability researchers and i used to at the beginning of my career add these kinds of questions without putting too much thought into it and start i started diving into some of this psychology research the past few years and questioning why we really do this so in my in some of my studies i've been experimenting with adding the preference questions to studies along with other measures and then seeing how the responses compare to see whether they do have any relationship to the the things that we're trying to measure so that's what i want to do next is to share a couple of case studies for you so the two in these two examples both of them were designed to be as close as possible to the sample size you would need to detect at least a 10 difference between options if one exists and i wanted to have people evaluate the options both independently and based on usability principles and then also comparatively so we had both sets of data so for the example number one we had three design options that were presented for guiding users to products that match their needs and it was a within subject study with 129 participants and within subjects if you're not familiar with that means that each participant saw all the versions in the same study and then the results on this is when asked to choose the one they preferred option one as you can see here was the clear winner over the other two options however when i measured ease of use understanding and satisfaction after each option what i found was that option two was rated on average more highly than option one even though they chose option one yeah that was it and so i wondered why why was that different why did they why did they rate option two hot more highly and chose the other one so the first thing i did is i looked more closely at the data set to see if more investing available showed the disparity and then they more strongly preferred option one so that bumped that preference up more and then i looked at the reasons given for that set who preferred option one but rated two more highly and what i found is that the reasons were largely visual like having icons they liked having icons i liked the overall layout they liked a particular interactive behavior that it had they said it was cooler than the other one and so here you can see where some of those decision making effects that i talked about earlier could have a play so we could have some type of framing involved where one type of layout or interaction versus the other was more positive to them you have the aesthetic usability effect where people are talking about aesthetical things that are over overwhelming their choice um there could be a decoy effect here we had three options and although that number three wasn't specifically meant to be a decoy it was the existing design that we knew had some issues so it could have had a play there as well and then let's take a look at another study in this one again we had three design options presented for displaying projected income from an investment over a given time frame and we were challenging the assumption that a line chart was the best way to do this again this was a within subject study and we had 160 participants with this one and then when we asked to choose which one they preferred option one was the clear winner in here because of the data type i was able to show the confidence intervals so you can see that this option one significantly was chosen over the other two again we found when asked for each option to independently evaluate on understanding confidence satisfaction options two and three were rated significantly easier to understand option three inspired significantly higher levels of confidence in option one option two is also higher but didn't show a significant difference and then option three had non-significant higher levels of satisfaction so on almost every measure on every measure besides preference the other options outperformed option one so again i looked at why the difference and i looked at the segmentation on the data on key variables i couldn't find anything that seemed to break the pattern and so again i looked at the reasons why they chose the one over the other and here i found that the choices were largely based on their preference for data being shown like this in a chart format so they just have a bias towards seeing things in chart but they didn't understand the content as well when it was presented that way so again here we can see possibly several effects coming into play the framing the the type of interaction that's presented aesthetic usability effect um could have played out here as well and we also had kind of a decoy in here it wasn't meant to be like for sure the worst one but we thought it was probably not as good as the other two so in conclusion this is only two studies so i wouldn't cut i wouldn't call that solid proof but in both cases one design was rated higher on independent variables related to usability than what users preferred providing support that preference doesn't correlate with usability and maybe future studies will find positive correlations with maybe with usability or with other outcomes that would be considered desirable so that's an area of research opportunity so maybe there's higher conversion rates and then that'd be kind of a question because we'd wonder why something less usable was converting in a different way but there's definitely some opportunity to do more research here so i've shared a lot of information for you about why i've concluded it's not a good idea to include preference questions and usability testing so i wanted to go back to that poll question we asked at the beginning and ask the same one again and just see whether any of this has changed your mind so no pressure you know if you said yes before and you still think that go ahead and answer the same if you've changed it provide another answer and i'll wait another minute this time for everybody to answer the poll how are we looking bill oh looking good we've already got over two-thirds of our attendees voting everybody's quick on the draw here okay well let's let's hear the results and move on then all right we've got the yeses at 33 and the nose at 67 percent zero i don't know is this time great well that uh that makes me happy i've i've changed the minds but i'm really interested to hear from the people who still think yes afterwards because that will be a great conversation in the q a so before i go to that part i just wanted to provide some guidelines so if you're still if i haven't convinced you not to include these in your tests or if you're unable to like win an internal battle and you still have to do it here are some guidelines to make it work the best for you so the first one is just asking for people to tell you why they made the choice so just ask them to explain and then carefully analyze those results to separate responses that don't relate to anything that truly impact usability and that likely come for from one or more of the decision-making effects we discussed so that you have that information and just a little warning as i mentioned people don't always know why they make a choice so they might not be able to explain it properly but this is a better than not doing it at all second guideline is to measure other dependent variables so not just asking which ones users like or prefer but making sure that you do tasks and actually do tasks with with the designs and measure other dependent variables that you want to impact with your design options and some of the common ones in usability is like we measure efficiency by time on task how long it takes for someone to complete a task we measure effectiveness by the task success rate or the error rate we measure overall usability and ease of use with things like the sus score and the seq score and we measure satisfaction with things like the csat and so guideline number three measure the options independently so rather than say measuring ease of use by asking users to choose which one was easier than the other you want to ask that single ease question after each option and then compare the results to look for significant differences and a warning that this typically requires unmoderated testing for you to actually find significant differences so that really doesn't work as well with um you know qualitative testing and then the last guideline i'd like to propose is to make sure that independent measures overrule the the comparative measure so if comparative results either show some significant difference between the options that the independent measures don't or if they show a significant difference that is not as the same as a significant difference found in the independent measures then the independent measures should really be the deciding factor and this is something that you should make sure your team agrees on beforehand so if you've planned the research properly you've identified the dependent variables that you want to measure and you've agreed on those so those should be the most impactful results um so in summary the definition we covered the definition of preference testing how valuable it is case studies and preference testing guidelines and one last thing i want to mention i don't know how reliable or valid preference testing is in relation to many other applications like in market research it may turn out to be really valuable in those uses although one story that gives me pause in the taste test realm is the pepsi challenge taste test and how that worked out and so for those of you who don't remember or just need a refresher shoppers and taste tests were asked to to compare a taste of pepsi to a taste of coke and coke was by far the market leader at the time you still are and they found all these people picking pepsi even people who say they love coke and so there was a huge ad campaign around this which was initially very popular and they found that while there was an initial jump in sales it it didn't last and they they didn't take over the market as they were hoping for and so when looking into why researchers found that because pepsi was slightly sweeter it had the first sip preference so people initially thought this is the better one but then when they drank a whole serving they didn't like it as much because it was too sweet so this also led to the new coke debacle if you remember that where coke responded by creating a sweeter version of coke and that tanked in the market and was hastily yanked from the shells so that makes me wonder whether this is as valuable in taste testing even as it might be in as i'm saying it is not in the field of user experience testing so that's my presentation thanks for hanging in there with me and now i'd like to open it up to any questions or comments alright questions for from molly um and as those are rolling in also anybody from team yes on the poll want to log your your answers as to why you chose that um i was actually getting ready to to give you my why sometimes yes molly but you you kind of summed it up as in your in your conclusions you know great for tiebreaker um not to be relied on sold me um let's see we've got uh it looks like amy has given a question uh the most common questions i receive from partners are do they like it would they use it um all right do they like it would they use it how do you respond to that especially when there isn't time for me to pull out your wonderful presentation and walk them through why i don't want to be asking those questions amy that's a good question um i think the short answer is i would ask them do you do you believe that if someone likes that that one's going to be easier to use and then if they'd say whatever they'd say then you could say can you can you show me some data on that because i could tell you a lot of research that demonstrates that the best way to measure usability is to have people actually doing things so i think maybe something short like that might might be somewhat effective and you know i i think i kind of counsel people and even myself if you really have people pushing for it then i think there's not a gigantic harm to adding it if you do use the guidelines that i provided at the end where you can just frame it and then you know i've shown my my groups that's this disparity and people saying what they like and then what works better and we all agree we'd rather have it work better so maybe it's a good idea to try it and just kind of show that kind of thing maybe you don't get the same results i do but it's at least if they if they if you're measuring independent variables and they match up then at least you have measured what usability is supposed to measure uh looks like rebecca has two questions which we will allow um if possible should usability and aesthetic preference both still be tested but perhaps completely separately and then should aesthetics look feel colors photos be stripped from a usability study as much as possible or does that add to the reality of the experience yeah both really good questions um i i tend to feel that we have a lot of data on what are usable and not usable aesthetic elements so what you know we have a lot of information on what colors are accessible which thing which things are accessible and usable and if you're applying those patterns and it really comes down to you know one over the other my thought is if you really find it's important then put it in your test plan a lot of times i i tend to focus on can people do the thing because if there are aesthetic issues that are getting in the way of them doing a task then we'll see them in a test and if you're if you're following just good design patterns and good design principles and making things accessible those those things shouldn't matter quite as much but if it's kind of like you know is this button over here over here going to result in more conversion that's what we have a b testing for so i think that's perfectly valid for things like that from the terms of aesthetic in terms of getting at people's emotions emotional impact for for aesthetic kinds of things i i don't have as much experience with that honestly and i i would be somewhat questionable as to how again why does that matter so if someone says oh this this color makes me feel more trusting than this one does it really play out in any real scenario that helps us enough you know is the difference significant enough so that's probably remains to be answered by more research i personally tend to kind of go away from that as far as how high fidelity the test can be i think that really depends on what questions you're trying to answer so early stages i think it's great to keep all that out so you just don't have to have people talking about oh i don't like that blue or things like that you're really focusing on the wireframes and the interaction and the process and as you get to higher fidelity you also want to make sure that you know the way you've applied this particular pattern and the way it's styled is actually working for people so it still goes into usability and then some we've all seen where things get into coding and they've totally changed from the way that they they were designed so i always think it's a good idea to try to get testing in the code environment as well if you have the kind of place where you have designers and then you have coders and they're separate so i think there's places for having lower and higher fidelity depending on where you are in the process and what questions you're trying to answer does that help rebecca i guess she can't directly answer but well we've got a thank you exclamation mark i'm going to take that as a thumbs up we had an early request during your presentation to uh list the three books that you alluded to early on list the books that what uh you mentioned three books early in your presentation if you could list what were they what those were yeah if you don't have those at hand we can share those in the archive well there's four of them but i i really really love thinking fast and slow by daniel kahneman that's about that's a big compendium on a lot of these different choice phenomenons and just just a whole ton of about how people think that's really really valuable for anybody in user experience and a couple of the other ones were anything by malcolm gladwell is great but blink was one that he wrote and that's that's where i got the example about the cope pepsi um richard thaler and cass sunstein's nudge is talking about how people make decisions and how you can possibly help them make good decisions if that's what you think is a good thing to do and then the last one was predict predictably irrational by dan ariely these are all on the new york times best seller list and that one is about um just the ways we make decisions and how irrational they can be all really good books awesome thank you for that i'm just going to do one more round of checking for questions on all of our platforms last chance for anyone to submit i think everybody's satisfied molly really appreciate it um thank you so much any any parting shots um i just i'm just thankful that you all came to listen to this and i hope i've changed your mind about how people make decisions i know i used to think i was a rational decision maker until i read some of the stuff and i realized oh yeah i've done that i do that and so i think it's really eye opening when we when we really we also have a tendency to think we make better choices than other people which is really not true and so i hope you all think about how you make choices and how our users make choices as you move forward in your jobs indeed i think the lesson is job security for us researchers because uh people are crazy irrational decision making that's for sure all right um molly thank you very much uh everyone who attended thank you so much um really appreciate it uh be on the lookout for archives of today's presentations being posted and uh be on the lookout for schedule for upcoming events um if anyone is interested in presenting definitely feel free to to reach out to me submit a form send a smoke signal whatever you like thanks a lot everybody and have a good one

Get Involved

Present at the next ARVIC

Share a method, a study, or a hard-won lesson with a room of senior research practitioners.