Skip to content
March 2023 ARVIC

Does the Respondent Task or Number of Attributes Affect MaxDiff Results?

Researchers from an MMR graduate program ran a controlled experiment to test whether two methodological choices in MaxDiff design, the respondent task type (reorder vs. best/worst selection) and the number of attributes shown (three vs. five), produce meaningfully different results. Across two product categories, flights and dining, the top and bottom ranked attributes remained stable regardless of design condition, but mid-ranked attributes shifted depending on how the question was asked and how many options were shown. The core practical message is consistency: pick a design and stick with it, especially when tracking over time.

Key Takeaways

  • End-cap attributes (most and least important) are robust across design conditions. Mid-ranked attributes are where method choices create measurable variation.
  • Task type matters: best/worst selection and reorder interfaces produce different rank orders for middle attributes, even when the same items are shown.
  • Attribute count matters: showing three vs. five attributes shifts mid-tier rankings, partly because probability of selection changes with the number of options.
  • No evidence that gender differences explain variation in attribute order or preference levels across conditions.
  • There is no universally 'better' design. The key is to choose one approach and apply it consistently, especially for tracking studies where changes over time should reflect actual preference shifts, not method changes.
  • MaxDiff is scale-invariant, making it useful for cross-cultural and segmentation work, particularly when hierarchical Bayesian estimation is used to extract individual-level utilities.

Questions & Answers

Is AYTM a free service and how accessible is it?
The presenter described AYTM (Ask Your Target Market) as a full-featured survey and analytics platform with MaxDiff and conjoint capabilities, its own panel of over 40 million people, and automated analysis outputs. The program has a partnership that includes full license and panel access. The presenter did not know the current public pricing or whether a free tier exists, and recommended contacting AYTM directly.
Do you think the device used to take the survey (computer vs. mobile) has any impact? Is it better to use three attributes for mobile and five for desktop?
The student presenter noted that AYTM optimizes for mobile, so the interface is viewable on smaller screens. The team did not have device-type data in this study but acknowledged it would be a useful variable to add. The lead researcher noted that identifying which condition is 'better' requires a benchmark or standard of comparison. Without one, you can observe differences by device but cannot conclude which produces more accurate utilities.
Was there anything that surprised you while conducting the research?
The team expected to find gender differences in preference order but found none. That was noted as a notable finding. The lead researcher also highlighted that the three-way interaction across task type, attribute count, and industry was not significant, meaning the overall number of attributes in the category did not appear to moderate the results.
Is there anything you would have done differently?
The team offered three reflections: adding covariates such as device type or operating system to test for additional sources of variation; conducting qualitative research before the study to better select attributes that more clearly differentiate preferences; and exploring a wider range of attribute counts (e.g., seven or ten attributes presented multiple times) to understand how set size scales.

Session Notes

Research Question and Motivation

The study originated from a classroom question: does it matter whether you show three or five attributes in a MaxDiff? The team expanded this into a structured experiment, motivated also by the rapid growth in MaxDiff use and the increasing accessibility of do-it-yourself research platforms. As more practitioners without formal training run MaxDiff studies, understanding the method's sensitivities becomes more important.

Experimental Design

The study used a 2x2x2 mixed design with four total experimental conditions and 800 participants, 200 per condition. Participants were balanced on gender, age range, and region. Two factors were manipulated between subjects: respondent task type and number of attributes. Industry (flights vs. dining) was a within-subject factor, with order randomized.

The Two Factors

  • Task type: Respondents either reordered attributes by dragging them up or down, or selected the single best and single worst attribute from the set shown.
  • Number of attributes: Respondents saw either three or five attributes per MaxDiff question.

The Four Conditions

  1. Reorder task, three attributes
  2. Reorder task, five attributes
  3. Best/worst selection, three attributes
  4. Best/worst selection, five attributes

Industries and Screening

Two categories were used: booking a flight and dining at a sit-down restaurant. Respondents who did not dine out at least once a month, or did not travel by plane at least once a year, were screened out. The survey was programmed and fielded using AYTM.

Analysis Approach

AYTM output utilities for each attribute within each industry. The team converted these utilities using a likelihood function to produce choice probabilities and indices, which were then used for significance testing. A repeated measures analysis of variance was used to test for interactions.

Key Findings

Effect of Task Type (Reorder vs. Best/Worst)

There is a significant interaction effect between task type and attribute order. When comparing reorder to best/worst conditions, the rank order of attributes in the middle of the preference scale shifts. Attributes at the top and bottom of the ranking remained consistent across conditions. This pattern held for both the flights and the dining industry.

Effect of Number of Attributes (Three vs. Five)

Showing three versus five attributes also produced variation in mid-ranked attribute order. For example, the second most important attribute in the three-attribute condition was ranked third when five attributes were shown. Absolute index scores were higher in the three-attribute condition, which is expected: with fewer competing options, the probability of any given attribute being selected is higher.

Stability at the Extremes

The most important and least important attributes were stable across all four conditions. The variation was concentrated in the middle of the ranking. A plausible explanation is that respondents hold stronger, clearer preferences at the extremes, making those positions less sensitive to how the question is framed.

No Gender Differences Found

The team tested for differences by gender and found no significant differences in preference order or attribute levels across male and female respondents.

No Three-Way Interaction Across Industries

An initial repeated measures ANOVA found no three-way interaction between task type, attribute count, and industry. The overall number of attributes in a category did not appear to moderate the pattern of results.

Implications for Practice

  • Be aware that the task format and the number of attributes shown can change the rank order of mid-tier attributes.
  • There is no evidence that one design is universally more accurate than another.
  • Consistency is critical. If you are tracking preferences over time, changing the task format or attribute count introduces a confound that makes it impossible to know whether shifts reflect real preference change or method variation.
  • Attribute ranks at the ends of the scale are reliable regardless of design choice. Mid-tier ranks should be interpreted with more caution.

Broader Context: Why MaxDiff

The lead researcher noted several reasons MaxDiff deserves continued methodological scrutiny. It forces respondents into trade-offs, which produces richer preference data than simple attribute ratings. Unlike rating scales, MaxDiff is scale-invariant, meaning results are not distorted by cultural differences in scale use, making cross-country comparisons more defensible. When hierarchical Bayesian estimation is applied, individual-level utilities can be extracted for each respondent, enabling preference-based segmentation. The technique sits between simple attribute ratings and full conjoint in terms of complexity and respondent burden.

"Don't build a wall just so you can use your power drill. Make sure you have a wall that needs to have a hole in it."

Limitations

  • AYTM is an online panel, limiting participation to those with technology access.
  • The survey window was short, which may have introduced availability bias.
  • Fielding in January 2023 may have influenced preferences for dining out and air travel in ways specific to that time period.

Future Research Directions

  • A longitudinal study to test whether the same patterns hold over time and whether a clearer pattern emerges for which specific mid-tier positions are most affected.
  • Testing with more polarizing or niche topics where respondents hold stronger opinions, to see whether stronger preferences reduce mid-tier instability.
  • Adding device type (mobile vs. desktop) as a covariate to assess whether interface differences moderate results.
  • Incorporating qualitative research beforehand to better calibrate attribute selection and ensure the attribute set spans a wider range of preference strength.

Transcript

Read the full transcript

All right. And today we will be discussing does the respondent task or number of attributes affect the results of a max diff. Uh your presenter is Dr. Kuna as well as some MMR candidates. So I will let you guys take it away and I will enable the share screen. So you should be able to share your screen now. Can you see our screen? I can. You're good. Is that in the way the window? Okay, cool. So, we're good to get started. You are good to go. All right. Thank you, Amber. And thanks, Bill, for uh inviting us. This has become a tradition. Now is the third year that we're presenting uh this interesting research on researching uh research on research projects which you have present on scales and also whether collecting the data uh just in one day or multiple days affects the results of your research. So this year we're going to be talking about max diff. Uh here's the team that work with me uh on this project, our MMR students. And today, Jackie and Jod are going to be leading uh the presentation. For those of you who are not familiar with the MMR program, so if you want to further your career in marketing research, that's going to be a great option for you. Our program was founded in 1979 and it's considered a standard uh in the industry and we are the premier program that develops the talent for both client and supplier u insights and research uh jobs. So I will uh now hand the presentation over to Jodi and uh Jackie. Thank you. Awesome. Thanks so much. All right. So how do we even get to this research question of does the respondent task or the number of attributes affect the results of a maxiff? So Dr. Huna was teaching a class on applications and marketing research and he was introducing maxiffs to us and I raised my hand and I was like wait a second does it matter if I put in three attributes or five attributes for a maxiff and he says you know what that make a great research on research presentation so here we are um and while I'm here I'd also like to give a shout out to AYTM for being a a great platform partner to us um so all the research was programmed and fielded using AYTM All right, so let's get into the nitty-gritty details of the experimental design. So it is a 2x twox two mix design. So it has a mix of both um manatic and sequential monatic um aspects to it. So there's four total experimental conditions. We'll go into that a little bit more on the following slides, but there are basically two ranking types. So either the respondent's task is to reorder the attributes that they're shown or select the best or the worst of the attributes. Um and then there's either they're shown three attributes or five attributes. So again, this will make a little bit more sense when we show some examples. Um but this is basically how we set it up. And then we have two different industries um that they were shown. So that part is sequential manatic meaning that they saw both industries. So one industry is air travel um or flights and then the other one is dining at a restaurant. So talking a little bit more about our sample and randomization in terms of the sample we had a total of 800 participants um and 200 of within each experimental condition. So once the respondents were qualified into our survey they were randomly assigned to one of the four um experimental conditions um and then uh the order of the max diff was randomized. So whether they saw the max stiff on the flights or the air travel or whether they saw the one about dining at restaurants first that aspect was randomized. All right. So here are the four- exp experimental conditions that I mentioned earlier. So in our first condition um the respondents that qualified and came into this saw um either flights or restaurants first for their maxdiff and they either had to do the task of reordering um the attributes based on their preference for either three attributes or five attributes. Um and then in the second condition um it was similar task of reordering except this time they saw it with five attributes. Going on to our third condition, um it was they had to select the best or worst um within the maxiff that they saw and they saw this with the three attributes. And then lastly, our fourth condition was they selected the best or worst with five attributes. Um and something else to keep in mind that the respondents were balanced based on gender, age range, and region. All right, so what did our survey actually look like in AYTM? So before the respondents were qualified into our survey, they were asked to pre-quall or screen our questions. So the first question was on average, how frequently do you go out to eat for a sitdown dinner at a restaurant? And for the users that selected um I don't dine out or I dine out less than once a month. They were um kicked out of the survey at that point. And then the second pre-qualification question asked about in a typical year, how many times do you travel domestically by plane? And for those people who selected I don't travel via an airplane or I travel via an airplane less than once a year, they were also kicked out of the survey at that point. And then the actual maxive questions, what did they look like? So um the question regarding dining at restaurants that kind of appeared as when dining out what features are most or least important in influencing the overall experience and then when the respondents saw that they either saw that again with the three or five attributes and then either their task was to reorder it or select best or worst. Um and then coming to what that looked like for our air travel um or flight industry um type of the maxiff for that it said when booking a flight what factors are most or least impactful when deciding which flight to choose. So let's jump into what that actually looked like um on the respondents end. So for the first condition when they saw those three attributes and they saw um uh they had to reorder it. This is kind of what it looked like. So um for the three options they had the option to like select it and drag it to be able to reorder. Um so that's what those dots on the side um represent and then it says that um instruction of drag up or down to reorder. And then for our second condition very similar interface except this time it has those five attributes. And then coming to our third condition is they had to select the best or worst. So obviously this looks a little bit different than the reordering task, but here basically when the respondent hovers over to select what their best is. So they uh indicate that thumbs up with that. And once they have selected their best um then I'll leave the other options to be able for you to select that as a thumbs down to indicate your worst. So um these examples for are all for condition three in terms of uh the three attributes shown and the interface will look exactly the same for those who qualified into condition four except they would have five options to select their best or worst from all right now that we've made it through all of this. I'm going to move it over to Jackie to tell you what we actually found from experiment. All right, so getting into the data just quick background AYTM after running the max diff outputs the utilities for each of the attributes within each of the industries, the flights and the restaurants. So we took those those utilities and we plugged them into the likelihood function to get probabilities of choosing each and then converted to indices as well so we could see how they compared against each other. And that was what we ended up working with for our significance testing. So first we wanted to look with interactions between the selection style and the attribute. And just for visual representation this is sorted in descending order by the best worst category showing three options. The first is the highest index for proximity lowest being value. But what we wanted to test was was there a difference based off of best worse or reorder. So when we bring in the reorder, you can start to see there is some interaction in the pattern as they go along. Um and that's represented in these middle options here. Service, menu, size, ambiance. These numbers here are the rankings of the order that they're presented in. Again, sorted by the best, worst option with showing three. And we saw that this was true for both showing three and five. There is an interaction effect in that the order difference throughout um the attributes in their order. And we also did this for the flight industry and saw the same effects. And you can see these start to present themselves in the middle but um not really any indication as to where in the middle just not on the end caps. So the interactions are shown when the crossovers in the lines here as well as um the third option is ranked as the fifth option here and so on. So there is a significant effect on the selection style being best worst or reorder on the attribute order. We also wanted to look at is the number of attributes causing an an interaction effect on the attribute order. And so just again sorting in descending order for best worst with three options. We also added in you know the best worst five options. And you can see there is some variation in the order in the center as well. So the second most important attribute being company was the third most attribute most important attribute when showing five different categories. And same is true for three and five within the reorder section. Um, you'll probably notice that all of the options for the best worst with three options shown are a little bit higher, which makes sense because when only showing three options, the probability of selecting one of these attributes is higher amongst only two other options versus when showing four other options in the case of being five total attributes. And again, we saw that this was the case within the flight industry questions as well. So there is some variation in the order throughout these these in between attributes um as the order goes on for their index score. So to summarize the attribute ranks at the end points tend to be in the same score. They aren't changing depending on whether we're showing a rank order or a best worst. And the in between attributes is where we really see some variation. And this might be due to the fact that those NCAA ones people feel more strongly about, but because people don't have as strong of preferences in the internal and in between options, um they're more likely to be affected by the way that they are asked the question. How could this impact you in your work? Really, you just need to be aware that these results can vary based on how you're asking the question, how many options you're showing. Um there's no one best way to do it. There was no indication as to five options being better than three or it's better to show a reorder versus a best worse option. But it is important that you be consistent whichever one you choose to go with in case the um attribute levels orders would change over time based off of your how you're asking the question versus what the actual preferences are. Some limitations we did factor in when doing our analysis. There might have been some sample bias. One in that AYM is an online panel, so only accessible to those that have technology, but more so that we only had our survey open for a short window of time. So some availability bias and that if it was only open for a few short days, that might have limited some of the respondents. as well as we fielded our survey in January 2023. So this might have impacted people's preferences on dining out or if people are ch traveling more in January, how they feel about certain attributes. So going off that, if we were to continue with this research, we think it would be really interesting to make it into a longitudinal study and see how these trends happen to change over time. We noted that all of the changes happened in the center variables, not the end cap. So not the first most important or the least important attribute, but just in the middle. And while there was no distinct pattern that we could clearly state now, it might something might emerge over time and that um the attributes only change in the very middle or only right after the least important, the most important. Um and then we also think it would be really interesting to look at more niche or polarizing topics for max stiff study. So if there is a topic that people feel more strongly about and have um stronger opinions then how would that impact the results in the order of each attribute level. So that is all we have. Does anyone have any questions? So uh before uh before questions I just wanted to add that um the reason why uh we decide also to focus on max diff is because there's you know really really a great amount of interest in this technique. There's a lot of growth of use of max div. So it's important that we develop our knowledge on uh research on research uh so people can understand what are the potential uh pitfalls and caveats uh from these different techniques or the different type of surveys that you're using. So we'll continue to do this kind of a research. Also the fact that technology is becoming widely available to uh many people that perhaps do not have the adequate training to understand uh maxd. So platforms like AYTM that allows for do-it-yourself maxd and they do a great job of doing that. So it's important that people get educated uh on the process and also you know educated on the implications of your interpretations when you are considering doing a max div which is a much more accessible technique that takes into consideration the idea of tradeoffs which is a lot more complex uh when you do a a conjoin right but has the great advantage of being a scale invariant right So relative to when you ask people to rate attributes right the scales could be uh could vary as a function of different factors. So for example some countries might be generally speaking more optimistic than others. So if you ask the same people from different cultures to uh rate attributes on a one to seven scale, you might get you know means that are very different because they the ratings and the scales used they are scale variant. They will depend on the characteristics of the respondents. uh max div tends to be a technique that you know kind of u standardize uh those ratings and then you can look at the data more globally also in that case and also it's great that with the development of hierarchical basent max diff you can now get individual utilities for everybody that takes your survey. What does that mean? You know that mean like we saw like the aggregate utilities here for um different attributes for restaurants and flights but we also we could if we look into the data granularly I could see that the utilities for Jod in terms of flying are very different than those for um Jackie right so they might overall you know we know the utility of the attributes but per respondent those utilities might vary greatly So and what's the implication of that? Well, that's a great opportunity for segmentation, right? So we can create now segments based on the individual utilities if you use a hierarchical basian max technique from which you can extract individual uh utilities. So it's a great technique I think better for the most you know I'm not saying like don't use attribute ratings but you know in some instances is might be better than attribute ratings because force people to into some sort of trade-offs and we just discussed some basic max diffs. They're like anchor max diffs and variations of max div that improve the estimation of the utilities and uh but also it might not be as time consuming from the respondent stand point standpoint or resource cons consuming from the people who are paying for the study when you compare to a conjooint conjoint study. So great technique, add that to your uh toolkit or marketing research. But as we try to teach the students here, uh don't build a wall just so you can use your power drill, right? Just make sure you have a wall that needs to uh have a hole in it. So you you use the tools that you uh you have available. So I think um with that, we'd like to open for uh questions from the from the audience. Okay, I see some coming in chat. Someone was just asking for a little bit further information on AYTM. I know you were kind of starting to get into that a little bit, Dr. Nuna, but like is this a free service? Um what does it offer? How accessible is it, etc. Okay, so uh we have a partnership with AYTM that stands for uh ask your target market. So aytm.com they are a great platform to develop your survey max diff uh conjoint you have like the same capabilities you'd get from Qualrix I think one thing that they differentiate themselves relative to Qualrix and they do a very good job at the automation uh on the part of the analysis so they do like you know immediately when you have your um uh data collected you can get a you know a great amount of analysis whether it's a conjoint ax or or just uh you know frequencies uh and they also allow you to export the data of course to SPSS and Excel uh and they have their own panel and it's it's a great partnership with our program because we not only have the access to their full license, but we also have access to their panel. Prior to our partnership with AY and they have a panel, I think of over 40 million people on their panel. Prior to the our partnership with AYTM, we used to have students would have to ask their friends, you know, post on social media to uh you know uh complete their survey so they could complete their project. Now with our partnership with am we can expect specified demographics and they also it starts giving you a sense they say like hey with this many questions and this many participants your cost per interview is going to be $5. So also students start learning and you're going to have your survey by 5:00 pm today. You make the survey more complex, you will adjust that make it more simple, you adjust that estimate. That also help our students to learn about the tradeoffs that they going to be facing uh in their careers. Uh I have no idea of what their pricing uh is and I don't know if they have a free uh premium version of their product. Uh but um you can reach out to them. They're great at customer service and the person that work with us with their edu person Morgan uh she's great working with our students. So it's a it's another great tool uh for you to to take a look at and if it works for you uh I would suggest you you sign up for it. Thank you. Someone else asked do you think the device used to take the survey has any impact i.e. a computer versus a mobile. Is it better to use three attributes for a mobile and five for a desktop? So I think I don't know if you guys want to answer I can start answering which um I think so I think my perspective on this would be that because AYTM already has it like optimized for mobile devices as well like it's at the point where you can see all the options um in your screen in your frame itself. So when you're actually previewing and testing it, you can see how that looks um on the mobile device uh versus like a computer. So personally, I don't see that there would be um that much of an impact between the two. Um but I don't know what you want to add. I think I think when when they we get their data, do we get that kind of I don't think we get what they take it on. They can we can probably get that. That would be a great question. We could we could add a variable to our analysis and see if there is differences. What you can tell is if there is differences, it's practically impossible to say what's better, right? So to figure out what's better, we need to have a standard of comparison. What's the ideal? So for example, if if you know something is an error, right? Then we would um we would see uh oh yeah, there are more errors uh when you do three attributes versus five. But without extend of comparison or benchmark, it's hard to say which one is better. Perhaps you could see something that you know uh you know that for flights people would prefer uh prices most people would prefer prices to be lower than higher. Okay now if you see a higher proportion of people rating higher prices as more preferable for three attributes than for five then perhaps there is you might make the inference that because this can be assumed to be an error but there might be people that prefer pay more so they can go on first class that three versus five three or five is better based on that. the the important thing would be to define what's the extent of comparison against which three or five are being compared. So we then we can make a claim about uh which one leads to fewer errors or better estimation of utilities. Okay. And I just had a personal question. I always like to ask this when anyone's conducting research. Was there anything that surprised you while you were conducting the research? We also wanted to see if there were any differences across gender, and it turned out there there wasn't, which we thought there might have been, at least in the order, but there really wasn't any significant differences between males and females on preference, order, or um the levels of which things came out based off of how we showed it. So, I think that was an interesting finding. Is there anything else? I think I would say that also a question that also I mean I'm throwing this curveball with you guys. I always ask you guys in our classes after you do a project now that you've done the project and you learned through it the you know through the conducting the project is there anything you would have done differently? I I mean based off this last question would have maybe added in a few other questions to evaluate covariance and see if there were other factors we could test on. You know are you a Mac user versus an HP user? Is there any effect there on the results and maybe other things that we could have looked at? Um so doing more but making more work for myself. Totally. Um uh I know that we had done like a pre-ervey before we even started this to be able to identify which topic we should use and which attributes we should use. I think it would be interesting to go back and before we started this to have like looked at the different attributes that we put in maybe do some qual beforehand um to identify if there were any other attributes that would have maybe helped determine um more like more attributes on either points. So I know we had like our first and last pretty much stayed the same, but if we had um more researched attributes and that was they were more like preferred um by our respondents than if that would change the number um of attributes that kind of stayed at those end points. So I think that would have been interesting to see. Yeah, we also tried to do two industries that were to some extent different and also they had different numbers of attributes because we also we can also see like having 10 attributes presented you know five times versus seven attributes presented five times does that really matter? And then our initial analysis which we did a repeated measure analysis of variance there is no three-way interaction between the two factors of interest and the industry. So it doesn't seem like the overall number of attributes seems to have a impact on the overall wording of the attributes as a function of how they were they were presented. All righty. Well, thank you guys so much for your time and sharing us all these cool findings with us. We appreciate it. Yeah. No, thank you. And um I would also like to tell you guys that Jod and Jackie are going to be joining the industry. Uh as soon as they graduate, Jod is going to be with Merc Pharmaceutical and Jack is going to be with Osam Zelman. Uh they do the Zmat uh analysis in Pittsburgh. Congratulations for researchers. We love that. All right. Thanks for the opportunity once again. Thanks, Bill and Amber for supporting on this and we'll start working on the research on research for next year. All right. Thank you. You guys have a great day. All right. Awesome. Thanks. Bye.

Get Involved

Present at the next ARVIC

Share a method, a study, or a hard-won lesson with a room of senior research practitioners.