Trevor, a UX researcher with nearly 20 years of experience, walks through how to implement PURE (Pragmatic Usability Rating by Experts) audits in real organizations, drawing on eight years of hands-on application across fintech, tech, education, and government. He covers roles, top-task analysis, live scoring, and his expanded five-point scale, while Brett Kjuski of Accelerant adds context from scaling the method at Walmart. The session emphasizes PURE as a practical way to quantify usability data, benchmark designs over time, and communicate research value to stakeholders.
Key Takeaways
- PURE converts qualitative usability findings into scored, graphable data that stakeholders can track over time, making it far more actionable than a traditional heuristic review.
- Evaluators should never score their own product. Use three or more trained UX adjudicators separate from the lead designer or product owner to remove bias and increase credibility.
- A top-task list (10 to 15 key tasks broken into interaction steps) is the essential foundation. It can also be reused for moderated usability tests, surveys, and other research activities.
- An expanded five-point scale adds a critical 'user cannot proceed' score (5) and a clean 'no friction' score (1) that the original three-point scale lacks, making it better suited to complex interfaces.
- PURE creates strong second-order effects: shared UX vocabulary across teams, built-in cross-project collaboration, a structured critique process that removes personal bias, and an upskilling mechanism for early-career designers.
- Versioning PURE reports (v.1, v.2, etc.) makes benchmarking explicit and enables ROI tracking by correlating usability scores with metrics like NPS or adapted SUS scores over time.
Questions & Answers
- Top tasks are big in our org, but they often feel disconnected from UX research on the ground. Does PURE help bring them together?
- Yes. The top-task list exercise is the foundation of a PURE audit and directly connects organizational priorities to the research activity. Once built, the same list can be reused for moderated usability tests, unmoderated studies, and surveys, so it becomes a shared reference point rather than a disconnected planning artifact.
- Can you walk through the scoring in a more concrete way?
- The scoring spreadsheet has columns for task number, step number, page or wireframe title, and the action taken. Each adjudicator enters a score (1 to 5 on the expanded scale) and a short note. The spreadsheet calculates the aggregate color and score automatically. The final score for the full audit reflects the worst single step. Adjudicators converge after the SME leaves and must agree on each score rather than averaging.
- Are the one-through-fourteen heuristics shared in the chat?
- Yes, Trevor shared a link to his rubric document in the Zoom chat during the session. The document includes the four criteria he added plus Nielsen's original ten heuristics, each with a plain-language explanation written in his own words. Free templates and scoring spreadsheets are also available through links in his Substack articles.
Session Notes
What Is PURE?
PURE stands for Pragmatic Usability Rating by Experts. It was created by Christian Rohrer and colleagues in the fintech space and has since been adopted across organizations including Capital One, Walmart, Stanford, and the City of Raleigh. The method quantifies qualitative usability data by having trained UX experts score a product's interaction steps against a set of usability heuristics, producing graphs and benchmark scores that stakeholders can understand and track.
Trevor has been applying and extending the method for about eight years. His approach diverges from Christian Rohrer's original papers in several practical ways, particularly around scoring scale and team roles.
Why PURE Over a Traditional Heuristic Review?
- Traditional heuristic reviews often produce a static document that sits unused. PURE turns the same evaluation into a number and a graph that can be benchmarked over time.
- PURE is faster. Trevor reported 6x to 8x time reduction compared to a conventional heuristic review, even accounting for having four people instead of one.
- Results speak directly to stakeholders in formats they already recognize, similar to a Rotten Tomatoes score.
- It triangulates with other metrics such as SUS, NPS, and analytics, giving teams a more complete picture of usability.
Step 1: Assign Roles
Two role types are needed for a PURE audit.
- Subject Matter Expert (SME): Usually the lead UX person, product manager, or stakeholder on that product. They guide the session but do not score their own work.
- Trained UX Adjudicators: Three or more UX professionals who score the experience. At least one should have prior PURE experience. Non-UX professionals have been tested in this role and do not produce reliable results.
Keeping the SME out of the scoring removes bias, increases credibility with stakeholders, and opens the door to including engineers, product managers, or other internal colleagues as observers who then become more bought into the research.
Step 2: Build the Top-Task List
Before scoring begins, the SME works with the facilitator to define the top-task list: the 10 to 15 most common or most impactful tasks users perform. Each task is then broken down into individual interaction steps along the happy path.
- Tasks should 'criss-cross' the interface rather than all following parallel paths, so the audit covers a representative range of interaction patterns.
- The happy path is an acknowledged bias of the method. PURE is not a substitute for real user testing, but it is an effective step before a prototype or launch.
- The top-task list doubles as a foundation for moderated usability tests, unmoderated studies, and surveys, so the work done here pays dividends across other research activities.
- Versioning the list (v.1, v.2) makes it explicit that the audit is designed to be repeated and benchmarked.
Step 3: Live Scoring Session
The session runs over a video call such as Zoom. The SME shares their screen and walks through each task step by step. Adjudicators score each step independently and in real time using a spreadsheet.
- Adjudicators score in a vacuum, one step at a time. They are evaluating from UX best-practice knowledge, not roleplaying a persona.
- Adjudicators can ask the SME questions during the walkthrough, which is valuable when the interface is a paper prototype or early-stage design.
- After all steps are scored, the SME leaves and the adjudicators converge to agree on a final score for each step. This is not an average. Every adjudicator must reach agreement.
- The best-argued position for each score wins. Disagreements almost always fall between adjacent scores (e.g., 2 vs. 3), never between extremes.
Scoring Scale: Three-Point vs. Five-Point
Christian Rohrer's original system uses a three-point scale: (1) accomplished easily, (2) notable cognitive load, (3) difficult for the target user. Trevor recommends this for simpler interfaces.
For complex interfaces, Trevor developed an expanded five-point scale:
- 1 - No meaningful difficulty. Little to no friction. No heuristic violations.
- 2 - Minor issues. One or two impactful violations.
- 3 - Moderate issues. Three to four violations.
- 4 - Significant issues. Five or more violations or multiple reasons a single heuristic is breached.
- 5 - Critical error. Users would be unable to proceed past this point.
The five-point scale has been validated as statistically as reliable as the three-point system. The critical score (5) fills a gap that the three-point system leaves when a step is clearly beyond a score of 3 but has no higher option.
How the Aggregate Score Works
Scores and colors are treated as two separate but related dimensions. The overarching score for any full PURE review is defined by the worst single step in the entire experience, not the average. This reflects the real-world reality that one critical failure can derail an entire user journey.
Trevor's spreadsheet tool calculates this automatically and generates color-coded graphs (green, yellow, orange) that match the score levels. Raw adjudicator notes are stored alongside scores so stakeholders or designers can drill into any specific rating to understand the reasoning.
The Scoring Rubric
Trevor's rubric is based on Jakob Nielsen's 10 usability heuristics, with four additional criteria he has added from practice. The rubric is embedded in the scoring spreadsheet as hover tooltips, so adjudicators can reference it during live scoring without needing to memorize it.
Accessibility was intentionally excluded from the rubric. At the fidelity level of a PURE audit, accessibility evaluation devolves into QA-style checks (ARIA labels, semantic markup) that belong in a dedicated accessibility review process, not a heuristic-level audit.
Converting Scores to Percentages and Benchmarking
Trevor converts raw scores to percentages to make them easier to communicate and benchmark. Organizations can set a target, for example a goal that all tasks score above 75% of the maximum possible score. Repeated PURE audits at different versions of the product then show movement toward or away from that goal, enabling correlation with business metrics like NPS or SUS.
The aggregate number at the end is meaningless. It's only meaningful to itself. But if you're able to figure out some kind of benchmark, then you can start showing the correlation of the work that you're doing.
Second-Order Effects of Adopting PURE
- Shared vocabulary: Teams align on UX terminology and heuristics, reducing inconsistency between senior and early-career practitioners.
- Stronger first drafts: Designers know their work will be evaluated against the heuristics, so they apply them earlier in the design process.
- Structured critique: Feedback is grounded in an agreed rubric rather than personal preference, reducing unproductive disagreement in design reviews.
- Cross-team collaboration: Distributed teams gain a structured, recurring format for cross-project collaboration without unstructured critique meetings.
- Stakeholder credibility: Usability is presented in numbers and graphs rather than expert opinion, making it easier to justify design decisions to executives.
- ROI measurement: Benchmarked scores make it possible to calculate the impact of design changes over time, something difficult to achieve with activity metrics alone.
Practical Tips for Implementation
- Start scoring by gut feel, then retroactively map the score to the heuristics. Most experienced evaluators' instincts are accurate, and the mapping exercise reinforces heuristic knowledge.
- Include at least one person with prior PURE experience in the adjudicator panel to anchor the converge stage.
- Run PURE before high-fidelity design, before usability testing, and before any beta or launch. It works on paper prototypes.
- Templates and scoring spreadsheets are available for free through Trevor's Substack articles. No paywall.
- Agencies can package PURE as a productized service, as Trevor's former digital agency did.
Transcript
Read the full transcript
Awesome. Appreciate it. Uh Brett Kjuski, uh head of research and growth here at Accelerant. Um pleasure to have Trevor joining me. We'll be talking a little bit about Pure. Um for those that are on the line, yes, we are recording these webinars. Uh they will also be uploaded to YouTube. Um so in the future, you're able to, you know, search for Accelerant or even go to the resources page that Bill just flashed up. Uh and you'll be able to go to past conferences all the way back from 2020 till the one that we're hosting today. Uh but without further ado, uh wanted to introduce our expert that we've brought in. You know, uh the guy that's helped create Pure Method and also take Pure Method to that next step. Um those that aren't familiar, you know, prior to my time at Walmart, I wasn't exposed to Pure. Um but you know, how how great of a methodology it is and how practical and and usable it is for agile teams and and quick spaces. But Trevor, I'll let you do a quick intro about who you are and then we can kind of jump into what you have prepared today. All right. Thanks. Yeah. Thanks, Bill. Thanks, Brett. Can you hear me? Okay. Everything's good. All right. Great. Well, first of all, I just want to say thanks to uh you, Bill, and the rest of the team for uh you know, having me on today. Um, Pure is a uh it's a really powerful way to quantify qualitative data and I it's I've literally seen it transform many orgs and uh I've been kind of evangelizing it for the last I don't know seven eight years around places and the more I get into it, the more I dive deep into it, the more I realize how it's it's so needed right now and it has all these fun, you know, secondary effects as well. So I I love talking about it. I'm actually currently uh helping implement it across a large organization uh like in a UX research ops capacity. So that's like literally what I do during the day. So I I'd love to kind of talk about it today a little bit too. So um did you want me to share screen? Do you how do you want to do it? Let's let's do it. Yeah, I know you have those slides prepared. So I think flashing that stuff up. Um pure is is a newer concept to most people in the UX space or even in the research space. So, I think uh some visuals might might aid in helping those kind of follow along and and understand. All right, great. Yeah. So, I didn't really prepare much of a slide deck. It's just got a couple of uh uh key points here and I thought we could kind of just riff and talk. So, feel free. This is more of a conversation. Um I'm way happy to talk about how to implement this in in the real world uh in your own organizations. Um, I've been doing it, like I said, for a long time and uh I'd love to just, you know, I I I was looking at the attendee uh names as they came in via Zoom and I actually recognize some of these names here. So, uh, but I don't recognize many of you. So, if you have any questions about this, feel free to ask. I'll literally, you know, give you my take on how I would approach doing this in your own org. But first, let's just talk a little bit about it. And Brett, feel free to jump in whenever. So, oh, this is outdated. I'm at 19 plus years now. I gotta update my thing. So, uh, my name's Trevor. I've been doing UX forever. Uh, I've been doing it for almost I'm going on I'll be going on 20 years after this year. And, um, I started off as a UX research slash uh, designer, but then maybe the last about 14 15 years I've been just doing UX research. So, I've been around, you know, my first job titles didn't have UX in the job title. And I've been around through all uh, many different kind of ups and downs throughout the uh, the profession. So, I I come at it from a slightly different perspective than some people, but mostly just been there, done that kind of stuff. Um, I have most of my experience in-house teams, but this is also good for I I did a a two-year stint at a digital agency in which I was like the head UX strategist, and we we actually bottle Pure is so it's so like um like a modular thing that we were able to bottle that up as a product and sell it. So if there's people here at digital agencies or research agencies or or things like that, we can talk about that. Um but mostly I've I've seen the effects happen at scale within larger like orgs. Oh, and I also write a weekly blog. So check me out on Substack. There's the QR code. Um it is here. I'll show you. I got it. This is I do funny illustrations because the content is so dry. So, these are like funny retro illustrations, but it's hardcore deep dive 2,000 3,000word uh uh articles about um advanced UX research, including I have a bunch for about pure um so I can I can put those in the in the uh chat as well. So, I'll just keep going here. Interrupt me if you want, but so just what I was going to say that is one of our offerings, right? So, we we are verse in pure method. Um, so you know, we talked offline and and obviously connected over, you know, geeking out about methodologies and one of the things was one of our first conversations was pure method. Um, and you know, bringing that to Accelerant, but also, you know, my time at Walmart and and how we scaled that with a three-step method of heristics and to pure into usability and how quick and agile that can be. Um, but obviously focusing on pure method today and that is something that we do offer to many clients here at Excel. Nice. Yeah. How does that go? Like Yeah, I mean it really depends, right? It's uh kind of like what we talked about offline of you know how we pair pure with usability. So bringing in participants to do the pure method with a usability kind of task orientation. Um, so it could be remote IDIs where we're having them go through more of a guided session and having those pure method or that pure scale u on the back end or even you know it could be more feasible where it's more like that traditional pier where you're having that panel of three to five experts um and and having our our clients start learning how to do that methodology implement that methodology. It gives them that base um just like base marking um of just you know where their design or where their product is right now um and where they can focus in. Uh that's been a big thing over the past let's say two years of just like how do we just find where we're at in par in the in the ecosystem of products or designs or all those things right features um and how can we then measure that to start correlating uh outcomes from the outputs that we're doing and pure is a great agile and useful like you said it's qualitative at stat sig or not at stat sig but it's qualitative research in a quantitative uh digestible way where uh it gives them some kind of benchmark to be able to move forward and and show the improvements of of their whatever they're producing, right? Yeah. It's like a lot of teams that I talk with and including the the current one I'm helping right now, they talk to me about like, well, we do a pretty good job with satisfaction, you know, they have some kind of like sus thing going on, some kind of benchmarking in that way. Uh we have a good sense of, you know, time on task, you know, because they have analytics, but like what about usability? What about like a vast majority of what we you know the things we bring to the table uh um you know when we're differentiating ourselves as user experience professionals against product management type uh the the the one place where we really really can shine is underusability and this covers all the rest of that and I think of it as a really nice way to triangulate so yeah pure for those that you don't know it's an acronym like all the crazy stuff in our field right pragmatic usability rating by experts. Uh this was done uh Christian Roar and some colleagues um in the fintech space uh kind of came up with this idea. They evangelized it. Uh they wrote some papers about it and then it's slowly been dripping out into the real world. Um so I didn't know you you did this at Walmart. So I have a slide with just like a random set of organizations. I didn't put Walmart on there. I was going to add that last night uh as we went through it. Um yeah, but it was actually uh a colleague of mine from Capital One worked with Christian as well. Um he brought it from Capital One to Walmart and that was kind of his uh Devon Singh who I don't think is on today. Um but he he was the one that kind of set up that three-step method. So going into heruristics then going into pure then going into usability and obviously you can fall out at any point. Um, so it more made it uh an approachable way to interact with research, right? That's been a major issue I've heard from our clients being an in-house, being an agency, you know, research takes too long or, you know, we're a cognitive machine versus this was a scalable way to say like you can fall out of that process at any time. Um, but it also pushed evangelizing research within a a Fortune global one company. Yeah, exactly. Yeah, it's awesome. So, so I'm going to speak a lot today. I've taken what Christian has published. Um, and then I I have had a a few side conversations with Christian himself about these things. Um, but I' I've taken it even further. Um, and uh, so I'd love to talk to you about how to actually do this, like how to make it happen in the real world. That's the whole idea here. Um, so you get these kind of graphs at the bottom. So imagine a world in which you can, you know, take a a standard heruristic evaluation or heristic review and boil it down to graphs that your stakeholders see all the time. That's what this does. It it quantifies qualitative data and I think it's really specifically usability data and I think it's really really powerful. Um, and it's been proven to be as reliable as like the stuff that people already or more than u the stuff that people already kind of like adopt. So, you know, I we'll just go past this, but if we want to talk further about the reliability about I' I'd be happy to talk a little deeper about it. And then just these this is just a random group. It's not comprehensive of orgs that use this. Um this the reason why I knew these is I either know the people at the company that are running pure or I've literally implemented it myself at these places and I try to pick a random variety. So there's fintech it's pretty big in fintech but there's fintech there's you know just tech there's an education home science tools is educational um organization Stanford of course um uh let's see where there's a Raleigh. Oh that's that's the city of Raleigh. That's where that is. So, governments use it. So, it it really is something that when adopted can really kind of transform orgs and it's got it's starting to get more and more popular from what I've been seeing. Um, let's just dive right into it. So, the process for for those of you who don't know is you need Oh. Oh, by the way, I'm going to derivate from any of the like Christian Roar writings and just say how I do it. So, this is like how it works in the real world. Is that cool? If anyone has any questions, you can hit me up in the chat. Oh, I also just put in the chat all the all my writings on pure if you want to really dive in deep um on some of my thoughts. So this is how how do you make this this method this process how do you actually do it within your orgs? So what I did was I broke it down into these steps. So step one I've got this idea that you have a subject matter expert which is usually your product lead UXer on any individual product and then trained UX adjudicator. the the uh pragmatic us uh uh usability rating by experts. They're the experts. And the reason why you don't want your lead uh UXer on the project, by the way, this is for larger teams, but you can scale this out. It can be done uh it can be done differently for smaller teams. I can talk to you about that. Um you know, go to uh different places uh like accelerant that do this. So, um, so yeah, you get your lead UXer, they should not be scoring their own thing. That's the beauty here. And I'm going to explain a little later why that has all these great, you know, second order effects. But so, you know, if you're the lead designer and you have enough staff or enough trained people that can understand this or you have a great partner um, like Accelerant that can help with these kind of activities, you get three or more uh, trained adjudicators to adjudicate the thing through the system. I I liked your point there, Trevor, about like don't don't be doing it yourself, right? If if if it's your product or you're the one designing or researching this, you know, there's that removal. Um there's obviously bias. You also know the flows, you know, the task and how to get the success. And yes, this is a heavily guided methodology. um where that that lead UXR that's to me is is very much hands-on guiding those trained UX uh you know experts in the process um and it is very leading at times you might say right um but it it's you know that that separation of church and state right so when we would implement this or when we do implement this with clients um you know we try to tap into a diversity of uh you know people that come in as that that expert, right? And put on their hat of whatever group or persona that they might need to put on in order to act out the the task that that we're going through. So, um a lot of the time we recommend to clients both internal or external, uh don't be the ones that are evaluating your own. It also helps you get buy in, right? So getting some different um diversity of like engineers or product or design folks within the uh larger ecosystem of your company can show kind of what this methodology can do. Can also show the output where they're a little bit more bought into the the power of research. Yeah. So it's you're not scoring your own thing. So it's not the same person they see every day in their standups coming back and saying, "Hey, look, we did this thing. Trust me, blah blah blah." Another thing I put like uh subject matter expert there in uh in parenthesis because I've explored not theme role in this way being the lead UXer. It could be the PM. It could be a stakeholder. It could be literally anyone um that can go through the process and and the process as which I've defined it for uh expedience. Um so that's kind of a neat interesting thing. We I actually have tried training other non-UXers on the train adjudicator part, but it doesn't work. It's just like that's just our profession. It's what we do. It's it's the E of Pure. It's the expert part. Um, but we've played around with all, like I said, I've been doing it for about eight years in all different orgs and, uh, we've tried many, many different flavors. And this is the one that's worked time and time again for multi-billion dollar uh, uh, tech conglomerates all the way down to some small startups uh, where I've done this, including the uh, the firm that I talked about. Um, step two, this I have a little secret um, a little secret. Do like a a task analysis. I think it's this like UXR method that is old school that people rarely do for some reason that's super powerful and I have rebranded it as a top task list. Um but it's essentially your task analysis and you do your top tasks and um I I describe top task as usually your 10 to 15 most common um uh commonly done or most impactful um tasks. Um and you want to pick ones that kind of I always say like criss-cross applesauce across your interface. You don't want ones that are all parallel that are are doing the same thing. So, um that way you get it's not a comprehensive like a full screen by screen heristic review, but what you get is a really good sense. Um again, in a in an agile world where you need to move fast, you get a really good sense of of what's what works because these interaction patterns are usually pervasive across any one experience, you know. So, um, by having the subject matter expert, you sit down with the subject matter expert up first and you figure out the top tasks, that's a that's a great first step for any other research, by the way. So, I tagged this in the beginning of Pure just because it that you could do a a moderated usability test unmodderated. It could help you write surveys, it could help you do any other things by framing it in a user center way simply by doing that number one of step two, uh, top task list. And then you break those down into interaction steps based on the user's happy path. Um this is the one bias of well there's two biases but there's one one bias of pure is that it's based on happy path. So it's not like a real usability test right. So it's not it's not necessarily user centered as in real users will interact with this but it is user centered in the fact that we'll be framing it in top tasks. We'll do the happy task and then we'll be evaluating it from like just think of these as like heristic review light you know it's like a heristic review without the all the the red lines and the reason why I think it's better than a heristic review in most cases in the real world in in like larger orgs is simply because I I just went through my career so long before I saw my like met Christian and started thinking about pure was uh I just saw so many heristic reviews of my own and others like just sit on someone's ask it's just like a snapshot in time and they just like they were not actionable. It was like here you go here's a couple recommendations and then it's gone and never used again. So the beauty of pure is imagine you could take that and then make a number out of it and a graph and then use that as a benchmark to grow and change over time. I I usually equate it to Rotten Tomatoes, right? like the the site um for movies uh and it gives you some like you said it's actionable where it's giving you that relative score and gives you almost like a color coding that people's eye not only draw to it um but then you can start segmenting by those top tasks right so then you can start kind of targeting okay I know you know these are in the green um but then these higher ones are are the ones that I really need to focus in on and then you can continue to iterate and test and go through the flow again and again in order to optimize the experience in order to you know create more of a usability centric exper or user centric experience which then can go into usability testing right it precisely this is a great step before you launch anything before you get it in front of real users before you have a beta before you have a working prototype right this is a great first step um and I believe uh I believe one of the key factors of this that that make it so valuable too is because you're breaking it down um into these like each interaction step. Um what you're gonna find is well at the end because we'll be able to score it, we'll have like a prioritized list of stuff to fix, right? And you're going to find that many of the same like patterns that are pervasive throughout your entire app, website, whatever you're looking at interface um will be uh they'll crop up time and time again. So it'll be like, "Hey, fix this." It just it's just like a moderated usability test in the fact that at the end you're going to get a list of stuff to fix and you should just fix it period. But it the one thing it does a little better is it gives you a nice little prioritized list for those uh on a budget. You know trying to you know you got to triage this kind of stuff. So it will tell you what will have the biggest impact. Then uh step three this is how I actually do it. I I I set up the adjudicators and the subject matter experts after the top task has been uh put together all so you have a list of tasks. Each task has a list of steps and they share their screen in Zoom just like this and everyone divergently scores as the uh asme goes through each um each of the top tasks one step at a time and they score and we'll talk about the scoring later but they score you know color code style like uh depending on what kind of scale you're using and then um they they they score each individual individual step kind of in a vacuum. You don't want to be like when I when I train people on this typically a lot of people try to be like okay it's like a journey map right like okay I'm going to put my persona hat on I'm you know like Larry the lagard so no no it's not that you are an expert and you're evaluating this just from the knowledge of UX best practices and by breaking it down step by step when you're pretty much just looking right where you know right where the user is looking like we have a lot of eyetracking data that says users are looking right where their cursor is it's right at that point of impact we're really dialing in on like the interaction design part of this and the usability. Um and then the then the subject matter expert after they it's nice to have the subject matter on the line because then the experts can ask questions you know because you can do these on you can do it on anything but like it could be a paper prototype you know this this is happens before high fidelity um in my in my how I like to recommend it um and then the subject matter expert leaves the meeting and the adjudicators converge to agree on a final score for each of the steps and that's how it goes. It's not it's not an average. Everyone has to agree and kind of like the best the best argument for each individual score wins out. And that's how you get your final scores. That's just the simple step. And then you make this great it can give you the ability to make these kind of graphs. Um I won't go too deep into this unless people have questions about it. But uh you can see this is these are literally from real examples. I've just scrubbed the uh identifying information and changed the numbers and stuff. But you can see there like task one understanding the curriculum. So there were one, two, three, four, five steps in that task. Each one was ranked a different way. And the way pure works is, and what Christian says in his articles is like you want it to be like golf. You want to have low numbers and on the green. So you can see that first task scored a yellow, which is not green. So you'd like to get it back to the green. Um, sign up for email was all green. And then um and then you can see the orange down at the bottom. Now, and the way these numbers work is you just add them up. So, if it's a the the little greens here are like a that's a one. Like look at task task two. There's three little greens. So, those would be counted as one. And then the two one two three four five six. Oh, that's like miscalculated. That should be a seven. I told you I mix the numbers up, but imagine that was a little seven in the green square there. That's funny. I just made these graphics yesterday off of real reports. It's kind of funny. Um anyway, this imagine a world in which you could talk about usability like this and then imagine a world where the next time you go through Oh, you can also I can talk to you about how to convert these to a percentage. This is something I do differently than um anyone else. Uh but I can talk to you about how to do that. Let's see. We got a question here. Top tasks are big in our org, but they often feel disconnected from UX research on the ground. This method seems like such a great way to bring them all together. Yes, great great point. Yeah, I I think a lot of people are doing that kind of activity in our orgs. Um, sometimes and it's not it's not inherently a a a UX researcher's job to do this, you know, but why not? And why it seems to be fit right in our wheelhouse and I believe it's a really good foundation for all all other research. So when I when I train teams on how to do like uh moderated usability tests, the first step is I just tag on a top task list exercise to go through. So there's the there's the the my simple everything's steps, right? There's my simple four-step process. You assign the roles, you you know find the right people, you uh do a top task list and break it down so you know how to score it. You do the scoring live. I people argue with me on this, but when you do the scoring live, um it is hard to get all these people in a meeting at once, but imagine the uh the time savings here just in general. Like the turnaround time for a report is like super fast. Um and then you you I I even recommend people don't even worry about the the reports and scoring for now. Just do what Christian Roar says in his articles. You know, you don't need to get too deep, but there are ways to turn them into percentages and then to have those be benchmarks that are that are over time. So um inherently within a pure report that I make, I always put v.1 or v.0. So it's inherently asking for another revision. And you can always run another one of these. And another beauty of it is once you have that top task list, you can just use the same one to to run your usability tests. You can use the same one to do any of the other methods. You've already done it once. And then if you have um if you run that top task list exercise up front as well then you you'll be able to go back and again benchmark this against itself. Um by the way that's why I converted to percentages is simply because people were really confused about they had to really get trained on the method you know so like it just like each in the pure method the the aggregate number at the end is meaningless. It's only meaningful to its own to itself. But if you're then able to figure out some kind of benchmark like I I do here. See that goal? That's like a that's like a benchmark that this particular organization has has agreed upon. Uh they want things higher than 75% uh of the UX hero6 accounted for and above a you know a two. Um so uh when when we were doing this Trevor we would uh we utilize like we would correlate our other statistics right so MPS as an adapted SUS score um and as we released certain features or or products or made iterations on it we could then go back and do that same pure method or same uh pure usability score in terms of that product or feature and be able to continuously benchmark um that one design off of what the the uh response was, right? So then we can start showing the correlation of the work that we're doing. Uh which is for many times very difficult within the design and research space um to show what that that outcome is of what the output is, right? Um and then start being able to to measure and show uh to executive or seuite uh you know what we're doing and how it's impacting the organization. Yeah. And it's just like like they're used to seeing stuff like that, right? So, how many times did I before I was doing the pure method, how many times did I go into some meeting and I was like, you know, trying to explain like Jacob Nielsen like like invented this idea about and it's just like oh boy, they don't want to hear any of that. But that's where that transitions to this behind the hood. How do we keep this legit? It's based on the Neielson Jacob Neielson 10 usability heristics. At least the way I do it, you could do it off of any rubric you want really. Um, but I like the the I like Jacob's 10. Um, and I literally have like a uh Oh, I I covered I had a typo. It said 15. There's only I've added four things to my rubric, but I have it I can I can uh share this if you like. Here's the criterion rubric that I use. I'll put it in the chat actually. And it'll have to I think that the Zoom chat's not big enough to hold all that. So, I'll put it in in three different things, but um I could share this doc, too. There's nothing cool about it, but you can see the first four are added and then the the the five through 14 are just the Jacob Neielson. And then I have a a little blurb on uh what those things kind of mean in context here in my own words. And then I put that blurb in. So in I I use a a a sheets or a Google or Google Sheets or a Excel spreadsheet to do the scoring. Um that creates those little graphs which is why that one was off because it must have had some like like formula problem in the sheet. But anyway, but under you can hover over the little thing and you'll get the the little tool tip with the explanation. So this is another way to get this like evangelize across the across your orgs. Um, so here, let me just real quick, I'll throw them in there. I'm sure most of us are familiar with these, but I do find uh I mean, maybe it's a good time to start talking about some of the second order effects because some of the second order effects you have here is There we go. There's the first. It looks like I can get six in at a time. So then here's seven through. Sorry, I just I pushed your email way up in the chat. or whatever you put in there, there's the rest. Um, that's just how I think about it. Again, this is derivating from the the Christian Roar stuff. Um, but I've seen it work in the real world, so that's why I like to share it. Um, yeah, some of the second order effects is those three experts are UX people on your teams usually internally. So all of a sudden, have you ever been in like our profession is so kind of in the infantile stages still that everyone's kind of using different terminologies, you know, I I never know uh when I'm working with a an early career person how much they actually know or not about like UX best practices. You I I've worked with seniors like UX design seniors that had very little knowledge of of interaction design. um uh like you know like the Jacob Neielson 10 usability heristics you know they may have heard of them but they wouldn't know them off the top this is a great framework to get everyone kind of going in the same direction and and it really like changes everything with your teams because the people designing know so in a lot of these places they've adopted it as a requirement it starts off like grassroots you know ground swell up the stakeholders start liking the fact that we can triangulate with usability it's no longer just like some expert talking at them. They're looking at numbers that change over time. You can calculate ROI off of this. So, how many like UX managers are trying to figure out the ROI of the usability stuff when all you're tracking is like active time, you know, it's like, you know, it could be there's some perverse incentives for that. You know, like Turboax wouldn't exist, right? Because or whatever if you're doing click count, Turboax wouldn't exist because click counts the clicks went way up but the time went way down, right? So depending what you measure is what's going to show up as far as your discipline. And I just feel like I don't know about you Brett and I now I'm going to get my old Kromaginly hat on is like I feel like we're just we're kind of like gone away from the basics these things that still matter that still have huge impact. So this is a way for UXRs to kind of like help the design discipline kind of refocus back to the things that we know to be true, the things that we've have 40 plus years worth of research around. Um, you know, so there's my soap box. Yeah, I'd agree with you. I think we that's where we nerd out, right? It's the there's a lot of basic methods that have been lost to time or lost to the influx of uh or overflow of boot camps hitting our industry. um were those that didn't get the foundational methodologies that uh allow for you know quantitative testing right even um quantitative usability uh I haven't seen a lot of people do that in the right way uh and this is a portion of that right um or could be substituted for it so I I think it's it's it's such a loweffort um methodology that that gives you both qualitative feedback because obviously as you're going through this uh experience or going through one of these sessions sessions, people are talking and people are giving, you know, user sentiment or expert sentiment um which is all valid and you can take that why behind what their uh quantitative responses are, right? So, um you're getting the the why and the so what behind even the the numbers uh that helps you derive like what the recommendation, what the adaptation, what the design changes should be um and a justification for it rather than just these numbers. And you know, I know Trevor, you're you're a a huge critic of NPS, right? Um and and that's largely why, right? Like you don't know why things are trending up or down. You you don't really know how to correlate it or why the experience is happening in the way that it's happening. Um this gives you a little bit more in order to do that. But uh I think here what you're talking about the three-point versus the five point, you know, the the traditional way to do pure is that three-point scale one to three. Uh but then you know would love for you to spend time on how you've even expanded that for that five point. Yeah I if you want I can talk a lot about that. I didn't know if we'd go this deep on it or not and in the chat say if you want to talk about something else but so I here's the article that I have about it. This is I've I've advanced it from a three-point scale which you'll find in Christian Roar's uh here. Let's just go to NG. It's on here. If you just search pure, you'll find it, my guess. There it is. Boop. All right. So, here's Christian's There's Christian Roar, by the way. Baller. So, there's your Rotten Tomatoes. Um, this is that graphic. Christian Roar's This paper is a one to three system, and it just says accomplished easily. like you believe that this would be accomplished easily. It would have notable degree of of cognitive load and it would be difficult for the target user. I found that to be really good for more simplistic interfaces, a very very good um system, but I've only really worked on super complex stuff for the most part. And I always it kind of always didn't sit well with me. So, the first couple years that I was doing the three-point system, I was like, "Huh?" So, I expanded to a five-point system and essentially, well, this was my first iteration of it. I'll skip past that. This just explains how I got there. Um, essentially, it was like I wanted a I ones I reserve for there's no meaningful difficulty. So, there's no heristic like violations within it. Um, there was like a one. A one means, you know, under the three-point system, it was like maybe a little bit of trouble, but not too bad. Well, I wanted a one to be like nothing. I wanted a one to be like this is there'd be zero, little to no friction at all. And then I the main thing I missed was the five that I have here was a critical error. There was no way to say the user users would be unable to do this right? And there's all kinds of stuff like that. And it always got in this weird gray area in which, you know, it was a like I I'd give it a three under the three-point system, but it was really more. It was like a three plus, but there was no such thing. So, I expanded the scope or the the this um and I give little rules of thumb here. These are not always true, but they were true in the last handful of organizations. the rule of thumb of like if there's one or two violations that are, you know, impactful, uh, then I would give it a two. You know, three to four would be a three, five or more would be a four, that kind of thing. Now, that's really deceiving and it gets it really muddies the water on the on on the pure kind of method. Um but this has been proven to be uh consistent and uh able to be like quantified as um just as valid as any of the other uh as the three-point system from a statistical reliability. And what I found is mostly it's just like all the arguments in the converge stage are just around is it a two or a three or a three or a four, right? And usually people kind of know what a no one ever argues a one or a two or a one or a five. So really you get down in in the weeds here. And we found that like when something gets mis scaled or miscored it comes out in the wash because of uh just the amount of scores that you do within a typical pure review. So that's my kind of like that's my kind of like uh highle pitch there. Um but you can still just do the three-point system. It it is it is a little more simple. Uh, I just prefer to have that ability to say a user would not be able to go past the past this point you know. Um, yeah, again, here's I just wrote this out one more time. Uh, so you can actually see it. I thought this might be a good slide for people just, you know, they share their screen, the adjudicator assigns the steps, adjudicator writes a short note um, in the in some some way. I do this in a spreadsheet, but you can do it at all. I have a one one team that does it in Muro like on stickies. They do it like a workshop which is kind of cool. Um and then once all the scores uh and notes are in there then the the subject matter expert moves on to the next step. And then you just do that over and over and over throughout the Zoom call. And at the end you've got the score. Um I I typically will facilitate these while the expert adjudicators are kind of learning about this um after they've been onboarding onboarded to the system. Um, oh, another thing about when you're putting together your expert panel, I always try to have one advanced person on there, like one at least one person that kind of has dumped here before, is familiar, and is the most senior. Um, I would highly recommend that. Um, it does help in the converge stage at the end. Um, and then, uh, I forget why I showed this just because I thought it was a nice way to synthesize everything into like another five-step process, right? So once you figure it all out, this is how you can actually run one of these. And by the end, you will have all the data you need to make a report with the graphs and stuff. It's pretty cool. And then here's all the second order effects. Is there anything you want to talk about the scoring? Is there any questions anyone has? Because this stuff gets I'm giving such a higher overview right now. It's like it's like ridiculous. But I've written, like I said, I think there's I think that article about expanding the Pure system is probably like a 3,000word article. There's probably 20,000 words in that list that I put together about thoughts on Pure if you want to really dive in. And this one too. Did you Did you share this in the this link in the chat? Sure. Share that in the chat. Yeah, we have a question of hey, can you go back to that previous slide with the scoring? Um seems to be some questions about that. Uh I think what would be helpful Trevor? Um nope. I believe the uh your article um with the actual scoring. Yeah. There. Sure. Um I think what would be helpful is like if we were to simulate right um imagine that we are on a pure and and you know uh giving a guided task and then how the expectation of what you've seen um you know those participants going through and how they go about you know giving those ratings and then from there um how would you how would you make sense of it right we've used an excel sheet in the past where it just calculates it for you so easy you know particip spent one two three four five etc. Um that's able to to actually go through it. But then uh you know if you could just give give a quick example to to walk people through so they can have something a little bit tangible to to kind of concrete themselves with. Right. Yeah. You mean something like this? Exactly. Yep. So this is what I actually use to create those graphs. Um here this is this is this how we really do it. Uh so this is this column represents the task number. This represents the steps. These were derived from the subject matter expert top task conversation. This is wireframe/page title whatever you want to call this um you know in your if it whatever uh and then this is the action that's done. So, login, find, learn, da da, and then you just type in your Oh, here's the spreadsheet, right? So, you hover over uh will the customer realistically do this? You can hover over it gives you the context from this sheet. It's literally the words that are written under here. So, evaluators can do this on their own time. Um, and then they come in. See, if you if you look, this is a step one is a five. Let's give it a one. Changes to a one. And then adjust the score. the the aggregate score always results the darkest color. Now, this is a fivepoint scale. I have one with a three as well, but you get the idea. And um so this Oh, uh don't think about I'm going to hide this. Don't think about how do you hide in sheets? Geez Louise. Just pretend you can't. There we go. Uh you don't really The average friction score is not helpful in most cases. Um, so, uh, so that the whole the whole score at the top is an orange in this case because the darkest color is orange. So, there's one thing I totally skipped over is you adjudicate. It's it's a two-fold adjudication. And you think of the color and the number as separated, but but you know, they're the same, but they're different in the fact that, like I said, if this were uh if if we turn this down, this one to a three, everything will change to yellow because it only reflects the worst score of any one step. The idea here is the overarching um the overarching score for any one given um uh full aggregated pure review is is uh defined by the worst moment within the user's experience. That's why it works that way. So what happens? So when you look here, what happens in the case? So why is this a rule of thumb? And why is it why is it not like exactly three to four? Why isn't it like simple math like that? Because the colors are separated in think of the colors as impact. So when I train people on this, I kind of have them do a little mind trick of their own. It's a great question by the way, whoever asked that. So what I have them do is I first have, you can see here, step two, adjudicators assign that step a number. I first had them assign in a number and then retroactively go back throughout the heristics and try to map what they found to the heruristics because I found that many people's guts on this are fairly accurate. They're fairly accurate and then they go through and then the the just the knowledge that they're going to have to go map it to a heruristic is what really matters there. And furthermore, at the end when you have all these things done, um the the the subject matter expert has this as a document too. So in my reports, I always put a link to this raw data and they could they could always go through and you know, oh there's a three. Why is there a three? And they they can scroll along here and say, oh, and there's usually notes in here from every one of the adjudicators like like, hey, you shouldn't have made that a hidden button. Why is there a hamburger on a full, you know, why are you using a hamburger on the on a wide uh viewport? Like stuff like that. So I kind of answered the question but not really. I just wanted to clarify the fact that the colors are more so there there are many cases in which it's a four but there's only one real heristic that's broke that's been violated. Right? So this says five or more. It's almost like five or more just normal five or more places in the rubric that you can in this scoring rubric that you can make little notes. You can think of it like that. Maybe it's five bullets under one thing. Does that make sense? So like visibility of system status with the the hamburger menu, maybe there's five different reasons why that violation uh has breached has been breached and will create problems and then that results into your four number. Does that make sense? So but my recommendation for the real world is just have them pick a number, map it to the rubric, and then um and then have them hash it out in the converge stage. And that's really how we get good valid results over time. And I' I've tested this like I know I know for a fact that the numbers are you know they're they're good valid and helpful and you don't get this wide uh wide bias swing is that how is that it's kind of complicated and it gets into weird sticky situations but I I I oh here let me read the actual are the one through 14 shared in the June chat. Oh I can share this document out there. So, Chris Ring asked if there's one if if this I can just share this doc. How do I do that? Share with anybody viewers. How's that work? Um, also Brett, you could make a version of this your own. You could house it on on your website or something if you want. Oh, I see people streaming in. Hello. Feel free to take a look at that. That's just how I think of it. Um, I will do one note. I have had people say, "Can we add accessibility to the end here?" I've tested this out. I've made a good faith attempt to in include accessibility into the rubric and it it it frankly it's at the wrong level of fidelity. So it doesn't quite work at this level because what it does it transforms a pure audit into a like a a QA session, right? People are could you rightclick on that? Could you inspect the element? Is there an ARA label? Like stuff like that like how is this is this done semantically? Like so um I believe that uh accessibility should be a requirement and floodgate that happens at a higher level than a heristic review or or a peer review. So just like you wouldn't, you know, handle that in other ways, it should be a a different set of criteria. It should be part of the QA process. I believe if anyone had if anyone was thinking about, hey, where's accessibility in here? Oh, and here's the beauty of this. At the last few organizations, we had anywhere from 6x to 8x time reduction from an old school heristic uh review to a pure audit. And that includes that you have four people instead of just one. So it's four people spending that much time is still less time than a single person doing a old school heristic review. And I believe they're much more useful in the real world because it creates these graphics. You speak straight to stakeholders and and all these second order effects here. Yeah, it's the digestibility. I think you you spoke to it well. It's uh a lot of the rooms as we go into as researchers, this is how a lot of our executive teams like to be communicated with, right? Show me the numbers. Show me what we're doing. Show me what the improvements are. Um and Pure does a good way of doing that. Uh and and you know, like Christian says in the the article, right, Rotten Tomatoes, um very real world aspect of being able to grasp that. Um, in terms of, uh, you know, learning, uh, pure and and being able to upload your your teams on this or or download your teams on this, uh, we're more than happy to set up time, uh, with anyone that might be on the call today to to go over how to do that. Um, train your research teams in order to do that and then, you know, you can take it and implement it into your organizations uh, as you see fit. Nice. I I also within these articles I have links to these templates for free too so you can find them in there like you know one of those one of those links that makes you it makes you make a copy. Um yeah so I mean that's kind of the overview. I just want I just want to make it clear that first of all this is just by this like 45 minutes that we've been talking about it is probably not enough for you to actually do this yourself. Um but it gives you an idea and this is the kind of like like if you're going back and looking at this recording and and you're you think this may be helpful for your organization just feel free to show them this um this recording. This is a nice overview. it kind of explains some of the things. Um, and and I've literally seen all these second order effects happen in the real world. Like there there's nothing here that's hypothetical. Um, that's my whole shtick is like what what is what does this really mean? Right? Let's go past the theory and into reality. And when we do these things, people are starting to use the same language. Stakeholders are starting to look at usability as a as a real thing simply because it's being uh talked about in ways that they understand. um early career colleagues are just like built-in upskill mechanism because they're look the first drafts are stronger of designs because everyone knows oh I my design will eventually be uh held to the standard of the the 10 heristics you know it's just like this really great virtuous cycle and I think the number one thing that I see when people adopt pure as a system across a larger team is the cross collaboration I've spoken to so many managers that are like UX managers that are like um you know if your team is distributed they spend so much time in like like meetings in which people like crit meetings and meetings where people just show what their work is you know like it is a built-in cross uh cross project kind of like like uh uh collab session right and it's very structured and it's structured around a UX. I don't know if you've ever been in a critique where like a design critique where people like basically arguing their their thoughts, feelings and not basing any kind of rubric. It just cuts through all that noise. And then it also uh it's good for morale because like people no long you're you're no longer saying, you know, like like like Brett didn't like that design in the crit, you know. It's not that. It's not like Brett has criticism or Trevor has criticism. It's it's the team of expert evaluators have, you know, gone through this process that we've all agreed upon and have made a judgment. So, it it kind of gets rid of all of that uh um uh kind of interplay and it allows for proper critique. It allows for a proper um feedback loop in a way that's not um destructive but can also be super super beneficial. Awesome that time. Uh I appreciate it. Obviously uh you know we always like to show Trevor's expertise uh and also his his blog. I'm an avid subscriber and and I think you know uh he feels he he feels the love from us coming to uh all the stuff we we like or love uh every week when he produces it. I think Fridays are usually the time that uh I see the most action. Um but you know go go uh follow UX in the wild. Um obviously reach out to to us at Accelerant if you want to learn more about Pure. uh we can share some of the resources that we went over today uh but also we can step in and help teach and evangelize what this methodology is that's almost been lost to time, right? So, uh Trevor, appreciate the time, appreciate the the conversation. Um I know this is a a little sliver maybe 5% of what Pure really is and we can go into a lot more detail. We can go into the nuances of it. Um, and you know, from from the scaling system to just pairing it up with usability or or other things that are just much more meteor, but um, appreciate your your expert opinion and expertise. Yeah, thanks. I've spent about eight years on this, so you know, it's like, hey, here's my life's work in like a couple slides. Um, so feel free to reach out to me. I'm I'm very active on LinkedIn. If you go to UX Research in the Wild, find me on my my Substack. Um, I all the stuff's for free. I don't put a payw wall. I just use Substack because I just used it for no good reason except because it was like easier to publish from. So, um, nothing's under a payw wall. And, um, I do have some people that just pay me randomly though, which is kind of cool, but that's whatever. Um, I I I messaged them, you everything's for free, and they're like, "We're going to pay you anyway." I was like, "Yes." Okay. But, um, uh, I do have a day job, so uh, it is um, it is uh, I'm available to help. you can just reach out at any time, but just know that, you know, I am kind of like in infrequent. Like if I don't get back to you right away, it's just because it's not like I'm selling this system or pure, you know, it's just like this is just what I've done that people seem to like and now people are asking me to talk about it. So I appreciate the opportunity and um yeah, this has been great. So feel free to hit me up and we'll keep the conversation going. Awesome. Uh I think we are at the time. So I'll hand it back over to Bill um who's going to be doing the the next session.

