Human-Computer Interaction
17 Aug 2026
Part VIII of Hornbæk et al. (2025)
Last week we were in engineering territory: systems, safety, verification and validation. We asked whether a system was built right.
This week’s question: how do we decide whether an interactive system is actually good, and whose judgment counts?
A patient arrives on time for an appointment, cannot get through the self-service kiosk, and is not even sure whether they are checked in.
Note
Evaluation is the attribution of value: assessing a design against explicit criteria (yardsticks) to conclude, in a systematic and defensible way, how good an interactive system is for its end-users.
Note
A formative evaluation is run to shape and improve a system in progress, feeding the next iteration. A summative evaluation is run to judge a finished system against fixed objectives, for procurement, contracts, or a go/no-go decision.
Pause and think
Your team ships a feature and a PM asks for “a usability test to prove it’s good” the day before launch. Is that a formative or summative request, and what is the risk of answering it with an early-stage method?
Note
A yardstick is the standard against which value is judged: it may be absolute (meet a fixed target, such as a System Usability Scale score near 70) or relative (beat a baseline system on some measure).
Pause and think
A client says the new checkout should be “faster.” Give one absolute operationalization and one relative operationalization of that yardstick, and say which method each implies.
Note
Validity: do the evaluation’s findings reflect the real value of the system for real users doing real tasks? Reliability: would you get the same findings if you repeated the study or swapped evaluators? Impact: do the results actually get used to change the system?
Is evaluation always worth it?
Pause and think
Two evaluators run the same heuristic evaluation and report almost no overlapping problems. Which of validity, reliability, and impact is most threatened, and would adding a third evaluator help?
Note
Analytic evaluation methods assess a design without collecting data from real users: an expert compares the interface against guidelines, principles, or performance models to predict problems, focus areas, or task times.
Note
Heuristic evaluation is an analytic method in which evaluators inspect an interface against a small set of usability heuristics (rules of thumb) and record where the design breaches them. The most cited set is the ten Molich and Nielsen heuristics.
Pause and think
You flag a “consistency” violation but your teammate does not see it as a problem at all. Is that a validity issue, a reliability issue, or both, and what does the disagreement tell you about heuristic evaluation?
Note
Detecting a usability problem can be modeled as a Bernoulli trial: each evaluator finds a given problem with some probability p. The chance that at least one of k evaluators catches it is 1 minus (1 minus p) to the power k.
Do five evaluators really find 85% of problems?
Pause and think
Your budget covers three evaluators. They are CS students with a roughly 30% detection rate. Using the intuition from the Bernoulli model, should you promise a stakeholder “thorough coverage”? What would you say instead?
Note
HEI is an analytic method that starts from a task analysis, enumerates the states a system can be in, and builds a matrix of legal, illegal, and unavailable transitions between them. A user is then simulated moving toward a goal to expose where they could take a wrong step.
In the wild
🎤 Therac-25
Pause and think
At a self-checkout, users constantly pay before scanning the last item and walk away. Which of HEI’s three error types is that, and what state-transition change would prevent it?
Note
A cognitive walkthrough is an analytic method that mentally simulates a novice exploring an interface: for each step of a task, the evaluator asks whether the user would try the right action, notice it, connect it to their goal, and see progress after doing it.
Pause and think
A new user stares at a screen with a valid button right in front of them and never presses it. Which of the four walkthrough questions most likely failed, and what redesign tactic addresses it?
Note
KLM is a simple mathematical model that predicts how long an expert, error-free user will take on a sequential task by summing standard time estimates for each operation: keystrokes, pointing, homing, mental operators, and system response.
In the wild
🎤 4000 Clicks
Pause and think
Two login designs do the same job; one keeps hands on the keyboard, the other forces repeated mouse-keyboard switches. Before doing the arithmetic, which KLM operator predicts the difference, and what would KLM completely miss about the two designs?
Note
A think-aloud study asks participants to verbalize their thinking while they use a system, so the evaluator can capture what they are trying to do, how they read feedback, and where they get stuck. Nielsen called it the single most important usability engineering method.
Note
The instructions and tasks are the biggest levers on what a think-aloud study finds. Classic thinking aloud tells users to “keep talking” and behave “as if alone in the room”; tasks should be representative, and worded in the user’s terms, not the system’s.
Pause and think
You want to know whether users can figure out how to share a document. Write the task two ways, one with hidden help and one without, and say what each version would and would not reveal.
Note
Analysis is a distinct step: the raw verbalizations are transcribed and segmented, then interesting moments are identified and classified, ideally by a different person so inter-rater reliability can be checked. Classification runs bottom-up (let themes emerge, as in affinity diagramming) or top-down (apply an existing code set).
Pause and think
A user sighs, says “I guess I’ll just try this one,” clicks the wrong menu, and eventually gives up. How many distinct usability problems is that, and which of the three types does each fall under?
Note
The evaluator effect: different evaluators, given the very same sessions, surface markedly different sets of usability problems. It hits both analytic methods and think-aloud, and it is fundamentally a reliability problem.
Can two analysts trust the same tapes?
Pause and think
Your team runs think-aloud sessions on the check-in kiosk, and a single analyst codes them into one problem list. Given the evaluator effect, how far should you trust that list, and what is the cheapest change that would most raise your confidence in it?
Evaluation is treated as indispensable, yet a serious line of work argues it can strangle good ideas.
“Usability Evaluation Considered Harmful” argued that mandatory, thoughtless evaluation punishes early, visionary prototypes that are supposed to be rough. Where is the line between rigor and premature judgment, and who decides when a system is “ready” to be evaluated?
The evaluator effect shows the same tapes yield different problem lists depending on who looks.
Only a fifth of problems were found by all four evaluators in Jacobsen et al. If reliability is this low, is a one-person heuristic evaluation worth anything? What would it take to make you act on a usability finding, and is that bar even affordable in practice?
Analytic methods are cheap and fast, but carry high false positive and false negative rates.
A PM wants to skip user testing entirely and ship on the strength of a heuristic review because conversion pays the bills and the deadline is Friday. When is analytic-only evaluation a smart trade, and when is it negligence, especially for safety, accessibility, or inclusiveness?
Amershi et al. wrote 18 heuristics specifically for interactive AI because the old ten did not fit.
When the system’s behavior is probabilistic, learns from you, and cannot fully explain itself, do classic yardsticks like “visibility of system status” or “user control” still mean anything? And when an AI agent acts on your behalf, who is even “the user” whose experience you evaluate?
Task: In groups of 3 to 4, individually run a heuristic evaluation of the same real interface (a travel planner, a volunteer org homepage, or a settings panel) using the Molich and Nielsen heuristics, then combine your findings by:
Produce: Each group posts its individual problem lists on the wall; everyone walks and tallies which problems appear on only one list vs. all of them.
Time: 22 min group work · 13 min share-out
Debrief: How much did your individual lists overlap, and what does that overlap (or lack of it) say about trusting a solo evaluation?
Task: In groups of 3–4, pick one first-use flow (create an account and complete one core action). Half the group runs a cognitive walkthrough: for each step, ask the four questions (right goal? notice the action? associate it with the effect? see progress?) and log every “no” you can’t explain away. The other half runs a think-aloud with one recruited outsider (a student from another group who hasn’t seen the flow), using neutral instructions (“keep talking,” “behave as if you’re alone”) and tasks worded in the user’s terms, not the system’s.
Produce: A merged problem list. Tag each problem: found by walkthrough, by think-aloud, or by both. Classify each as failure-to-reach-goal, misunderstanding, or negative-experience. Mark which problems only one method caught.
Time: 28 min · 12 min share-out
Debrief: Which method caught what — and given the evaluator effect, how much would you trust either list from a single analyst?
Task: In groups of 3 to 4, audit a real AI feature (autocomplete, a chatbot, an AI photo tool, or an AI-generated UI) against Amershi et al.’s 18 AI heuristics, and pick one task to model with KLM or probe with a quick think-aloud.
Produce: An annotated screenshot flagging at least 5 heuristic violations, plus either a KLM estimate for one task or three think-aloud-style user quotes you would expect. Include one position statement: name one AI heuristic that is impossible to satisfy with today’s opaque models, and argue why.
Time: 25 min group work · 15 min share-out
Debrief: Which failures were unique to the system being AI, and which are just old usability problems wearing a new coat?
This slideshow was produced using quarto
Fonts are Fira Sans, Fira Sans Light, and Victor Mono Nerd Font
Math is set in Fira Math via MathJax 4