Human-Computer Interaction
17 Aug 2026
Part III of Hornbæk et al. (2025)
Interviews · Field Research · Surveys
For three weeks we’ve been building a general model of the human, which is powerful and a good starting point, but it’s not specific to your users.
This week’s question: you are not your user, so how do you find out who is, and what they do?
Uber got off to a rocky start with women. Because the all-male cofounders designed for themselves.
This is the founding commitment of user research. Three reasons it’s hard, and each one previews a method this week.
Note
User research is the set of empirical methods for obtaining, analyzing, and representing knowledge about users, their activities, their contexts, and the systems they already use, in order to inform design.
If you are not the user, your first job is to say who is. That is a three-step commitment, before any method.
Note
Identifying the user is a three-step commitment: specify the target audience, map the other stakeholders, then sample representatively. A failure at any step injects bias that is hard to detect and often only surfaces at deployment.
The people who wreck your design are often the ones who never made it onto your participant list.
Pause and think
You are designing a homework app for a school. Name three stakeholders who are not the students, and for one of them, argue why they might be more urgent than the students themselves.
Every method trades off three things you cannot maximize at once (McGrath).
Note
Because every method is biased in its own way, the fix is triangulation: combine methods with different weaknesses so they cover for each other. Hold this throughline: interviews, field research, and surveys each buy one corner of the triangle and pay for the others.
Choosing a method is only half the job. The other half is judging whether a study was done well. The textbook names four quality dimensions.
Note
Validity asks a different question than reliability: not “would I get the same answer again?” but “is this answer about the thing I think it is about?” A study can be perfectly reliable and completely invalid.
Reliability is a tight cluster of darts. Validity is whether they landed on the right board at all. You can measure the wrong thing very consistently.
Pause and think
A team measures “user delight” with five questions and gets a very high Cronbach’s alpha, so they call the scale trustworthy. Which quality dimension have they shown, and which have they not? What would you check next?
Note
A leading question is one that implies a “right” answer and injects the interviewer’s prejudice into the interviewee’s response, for example “Don’t you think autocorrect should be on by default?”
The rule is easy to state but brutal in practice, especially when you built the thing you are asking about.
In the wild
Pause and think
Rewrite this into a neutral, single-topic question you could actually ask a user: “Don’t you find the new onboarding both faster and clearer than the old one?” Name every flaw you removed.
Note
Contextual inquiry trades generalizability and precise task timing for realism and a full understanding of a few users’ activity in situ.
Contextual inquiry exists because asking alone hits the say-do gap and the tacit-knowledge wall. Being present lets the artifacts, the room, and the interruptions jog what a bare interview would miss.
Do interviews tell you what users do?
Pause and think
Your team wants to know how often people actually use a rarely-touched “export” feature and how they feel about it. Which part should you get from an interview, and which part should you refuse to trust from the interview? Why?
Collecting the interview is only half of it. Then you have to make sense of it, from the participant’s perspective, not by jumping to a design.
Note
Analysis is where interview data becomes findings. The aim is to make sense of the participants’ world in their own terms; any clever design idea is secondary.
Pause and think
You have 30 transcripts and 400 sticky notes on the wall. In your own words, what is the difference between a code and a theme, and why can’t you just count which words came up most often?
Note
Reactivity is the change in people’s behavior caused by the observer’s presence. Imagine your employer hired someone to sit and watch you work.
Field research buys realism: it captures collaboration, workarounds, and the gap between the official procedure and how work actually gets done. The cost is reactivity and effort.
Pause and think
You have two hours to observe a busy call center. Would you go etic or emic, and name the one thing your choice will systematically cause you to miss?
Note
The classic ethnographic example: to a newcomer a print room is undifferentiated noise; to experienced operators the machine sounds decompose into a rich resource for monitoring each other’s work (Button & Sharrock).
Here is the catch the chapter is honest about. Realism makes field data vivid and contingent, which is exactly what makes it hard to turn into a general design implication.
Does watching users actually inform design?
Pause and think
You run a two-hour observation of nurses using a medication cart and it is spotless, no errors, everyone by the book. Name two reasons this “clean” session might be misleading, using terms from this section.
Note
Self-selection bias is the distortion introduced when the people who choose to respond differ systematically from those who do not, over-representing whoever had a reason to answer.
Generalizability is the selling point, and it is exactly where surveys quietly fail. Scale multiplies a biased question or a skewed frame; it does not fix them.
In the wild
Pause and think
You need results that generalize to all users of your app, but you can only recruit from your existing mailing list. Name the population, the sampling frame, and the single biggest threat to generalizing from what you collect.
Note
Response biases distort answers systematically: acquiescence (agreeing regardless), social desirability (looking good), response-order and neutral-item effects, and demand bias (answering to fit the study’s apparent purpose).
A validated scale measures a construct with several items and checks that they move together. Useful, but easy to over-trust.
Does a validated, high-response survey give you the truth?
Pause and think
A Product Manager shows you a survey: 12,000 respondents, 92% “agree the new feature is useful,” recruited via a banner inside the feature. Give two reasons that 92% might be worthless, naming the specific bias each time.
User research assumes empirical data beats the designer’s intuition, yet the chapter admits empathy tends to fail and Norman warns that following users too closely breeds feature creep.
Is rigorous user research a genuine correction to designer bias, or does it just launder the same biases through a method? Where is the line between listening to users and letting the loudest signal design your product?
Field research is strongest when people do not modify their behavior for the observer, but reactivity is reduced most by watching people who did not fully consent.
Weigh the realism you gain against the autonomy you take. Would you run a covert study of a vulnerable group if it were the only way to surface a real harm? What would make it defensible, or never defensible?
Teams increasingly prompt a model to role-play respondents or interviewees instead of recruiting people, arguing it is faster and cheaper at scale.
A synthetic user has no tacit knowledge, no say-do gap, and no future it cannot imagine, because it has no behavior at all. Does simulating users reintroduce the exact “you are not the user” problem the field exists to solve, or is there a defensible use? Name where you would and would not trust it.
Task: In groups of 3 to 4, take the provided set of eight flawed interview and survey questions (leading, double-barreled, vague, future-hypothetical, socially loaded) and diagnose then repair each one.
Produce: A table with at least eight rows: original question, the specific flaw named, and a neutral rewrite. Plus one forced position: pick one question the chapter’s rules would flag that you think is actually fine as written, and argue why.
Time: 25 min group work · 10 min share-out
Debrief: Which flaw was hardest to fix without introducing a new one?
Task: In pairs, one person performs a real, mundane task on their own phone (find and mute a specific notification; add an event to their calendar; split a bill in a payments app) while narrating nothing unprompted. The partner is the “apprentice”: observe, take timestamped notes, and use only concreteness probes (“show me the last time you actually did that,” “wait, why that tap?”). Swap. Then, from your notes, do 5 minutes of open coding on the observed behavior.
Produce: For each observation, one code (bottom-up) and one say-do gap you saw — something the user did that they would probably not have reported in an interview. Then name one code that is at risk of being your interpretation rather than their behavior.
Time: ~24 min paired work (12 observe + swap, ~5 code) · ~11 min share-out
Debrief: What did watching surface that asking would have missed — and where did you contaminate the data by being the apprentice?
Task: In groups of 3 to 4, choose a real AI feature (a coding assistant, an AI chat agent, an AI photo editor). Write three research questions about its real use, then assign each to a method (interview, contextual inquiry, field observation, survey) and justify the choice with McGrath’s realism, precision, and generalizability.
Produce: a one-paragraph research-plan memo written to a skeptical PM — “here are our three questions, the method we’d use for each, and the one place the say-do gap will bite us.”. Forced position: name one question where you would trust an LLM-simulated “synthetic user” and one where you absolutely would not, with your reasoning.
Time: 28 min group work · 12 min share-out
Debrief: Where did two groups pick opposite methods for the same question, and who was right?
This slideshow was produced using quarto
Fonts are Fira Sans, Fira Sans Light, and Victor Mono Nerd Font
Math is set in Fira Math via MathJax 4