Human-Computer Interaction
17 Aug 2026
Part VIII of Hornbæk et al. (2025)
Last week we lived inside the lab: think-aloud studies and controlled experiments, where we bought precision by stripping the world away. But a keyboard that wins in a tapping task can still lose on a real phone, in a real pocket, or on a real commute.
Three stops on one spectrum of system completeness: field evaluation of a prototype, then a pilot study of a partial-but-real system, then a deployment study of the fully shipped thing.
This week’s question: what do we learn when we drag evaluation back into the messy real world, and what does it cost us?
A device that dazzled in demos and delighted its own testers, then collapsed on contact with real streets, bars, and strangers.
Realism, and what it costs: ecological validity, degrees of the field, and experiments that leave the lab.
Note
An evaluation has high ecological validity when its people, tasks, context, and technology closely match real use, so that observed behavior reflects what users would actually do outside the study.
Is the field always worth the hassle?
Pause and think
A team deploys a new feed to whoever opts into the beta, then compares them against everyone else. They find the beta group is more engaged. Name the threatened assumption, and give one reason the difference might have nothing to do with the feed.
Putting an unfinished but real system into real use: pilot implementations, technology probes, and the minimum viable product.
Note
A pilot implementation is a limited-time deployment of a real but unfinished system into its actual context of use, run with real data and safeguards, in order to learn about and reduce the risk of the eventual full implementation.
Pause and think
Your team wants to test a new hospital scheduling system, but it is only 60% built. Would you run a lab usability test, a pilot implementation, or an MVP release, and what is the one thing each would tell you that the others would not?
Evaluating the fully shipped system, and deciding when realism is worth the hassle.
Note
A deployment study evaluates an interactive system after it has been fully released, gathering data from real use, such as feedback, logs, and behavior, to judge whether the system is achieving its purpose and how to improve it.
Pause and think
A banking app team wants to know how people navigate their newly launched site. They have logs from millions of sessions and a stack of one-star reviews. You already know to triangulate; now name the population each source structurally cannot see, and which one your specific question depends on.
Field studies buy you the truth of real use and hand back your ability to say what caused what.
Kjeldskov found the lab caught the same problems for a fraction of the cost, yet the nurses’ collective reading only ever appeared in the field. Where is the line? For a given evaluation goal, students should be able to argue when the extra realism earns its price and when it is expensive theater.
The users who opt into your beta are not a random sample, and that is exactly why the result may not generalize.
The Wikipedia Adventure died on self-selection: opted-in newcomers were simply different. But one could argue that the people who choose your product are the people who matter. Is filtering by motivation a bias to correct, or the real-world condition you should be studying?
Every shipped feature is a live experiment on people who never agreed to be studied.
Log analysis, A/B tests, and feed tweaks run on real users continuously, often without meaningful notice. When does routine product measurement cross into unconsented human-subjects research, and what would honest field-evaluation ethics actually require of a company?
When an AI agent personalizes itself per user and acts on their behalf, no two people see the same system, and “the task” is negotiated in real time.
If an AI assistant books, buys, and writes differently for every user, there is no fixed interface to evaluate and no shared task to standardize. What does ecological validity even mean here, who counts as “the user” when the agent acts, and can a controlled field experiment survive a system that rewrites itself for each person?
Task: In groups of 3 to 4, take a simple lab usability test (for example, a food-delivery app checkout) and redesign it as a field evaluation, deciding where to sit on each of McGrath’s four dials: people, activities, context, technology.
Produce: A single dial diagram — draw McGrath’s four dials (people, activities, context, technology) as sliders, mark where the lab sits and where your field version moves each, and write one sentence on the dial you deliberately left at “lab.”
Time: 25 min group work · 10 min share-out
Debrief: Which single dial, turned to field, would most change your findings, and why?
Task: In groups of 3–4, take a partially built system (a 60%-complete hospital scheduler, a new internal tool, a campus app). Design a pilot implementation, not a lab test: real but unfinished system, real context, limited time, real data, safeguards against breakdown.
Produce: A one-page pilot plan stating the one question the pilot must answer, what you’ll instrument to answer it, the safeguards you’ll run, and — explicitly — the list of things this pilot cannot tell you (so nobody over-reads the signal). Name one sociotechnical effect (à la Hertzum’s nurses reading the record aloud together) that a solo usability test would miss and your pilot might catch.
Time: 25 min group · 10 min share-out
Debrief: What’s your “shadow spreadsheet” — the workaround that, if it appears during the pilot, tells you exactly where the design misses reality?
This slideshow was produced using quarto
Fonts are Fira Sans, Fira Sans Light, and Victor Mono Nerd Font
Math is set in Fira Math via MathJax 4