A within-subject crossover design
Useably runs a within-subject, four-task crossover, not two separate surveys. Each participant experiences both conditions — the system under test and a comparison condition — so every comparison is made within the same person. Condition order is assigned by block-of-4 randomizationat consent and is one-way: once randomized, a participant’s assignment is fixed and database-enforced, so the sequence can’t drift mid-study.
This removes the dominant confound in tool comparisons — differences between the people, or the questions, rather than the systems. Same participant, same cases, different system.
The protocol, step by step
A server-authoritative step machine is the spine of the study. The participant’s current step lives on the server and is the single source of truth for routing; instruments are administered inside the task protocol rather than bolted on at the end.
- 1Confirm profile. The participant lands from a single-use registration link or an emailed six-digit code, then confirms name and training level.
- 2Consent & randomization. Informed consent is recorded; block-of-4 crossover order is assigned at this moment.
- 3Baseline survey. Background and prior-experience measures captured before any task.
- 4Orientation & practice. A guided orientation and an untimed practice task so the protocol mechanics never confound the real measurements.
- 5Four timed retrieval tasks. Each task runs a fixed micro-sequence: begin → scenario → prior-knowledge → timed retrieval (300-second cap) → response → NASA-TLX. Timing is measured and capped on the server; the on-screen timer is display-only.
- 6Usability instruments. A System Usability Scale is administered per condition; an additional usability questionnaire is captured for the AI-assisted condition.
- 7Post-survey & completion. Closing measures, then the protocol marks the participant complete.
Controlled case assignment
Tasks are driven by a versioned, per-cohort case bank. Every participant in a cohort sees the same clinical scenarios in the same protocol order, and the case content is version-controlled, so a result reflects the system under test rather than an easier or harder set of questions. Cohorts can diverge only deliberately, by editing their own case bank.
What gets measured
Each task and condition produces complementary, pre-specified outcomes:
- Time to answer. Server-measured and capped — a tamper-resistant efficiency outcome, not a client-reported estimate.
- Cognitive workload (NASA-TLX). The validated Task Load Index, captured per task across six workload dimensions. (The performance dimension is stored on the conventional reversed scale and flipped only at analysis.)
- Perceived usability (SUS). The 10-item System Usability Scale, scored per condition on its published 0–100 scale.
- Retrieval correctness (adjudicated). A trained reviewer adjudicates whether the retrieved source actually answers the clinical question — the difference between “found a document” and “found the right answer.”
Data integrity & governance
Submissions and audit records are append-only, enforced by database integrity triggers: instrument submissions, consent records, the event log, and task-timing rows reject silent updates or deletes. Corrections are tracked through an explicit soft-delete overlay rather than overwriting history, so the dataset stays defensible under review.
Free-text answers that could carry identifiers are automatically scanned and masked by default; a reviewer must explicitly reveal, mark safe, or redact each one, and every such action is logged. The platform currently operates under Institutional Review Board oversight for its active study, and is run as an independent research project — see the privacy policy for how participant data is handled.
Want your system evaluated?
If you build an institutional knowledge-retrieval tool, Useably can benchmark it on this protocol — controlled cases, real users, validated instruments, and adjudicated accuracy.