Benchmark knowledge-retrieval systems on evidence, not demos.
Useably is a research-grade task-simulation platform for evaluating institutional knowledge-retrieval systems — controlled case assignment, timed task completion, head-to-head condition comparison, validated workload and usability instruments, and reviewer adjudication of source-retrieval accuracy.
See Useably in action
A guided tour of the researcher workspace — cohorts, rosters, adjudication, analysis, and exports — showing how a study runs day to day.
Retrieval tools are sold on demos. Buyers have no controlled way to compare them.
When a clinician needs an answer fast, the question isn’t whether a system cansurface a document — it’s whether the right person finds the right answer, quickly, under real cognitive load, and whether the source actually supports it. Useably turns that into a measurable, repeatable protocol: the same cases, the same timing rules, and the same validated instruments applied to every system under test.
A server-authoritative protocol engine
Every participant walks an identical, controlled sequence — so the data answers the question you actually asked.
Within-subject crossover
Each participant completes the task set under both conditions, with block-of-4 randomized ordering. Comparisons are within-person, not between cohorts — removing the confound that sinks most tool bake-offs.
Controlled case assignment
A versioned, per-cohort case bank drives every participant through the same clinical scenarios in a fixed protocol order, so differences reflect the system, not the questions.
Server-authoritative timing
Task duration is measured and capped server-side (300s), not trusted from the client. Time-to-answer is a first-class, tamper-resistant outcome.
Append-only, audited data
Submissions are append-only with database-enforced integrity triggers; corrections are tracked, never silently overwritten. The dataset is defensible under review.
Validated instruments, not vibes
Outcomes are scored with published, peer-reviewed instruments — the same scales the human-factors literature already trusts — embedded directly in the task protocol.
System Usability Scale
The validated 10-item usability benchmark, administered per condition so each system gets a comparable, published-scale score.
Task Load Index
Per-task workload across six dimensions — mental demand, effort, frustration, and more — capturing the cognitive cost of getting to an answer.
mHealth App Usability Questionnaire
An additional validated usability instrument for the AI-assisted condition, for a fuller picture of the interactive retrieval experience.
Reviewer adjudication of source-retrieval accuracy
Speed and self-reported usability aren’t enough. A trained reviewer adjudicates whether the source a participant retrieved actually answers the clinical question — turning “found a document” into a graded measure of correct retrieval. Adjudications run through a dedicated queue, are append-only, and are audited like every other record.
Run the whole study from one place
The participant protocol is half the platform. The other half is the operational surface a real study needs — built in, not bolted on.
Live dashboard & roster
Enrollment, protocol progress, and instrument scores update as participants work — per cohort and per participant, down to individual submissions.
Cohort scheduling & invitations
Each cohort gets its own study window, case bank, and enrollment rules. Invite participants with one-time registration links, emailed sign-in codes, or issued access codes.
Adjudication queue
Reviewers grade source-retrieval accuracy in a dedicated queue, with every decision recorded in the audit log.
Built-in statistical analysis
Within-subject comparisons and instrument scoring computed in-app and validated against scipy — no export-to-notebook round trip to see where the study stands.
Privacy safeguards
Free-text responses are automatically screened for identifying information and masked until a reviewer clears or redacts them. Every reveal and redaction is logged.
Audited exports
Blinded or identified CSV/JSON exports for downstream analysis, each one recorded in a full audit trail.
In use at major research universities
Useably runs live, IRB-governed studies at major universities — including a within-subject, four-task crossover comparing AI-assisted retrieval against legacy institutional knowledge access in perioperative care, with per-task workload and per-condition usability scoring. It isn’t a prototype: it’s a protocol engine already collecting defensible data under ethics oversight.
Put your retrieval system to the test.
If you build an institutional knowledge-retrieval tool and want it evaluated on a neutral, validated harness — controlled cases, real users, adjudicated accuracy — start a conversation.