Neutral evaluation harness

Benchmark knowledge-retrieval systems on evidence, not demos.

Useably is a research-grade task-simulation platform for evaluating institutional knowledge-retrieval systems — controlled case assignment, timed task completion, head-to-head condition comparison, validated workload and usability instruments, and reviewer adjudication of source-retrieval accuracy.

Retrieval benchmark
Under testLegacy
Time to answer
Cognitive workload
Perceived usability
Source-retrieval accuracyReviewer-adjudicated, per case
Validated
SUS · TLX · MAUQ
Major research universitiesIRB-governedWithin-subject crossoverSUS · NASA-TLX · MAUQAppend-only audit
Platform walkthrough

See Useably in action

A guided tour of the researcher workspace — cohorts, rosters, adjudication, analysis, and exports — showing how a study runs day to day.

The gap

Retrieval tools are sold on demos. Buyers have no controlled way to compare them.

When a clinician needs an answer fast, the question isn’t whether a system cansurface a document — it’s whether the right person finds the right answer, quickly, under real cognitive load, and whether the source actually supports it. Useably turns that into a measurable, repeatable protocol: the same cases, the same timing rules, and the same validated instruments applied to every system under test.

How it works

A server-authoritative protocol engine

Every participant walks an identical, controlled sequence — so the data answers the question you actually asked.

Within-subject crossover

Each participant completes the task set under both conditions, with block-of-4 randomized ordering. Comparisons are within-person, not between cohorts — removing the confound that sinks most tool bake-offs.

Controlled case assignment

A versioned, per-cohort case bank drives every participant through the same clinical scenarios in a fixed protocol order, so differences reflect the system, not the questions.

Server-authoritative timing

Task duration is measured and capped server-side (300s), not trusted from the client. Time-to-answer is a first-class, tamper-resistant outcome.

Append-only, audited data

Submissions are append-only with database-enforced integrity triggers; corrections are tracked, never silently overwritten. The dataset is defensible under review.

Measurement

Validated instruments, not vibes

Outcomes are scored with published, peer-reviewed instruments — the same scales the human-factors literature already trusts — embedded directly in the task protocol.

SUS

System Usability Scale

The validated 10-item usability benchmark, administered per condition so each system gets a comparable, published-scale score.

NASA-TLX

Task Load Index

Per-task workload across six dimensions — mental demand, effort, frustration, and more — capturing the cognitive cost of getting to an answer.

MAUQ

mHealth App Usability Questionnaire

An additional validated usability instrument for the AI-assisted condition, for a fuller picture of the interactive retrieval experience.

Reviewer adjudication of source-retrieval accuracy

Speed and self-reported usability aren’t enough. A trained reviewer adjudicates whether the source a participant retrieved actually answers the clinical question — turning “found a document” into a graded measure of correct retrieval. Adjudications run through a dedicated queue, are append-only, and are audited like every other record.

Time to answerServer-measured, capped
Perceived workloadNASA-TLX, per task
Perceived usabilitySUS, per condition
Retrieval correctnessReviewer-adjudicated
Researcher workspace

Run the whole study from one place

The participant protocol is half the platform. The other half is the operational surface a real study needs — built in, not bolted on.

Live dashboard & roster

Enrollment, protocol progress, and instrument scores update as participants work — per cohort and per participant, down to individual submissions.

Cohort scheduling & invitations

Each cohort gets its own study window, case bank, and enrollment rules. Invite participants with one-time registration links, emailed sign-in codes, or issued access codes.

Adjudication queue

Reviewers grade source-retrieval accuracy in a dedicated queue, with every decision recorded in the audit log.

Built-in statistical analysis

Within-subject comparisons and instrument scoring computed in-app and validated against scipy — no export-to-notebook round trip to see where the study stands.

Privacy safeguards

Free-text responses are automatically screened for identifying information and masked until a reviewer clears or redacts them. Every reveal and redaction is logged.

Audited exports

Blinded or identified CSV/JSON exports for downstream analysis, each one recorded in a full audit trail.

Live, IRB-governed research

In use at major research universities

Useably runs live, IRB-governed studies at major universities — including a within-subject, four-task crossover comparing AI-assisted retrieval against legacy institutional knowledge access in perioperative care, with per-task workload and per-condition usability scoring. It isn’t a prototype: it’s a protocol engine already collecting defensible data under ethics oversight.

Put your retrieval system to the test.

If you build an institutional knowledge-retrieval tool and want it evaluated on a neutral, validated harness — controlled cases, real users, adjudicated accuracy — start a conversation.