Study Results

v1 derivation cohort — complete

These are the published results of the first (v1) derivation cohort, which is closed and fully analyzed. Results for the current validation cohort remain embargoed until the pre-registered analysis checkpoints.

50
Physician sessions
626
Forced-choice comparisons
~69%
Agreement with the index
Z ≈ 3.4
Signal vs. chance
Preprint — not peer reviewed

Read the full v1 derivation-cohort manuscript

Do objective circadian-alignment scores track emergency physicians’ own sense of a “good” schedule? Full methods, results, limitations, and declarations for the completed v1 cohort (50 physicians, 626 comparisons). The scoring model itself remains proprietary and is not disclosed in the manuscript.

Download PDF613 KB

What the v1 cohort showed

Fifty practicing emergency physicians completed sessions of blinded head-to-head schedule comparisons — 626 graded forced-choice judgments in total. Participants never saw a score, ranking, or prediction.

Physician preference aligned with the schedule the index ranked higher on approximately 69% of graded pairs. The directional signal was statistically distinguishable from chance (Z ≈ 3.4), confirming that the index was measuring something physicians genuinely respond to.

The result did not cross the pre-specified early-stopping efficacy boundary, so v1 was treated exactly as pre-registered: as derivation evidence rather than a validation claim. Nothing on this page should be read as a validated performance estimate.

What we learned and changed

A pre-registered exploratory diagnostic on the same 626 comparisons identified which fatigue and circadian axes the index was under-weighting or missing entirely. The revisions were material enough that the v1 sessions were permanently converted to derivation data and excluded from validation statistics.

Under the governance rule set before data collection, revising the index converted the v1 sessions into derivation data. They were used to refit the scoring weights and are permanently excluded from every validation statistic. The revised index — CAI v2.1 — is now frozen, and validation restarted from zero with a fresh, independent cohort.

Rank correlation between the v1 and v2.1 indices is high (ρ ≈ 0.95–0.97), but the changes are material enough that v1 responses cannot be pooled with validation responses.

Transparency note

This page reports high-level findings from the completed v1 derivation cohort. It does not disclose the CAI scoring model, feature weights, normalization, formulas, or any other proprietary implementation details.

Current cohort — enrollment only

The live figures below describe the ongoing CAI v2.1 validation cohort: enrollment and data quality only. Agreement between physician preference and the index is sealed for this cohort until the pre-specified checkpoints are reached.

CAI v2.1 cohort status

Live · interim

These numbers update automatically as physicians take part. They describe how good the data is, not whether the algorithm was right. The main result — how closely the algorithm matches physician preference — stays sealed until we reach our planned checkpoints, so that nobody (including us) can peek and cherry-pick a favourable moment to announce it.