Study Methods
CAI v2.1 cohortDesign of the ongoing CAI v2.1 prospective validation cohort for the CAIER circadian alignment scheduling algorithm.
CAI v2.1 cohort status
Live · interimThese numbers update automatically as physicians take part. They describe how good the data is, not whether the algorithm was right. The main result — how closely the algorithm matches physician preference — stays sealed until we reach our planned checkpoints, so that nobody (including us) can peek and cherry-pick a favourable moment to announce it.
Study type
This is a staged derivation-and-validation study of a proprietary circadian alignment algorithm — the scoring function used to rank candidate emergency-department shift schedules. An initial derivation cohort was collected and closed; its evidence was used to revise the index, and the revised index (CAI v2.1) is now frozen and being validated prospectively on a fresh, independent sample of practicing emergency physicians.
Participants are never shown the algorithm's prediction, ranking, or score during the task. The scheduling engine itself is not the object of validation; it is the generator that exposes the scoring algorithm to human judgment.
Participants
Board-certified or board-eligible emergency physicians (and urgent-care physicians) who complete informed consent and a short demographic questionnaire. No protected health information is collected. Participation is voluntary and uncompensated unless a participant arrives through a compensated recruitment panel.
Primary task
Each session is a short batch of forced-choice comparisons between two candidate monthly schedules; participants pick the schedule they would personally prefer to work. A subset of comparisons are internal quality checks rather than scored items, and are excluded from the primary endpoint. The exact composition of each session is withheld so that responses are not influenced by knowing which comparisons are graded.
Endpoints
Primary: agreement between physician preference and the algorithm's preferred schedule on graded forced-choice comparisons.
Secondary: inter-physician agreement on the same comparisons, reported as a validity checkpoint for the preference task itself.
How inter-physician agreement is measured. Any two physicians shown the same pair of schedules will agree 50% of the time by chance alone, so raw agreement is uninformative. We report a chance-corrected coefficient (item-weighted free-marginal κ — the multi-rater generalization of Bennett's S / Brennan–Prediger κ, per Randolph 2005), where 0 means coin-flip and 1 means perfect agreement. Every schedule pair rated by at least three physicians contributes equally regardless of how many people rated it; a ≥5-rater threshold is reported as a sensitivity analysis, and 95% intervals come from a pair-level bootstrap. Attention probes, low-delta probes and within-session repeats are excluded, and each physician contributes one vote per pair. Self-reported confidence is reported alongside as supporting evidence: agreement is far higher on pairs both physicians rated confidently. Per-pair breakdowns stay embargoed.
All agreement coefficients here are free-marginal κ (a fixed 0.5 chance baseline; Brennan–Prediger / Randolph), not Cohen's or Fleiss' κ. Free-marginal is appropriate because left/right position is randomized and the two schedules in each comparison have no stable cross-pair label, so there is no meaningful marginal distribution to estimate. The same 0.5 baseline applies to the within-person (intra-rater) repeat, the inter-physician agreement, and the per-pair values.
Within-person (intra-rater) reliability (test–retest)
This is a within-person (intra-rater) reliability: it measures how consistently a single physician reproduces their own choice when the identical comparison is re-shown with left/right position swapped. It is reported both as a raw agreement rate and, chance-corrected, as a free-marginal κ on the same 50% baseline used elsewhere.
Because it is a consistency ceiling, no external score — including CAI — can be expected to track stated preference more closely than the same physician's own repeat answer.
Analysis
Agreement is reported using free-marginal κ as the primary statistic, with a fixed 0.5 chance baseline and 95% confidence intervals from a pair-level bootstrap. Because left/right position is randomized and schedule labels have no stable meaning across pairs, free-marginal κ is the appropriate chance- corrected measure; Cohen's or Fleiss' κ are not used. The primary analysis uses each participant's first completed session, so repeat sessions cannot inflate agreement.
Stage 1 uses a locked sequential design with pre-specified interim looks. If the pre-specified efficacy boundary is crossed at the first look, the v2.1 cohort stops and a new index version is derived to add a social-recovery / weekend-synchrony construct; validation then restarts with an independent sample. If the boundary is not crossed, the cohort continues to the next look under the unchanged, frozen v2.1 index. The full stopping rules, boundary values, and conditional restart plan are held on file and available on request.
Versioning & freeze
The scoring algorithm is version-tagged, and every session records the version in effect when responses were collected. The version under validation is frozen: it is not iterated, reweighted, or renormalized while validation data are being collected. Approved candidate schedules are likewise frozen once admitted to the study pool.
These rules are enforced in software, not by convention. The frozen version and the results embargo are stored in an access-controlled configuration record and checked server-side. The study team cannot inspect aggregate human-vs-algorithm agreement until the pre-registered interim threshold is reached. If the index is modified after validation begins, the current cycle is invalidated and validation restarts with a fresh cohort.
Reporting
The study follows the TRIPOD-AI reporting checklist for prediction-model studies using artificial intelligence. It evaluates agreement between algorithm outputs and independent physician preferences without disclosing the proprietary implementation or optimization methodology.
The completed v1 derivation-cohort manuscript is available as a preprint PDF; it reports the same high-level methods and findings withheld here.
Language
All study copy is authored in English. Participants may read the interface in other languages via machine translation at display time; the reading language is recorded per session and non-English sessions are examined in a pre-specified sensitivity analysis.
Ethics & data handling
No identifiable patient data is used. Participant data is limited to demographic and preference responses, stored in an access-controlled administrative environment. This study is investigator-initiated and privately funded; it is designed to be IRB-approval-eligible should external submission become appropriate.
Full methodological detail — including the locked statistical analysis plan, amendment log, and engine change history — is maintained internally and shared on request with reviewers, collaborators, or journals.