Research agenda¶
The agenda turns gaps into research decisions. It prioritizes questions that could weaken a claim rather than only confirm it.
Priorities¶
| Priority | Falsifiable question | Minimum design | Progress signal | Risk |
|---|---|---|---|---|
| P1 · DI field | Does DI have its own community, methods and cumulative evidence? | preregistered systematic review and bibliometric map | reproducible criteria and search | mistaking sample absence for real absence |
| P1 · Human–AI | When does human+AI outperform both human and AI alone? | task-level comparison with allocation and defined outcomes | effect, heterogeneity, harm and cost | automation complacency or deskilling |
| P1 · Control | Which controls change behavior when an agent can act? | override, stop and reversal tests | intervention and recovery rates | ceremonial control |
| P2 · PULSE | Does PDAMR reduce material omissions versus the current process? | preregistered pilot with comparator | completeness, latency, burden and ex-ante quality | bureaucracy and gaming |
| P2 · Learning | Does the longitudinal record improve calibration or only documentation? | repeated comparable decisions and feedback | calibration, review and observable change | outcome bias |
| P2 · Decision Experience | Which combination of conversation, visualization and workflow works by task? | factorial or sequential study | accuracy, comprehension, time and accessibility | fluency mistaken for quality |
| P3 · Metaphors | Do biological lenses generate useful, distinguishable predictions? | observable hypotheses without metaphorical language | a prediction that can fail | anthropomorphism |
Minimum contracts¶
Every study in the program should state:
- decision and unit of analysis;
- population, context and exclusions;
- relevant comparator;
- outcomes and guardrails separated from adoption;
- measurement period and missing data;
- human authority and incident protocol;
- heterogeneity analysis and adverse results;
- materials, system version and deviations;
- criterion for strengthening, maintaining or weakening each claim.
Recommended first pilot¶
The B2B case can test instrumentation, but it should not move directly from simulated data to an effectiveness claim. The next decision is whether an owner, an authorized operation and a viable comparator exist. Otherwise, the case remains stopped.
A safer initial test would measure whether PDAMR and the Decision Brief reduce process omissions in historical or simulated decisions without allowing AI to execute actions. That result would inform usability and coverage, not causal sales impact.
Review cadence¶
- quarterly: standards, frameworks and future signals;
- for each candidate release: sources, claims and translations;
- after each study: null and adverse results, and deviations;
- annually: reproducible state-of-the-art search;
- exceptionally: regulatory change, material incident or new contradictory evidence.
Success criterion¶
The program advances when a claim becomes more precise even if its strength decreases. More pages or more portal usage do not demonstrate better decisions.