Skip to content

Research agenda

The agenda turns gaps into research decisions. It prioritizes questions that could weaken a claim rather than only confirm it.

Priorities

Priority Falsifiable question Minimum design Progress signal Risk
P1 · DI field Does DI have its own community, methods and cumulative evidence? preregistered systematic review and bibliometric map reproducible criteria and search mistaking sample absence for real absence
P1 · Human–AI When does human+AI outperform both human and AI alone? task-level comparison with allocation and defined outcomes effect, heterogeneity, harm and cost automation complacency or deskilling
P1 · Control Which controls change behavior when an agent can act? override, stop and reversal tests intervention and recovery rates ceremonial control
P2 · PULSE Does PDAMR reduce material omissions versus the current process? preregistered pilot with comparator completeness, latency, burden and ex-ante quality bureaucracy and gaming
P2 · Learning Does the longitudinal record improve calibration or only documentation? repeated comparable decisions and feedback calibration, review and observable change outcome bias
P2 · Decision Experience Which combination of conversation, visualization and workflow works by task? factorial or sequential study accuracy, comprehension, time and accessibility fluency mistaken for quality
P3 · Metaphors Do biological lenses generate useful, distinguishable predictions? observable hypotheses without metaphorical language a prediction that can fail anthropomorphism

Minimum contracts

Every study in the program should state:

  • decision and unit of analysis;
  • population, context and exclusions;
  • relevant comparator;
  • outcomes and guardrails separated from adoption;
  • measurement period and missing data;
  • human authority and incident protocol;
  • heterogeneity analysis and adverse results;
  • materials, system version and deviations;
  • criterion for strengthening, maintaining or weakening each claim.

The B2B case can test instrumentation, but it should not move directly from simulated data to an effectiveness claim. The next decision is whether an owner, an authorized operation and a viable comparator exist. Otherwise, the case remains stopped.

A safer initial test would measure whether PDAMR and the Decision Brief reduce process omissions in historical or simulated decisions without allowing AI to execute actions. That result would inform usability and coverage, not causal sales impact.

Review cadence

  • quarterly: standards, frameworks and future signals;
  • for each candidate release: sources, claims and translations;
  • after each study: null and adverse results, and deviations;
  • annually: reproducible state-of-the-art search;
  • exceptionally: regulatory change, material incident or new contradictory evidence.

Success criterion

The program advances when a claim becomes more precise even if its strength decreases. More pages or more portal usage do not demonstrate better decisions.