Hybrid human-algorithm systems to support situated clinical decision-making with an empirical study in dispatch of pre-hospital critical care
Abstract
Healthcare is expected to benefit from advances in Artificial Intelligence (AI). But the pace of translation into practice has lagged behind expectations. This AI Chasm in healthcare may be an expression of insufficient attention being given to how humans and such machines need to work well together. This thesis has a specific contextual motivation in the work of the pre-hospital critical care teams of the Welsh National Health Service. These teams use a continual and intensive triage process to arrive at decisions on activating their specialist resources via an air ambulance service or a fleet of rapid response road vehicles. The work here investigates hybrid human-machine combination in decision-making and makes four contributions. First, an approach to designing for hybridity that lays out the dimensions of human-machine combination drawn from an extensive review of relevant theory and of recent studies. Second, a detailed account of how the situatedness of human action in the given clinical context is shown to affect the most appropriate design and development of an algorithmic support system. In the third contribution, an empirical study, we show how clinicians in this context can make positive use of algorithmic support, assuring that under-reliance is not a barrier to securing benefit. Specifically, we run a three-arm study in which users are exposed to support from a human expert score on the one hand, an algorithmic score on another and a no-score condition on the third arm. The presence of either score condition has a statistically significant positive influence on critical care triage decisions (p < 0.05) compared to having no assistive score. But the influence of whether the score is presented as algorithmic or human in provenance is not statistically significant (p ≈ 0.8). User satisfaction withour interface design is entirely positive, while satisfaction with the assistive value of the proposed tool is largely positive. There is a positive skew to the predisposition of hub staff to embrace AI technologies. Nevertheless, predisposition towards AI systems does not appear to predict satisfaction with the proposed algorithmic tool. The fourth contribution, a further study, shows how over-reliance is countered by the highly skilled participants while their ability to discriminate between the quality levels of different machines is maintained. In this study, support tools with three very different performance levels are presented to users. The approximate sensitivities of the tools are 80%, 63% and 40%. The differences between decision results obtained with each tool are not statistically significant (p > 0.7 in all paired tests). However, when comparing results here with the no-score treatment from the earlier study, we find that the influence of each of the two higher quality tools is to improve decision performance significantly (p < 0.05 in each case). In a striking result, we also find that subjective judgements achieve statistically significant discrimination between the different tools, ranking them in order of sensitivity (p < 0.05 in all paired tests). Critically, we establish that a realistic baseline performance for a Random Forest classifier on this task is 63%. And at this tool performance level, a tool appears to have a significant positive influence on decisions. Importantly, from our results, if a tool were to have significantly lower performance, it does not follow that the quality of decisions would decline below that of a decision-maker without support. In other words, poor tool performances is not found to be detectably unsafe. Human decision-makers do express the greatest concerns, however, at the high False Negative Rate of a poorly-performing machine. The thesis identifies that practical development on a larger dataset is not only possible, but likely to yield measurable benefit for the service and the public it serves. It concludes by drawing together the significance of studying situated action as a precondition and situated evaluation as a success factor in order to put human combination with machines into real effect for human benefit.