← ireadpaper · 顶刊中的公共政策研究

诊断医生错误:低价值医疗服务的机器学习研究方法

Diagnosing Physician Error: A Machine Learning Approach to Low-Value Health Care
Quarterly Journal of Economics · 2021 · Sendhil Mullainathan、Ziad Obermeyer

中文摘要

医生诊断心脏病发作的有效性如何?为回答这一问题,我们将医生的检查决策与机器学习风险模型进行比较。当二者出现偏离时,我们利用实际健康结果数据判断算法或医生究竟哪一方正确。我们发现,医生存在过度检查:即使某些检查可预见地毫无用处,医生仍会实施这些检查。与此同时,医生也存在检查不足:许多被预测为高风险的患者未接受检查,随后以较高比率发生不良健康事件(包括死亡)。一项利用班次间检查差异的自然实验验证了这些发现:增加检查能够改善健康状况并降低死亡率,但仅对被算法标记为高风险的患者有效。过度检查与检查不足并存的现象难以单凭激励因素加以解释,反而表明其中存在错误。我们就这些错误背后的心理机制提供了提示性证据:(i)医生使用的风险模型过于简单,表明其存在有限理性;(ii)他们对显著信息赋予过高权重;(iii)他们对具有心脏病发作代表性或典型特征的症状赋予过高权重。总之,这些结果表明,医疗服务模型与政策不仅需要纳入医生激励,还需要考虑医生的错误。

Abstract

How effective are physicians at diagnosing heart attacks? To answer this question, we contrast physician testing decisions with a machine learning model of risk. When the two deviate, we use actual health outcome data to judge whether the algorithm or the physician was right. We find physicians over-test: tests that are predictably useless are still performed. At the same time, physicians also under-test: many predicted high-risk patients are untested and then suffer adverse health events (including death) at high rates. A natural experiment using shift-to-shift testing variation confirms these findings: increasing testing improves health and reduces mortality, but only for patients flagged as high-risk by the algorithm. The simultaneous existence of over- and under-testing cannot easily be explained by incentives alone, and instead suggests errors. We provide suggestive evidence on the psychology behind these errors:(i) physicians use too simple a model of risk, suggesting bounded rationality; (ii) they over-weight salient information; and (iii) they over-weight symptoms that are representative or stereotypical of heart attack. Together, these results suggest the need for health care models and policies to incorporate not just physician incentives, but also physician mistakes.
在 ireadpaper 查看全部 →