Impact of a deep learning assistant on the histopathologic classification of liver cancer

Amirhossein Kiani; Bora Uyumazturk; Pranav Rajpurkar; Alex Wang; Rebecca Gao; Erik Jones; Yifan Yu; Curtis P. Langlotz; Robyn L. Ball; Thomas J. Montine; Brock A. Martin; Gerald J. Berry; Michael G. Ozawa; Florette K. Hazard; Ryanne A. Brown; Simon B. Chen; Mona Wood; Libby S. Allard; Lourdes Ylagan; Andrew Y. Ng; Jeanne Shen

doi:10.1038/s41746-020-0232-8

Impact of a deep learning assistant on the histopathologic classification of liver cancer

Amirhossein Kiani, Bora Uyumazturk, Pranav Rajpurkar, Alex Wang, Rebecca Gao, Erik Jones, Yifan Yu, Curtis P. Langlotz, Robyn L. Ball, Thomas J. Montine, Brock A. Martin, Gerald J. Berry, Michael G. Ozawa, Florette K. Hazard, Ryanne A. Brown, Simon B. Chen, Mona Wood, Libby S. Allard, Lourdes Ylagan, Andrew Y. NgJeanne Shen

Research output: Contribution to journal › Article › peer-review

151 Scopus citations

Abstract

Artificial intelligence (AI) algorithms continue to rival human performance on a variety of clinical tasks, while their actual impact on human diagnosticians, when incorporated into clinical workflows, remains relatively unexplored. In this study, we developed a deep learning-based assistant to help pathologists differentiate between two subtypes of primary liver cancer, hepatocellular carcinoma and cholangiocarcinoma, on hematoxylin and eosin-stained whole-slide images (WSI), and evaluated its effect on the diagnostic performance of 11 pathologists with varying levels of expertise. Our model achieved accuracies of 0.885 on a validation set of 26 WSI, and 0.842 on an independent test set of 80 WSI. Although use of the assistant did not change the mean accuracy of the 11 pathologists (p = 0.184, OR = 1.281), it significantly improved the accuracy (p = 0.045, OR = 1.499) of a subset of nine pathologists who fell within well-defined experience levels (GI subspecialists, non-GI subspecialists, and trainees). In the assisted state, model accuracy significantly impacted the diagnostic decisions of all 11 pathologists. As expected, when the model’s prediction was correct, assistance significantly improved accuracy (p = 0.000, OR = 4.289), whereas when the model’s prediction was incorrect, assistance significantly decreased accuracy (p = 0.000, OR = 0.253), with both effects holding across all pathologist experience levels and case difficulty levels. Our results highlight the challenges of translating AI models into the clinical setting, and emphasize the importance of taking into account potential unintended negative consequences of model assistance when designing and testing medical AI-assistance tools.

Original language	English (US)
Article number	23
Journal	npj Digital Medicine
Volume	3
Issue number	1
DOIs	https://doi.org/10.1038/s41746-020-0232-8
State	Published - Dec 1 2020
Externally published	Yes

ASJC Scopus subject areas

Medicine (miscellaneous)
Health Informatics
Computer Science Applications
Health Information Management

Access to Document

10.1038/s41746-020-0232-8

Cite this

Kiani, A., Uyumazturk, B., Rajpurkar, P., Wang, A., Gao, R., Jones, E., Yu, Y., Langlotz, C. P., Ball, R. L., Montine, T. J., Martin, B. A., Berry, G. J., Ozawa, M. G., Hazard, F. K., Brown, R. A., Chen, S. B., Wood, M., Allard, L. S., Ylagan, L., ... Shen, J. (2020). Impact of a deep learning assistant on the histopathologic classification of liver cancer. npj Digital Medicine, 3(1), Article 23. https://doi.org/10.1038/s41746-020-0232-8

Kiani, A, Uyumazturk, B, Rajpurkar, P, Wang, A, Gao, R, Jones, E, Yu, Y, Langlotz, CP, Ball, RL, Montine, TJ, Martin, BA, Berry, GJ, Ozawa, MG, Hazard, FK, Brown, RA, Chen, SB, Wood, M, Allard, LS, Ylagan, L, Ng, AY & Shen, J 2020, 'Impact of a deep learning assistant on the histopathologic classification of liver cancer', npj Digital Medicine, vol. 3, no. 1, 23. https://doi.org/10.1038/s41746-020-0232-8

@article{4c0ab5c16ba84dc9aad65afa904e688f,

title = "Impact of a deep learning assistant on the histopathologic classification of liver cancer",

abstract = "Artificial intelligence (AI) algorithms continue to rival human performance on a variety of clinical tasks, while their actual impact on human diagnosticians, when incorporated into clinical workflows, remains relatively unexplored. In this study, we developed a deep learning-based assistant to help pathologists differentiate between two subtypes of primary liver cancer, hepatocellular carcinoma and cholangiocarcinoma, on hematoxylin and eosin-stained whole-slide images (WSI), and evaluated its effect on the diagnostic performance of 11 pathologists with varying levels of expertise. Our model achieved accuracies of 0.885 on a validation set of 26 WSI, and 0.842 on an independent test set of 80 WSI. Although use of the assistant did not change the mean accuracy of the 11 pathologists (p = 0.184, OR = 1.281), it significantly improved the accuracy (p = 0.045, OR = 1.499) of a subset of nine pathologists who fell within well-defined experience levels (GI subspecialists, non-GI subspecialists, and trainees). In the assisted state, model accuracy significantly impacted the diagnostic decisions of all 11 pathologists. As expected, when the model{\textquoteright}s prediction was correct, assistance significantly improved accuracy (p = 0.000, OR = 4.289), whereas when the model{\textquoteright}s prediction was incorrect, assistance significantly decreased accuracy (p = 0.000, OR = 0.253), with both effects holding across all pathologist experience levels and case difficulty levels. Our results highlight the challenges of translating AI models into the clinical setting, and emphasize the importance of taking into account potential unintended negative consequences of model assistance when designing and testing medical AI-assistance tools.",

author = "Amirhossein Kiani and Bora Uyumazturk and Pranav Rajpurkar and Alex Wang and Rebecca Gao and Erik Jones and Yifan Yu and Langlotz, {Curtis P.} and Ball, {Robyn L.} and Montine, {Thomas J.} and Martin, {Brock A.} and Berry, {Gerald J.} and Ozawa, {Michael G.} and Hazard, {Florette K.} and Brown, {Ryanne A.} and Chen, {Simon B.} and Mona Wood and Allard, {Libby S.} and Lourdes Ylagan and Ng, {Andrew Y.} and Jeanne Shen",

note = "Publisher Copyright: {\textcopyright} 2020, The Author(s).",

year = "2020",

month = dec,

day = "1",

doi = "10.1038/s41746-020-0232-8",

language = "English (US)",

volume = "3",

journal = "npj Digital Medicine",

issn = "2398-6352",

publisher = "Nature Publishing Group",

number = "1",

}

TY - JOUR

T1 - Impact of a deep learning assistant on the histopathologic classification of liver cancer

AU - Kiani, Amirhossein

AU - Uyumazturk, Bora

AU - Rajpurkar, Pranav

AU - Wang, Alex

AU - Gao, Rebecca

AU - Jones, Erik

AU - Yu, Yifan

AU - Langlotz, Curtis P.

AU - Ball, Robyn L.

AU - Montine, Thomas J.

AU - Martin, Brock A.

AU - Berry, Gerald J.

AU - Ozawa, Michael G.

AU - Hazard, Florette K.

AU - Brown, Ryanne A.

AU - Chen, Simon B.

AU - Wood, Mona

AU - Allard, Libby S.

AU - Ylagan, Lourdes

AU - Ng, Andrew Y.

AU - Shen, Jeanne

PY - 2020/12/1

Y1 - 2020/12/1

N2 - Artificial intelligence (AI) algorithms continue to rival human performance on a variety of clinical tasks, while their actual impact on human diagnosticians, when incorporated into clinical workflows, remains relatively unexplored. In this study, we developed a deep learning-based assistant to help pathologists differentiate between two subtypes of primary liver cancer, hepatocellular carcinoma and cholangiocarcinoma, on hematoxylin and eosin-stained whole-slide images (WSI), and evaluated its effect on the diagnostic performance of 11 pathologists with varying levels of expertise. Our model achieved accuracies of 0.885 on a validation set of 26 WSI, and 0.842 on an independent test set of 80 WSI. Although use of the assistant did not change the mean accuracy of the 11 pathologists (p = 0.184, OR = 1.281), it significantly improved the accuracy (p = 0.045, OR = 1.499) of a subset of nine pathologists who fell within well-defined experience levels (GI subspecialists, non-GI subspecialists, and trainees). In the assisted state, model accuracy significantly impacted the diagnostic decisions of all 11 pathologists. As expected, when the model’s prediction was correct, assistance significantly improved accuracy (p = 0.000, OR = 4.289), whereas when the model’s prediction was incorrect, assistance significantly decreased accuracy (p = 0.000, OR = 0.253), with both effects holding across all pathologist experience levels and case difficulty levels. Our results highlight the challenges of translating AI models into the clinical setting, and emphasize the importance of taking into account potential unintended negative consequences of model assistance when designing and testing medical AI-assistance tools.

AB - Artificial intelligence (AI) algorithms continue to rival human performance on a variety of clinical tasks, while their actual impact on human diagnosticians, when incorporated into clinical workflows, remains relatively unexplored. In this study, we developed a deep learning-based assistant to help pathologists differentiate between two subtypes of primary liver cancer, hepatocellular carcinoma and cholangiocarcinoma, on hematoxylin and eosin-stained whole-slide images (WSI), and evaluated its effect on the diagnostic performance of 11 pathologists with varying levels of expertise. Our model achieved accuracies of 0.885 on a validation set of 26 WSI, and 0.842 on an independent test set of 80 WSI. Although use of the assistant did not change the mean accuracy of the 11 pathologists (p = 0.184, OR = 1.281), it significantly improved the accuracy (p = 0.045, OR = 1.499) of a subset of nine pathologists who fell within well-defined experience levels (GI subspecialists, non-GI subspecialists, and trainees). In the assisted state, model accuracy significantly impacted the diagnostic decisions of all 11 pathologists. As expected, when the model’s prediction was correct, assistance significantly improved accuracy (p = 0.000, OR = 4.289), whereas when the model’s prediction was incorrect, assistance significantly decreased accuracy (p = 0.000, OR = 0.253), with both effects holding across all pathologist experience levels and case difficulty levels. Our results highlight the challenges of translating AI models into the clinical setting, and emphasize the importance of taking into account potential unintended negative consequences of model assistance when designing and testing medical AI-assistance tools.

UR - http://www.scopus.com/inward/record.url?scp=85085655108&partnerID=8YFLogxK

UR - http://www.scopus.com/inward/citedby.url?scp=85085655108&partnerID=8YFLogxK

U2 - 10.1038/s41746-020-0232-8

DO - 10.1038/s41746-020-0232-8

M3 - Article

C2 - 32140566

AN - SCOPUS:85085655108

SN - 2398-6352

VL - 3

JO - npj Digital Medicine

JF - npj Digital Medicine

IS - 1

M1 - 23

ER -

Impact of a deep learning assistant on the histopathologic classification of liver cancer

Abstract

ASJC Scopus subject areas

Access to Document

Other files and links

Fingerprint

Cite this