Sybil: A Validated Deep Learning Model to Predict Future Lung Cancer Risk from a Single Low-Dose Chest Computed Tomography

Peter G. Mikhael, Jeremy Wohlwend, Adam Yala, Ludvig Karstens, Justin Xiang, Angelo K. Takigami, Patrick P. Bourgouin, Puiyee Chan, Sofiane Mrah, Wael Amayri, Yu Hsiang Juan, Cheng-Ta Yang, Yung Liang Wan, Gigin Lin, Lecia V. Sequist*, Florian J. Fintelmann, Regina Barzilay

*Corresponding author for this work

Research output: Contribution to journalJournal Article peer-review

49 Scopus citations


PURPOSE: Low-dose computed tomography (LDCT) for lung cancer screening is effective, although most eligible people are not being screened. Tools that provide personalized future cancer risk assessment could focus approaches toward those most likely to benefit. We hypothesized that a deep learning model assessing the entire volumetric LDCT data could be built to predict individual risk without requiring additional demographic or clinical data.

METHODS: We developed a model called Sybil using LDCTs from the National Lung Screening Trial (NLST). Sybil requires only one LDCT and does not require clinical data or radiologist annotations; it can run in real time in the background on a radiology reading station. Sybil was validated on three independent data sets: a heldout set of 6,282 LDCTs from NLST participants, 8,821 LDCTs from Massachusetts General Hospital (MGH), and 12,280 LDCTs from Chang Gung Memorial Hospital (CGMH, which included people with a range of smoking history including nonsmokers).

RESULTS: Sybil achieved area under the receiver-operator curves for lung cancer prediction at 1 year of 0.92 (95% CI, 0.88 to 0.95) on NLST, 0.86 (95% CI, 0.82 to 0.90) on MGH, and 0.94 (95% CI, 0.91 to 1.00) on CGMH external validation sets. Concordance indices over 6 years were 0.75 (95% CI, 0.72 to 0.78), 0.81 (95% CI, 0.77 to 0.85), and 0.80 (95% CI, 0.75 to 0.86) for NLST, MGH, and CGMH, respectively.

CONCLUSION: Sybil can accurately predict an individual's future lung cancer risk from a single LDCT scan to further enable personalized screening. Future study is required to understand Sybil's clinical applications. Our model and annotations are publicly available.

UNLABELLED: [Media: see text].

Original languageEnglish
Pages (from-to)2191-2200
Number of pages10
JournalJournal of Clinical Oncology
Issue number12
StatePublished - 20 04 2023

Bibliographical note

Publisher Copyright:
© American Society of Clinical Oncology.


  • Humans
  • Lung Neoplasms/diagnostic imaging
  • Deep Learning
  • Early Detection of Cancer/methods
  • Tomography, X-Ray Computed
  • Lung
  • Mass Screening/methods


Dive into the research topics of 'Sybil: A Validated Deep Learning Model to Predict Future Lung Cancer Risk from a Single Low-Dose Chest Computed Tomography'. Together they form a unique fingerprint.

Cite this