

Most conversations about closing the lung cancer screening gap focus on uptake, and for good reason. Just 18.7 percent of eligible Americans were up to date with low-dose CT screening in 2024, roughly one in five, according to the American Cancer Society. The standard remedies follow: better outreach, more reminders, smarter scheduling, friction removal in the EHR. All of that is worth doing.
But it treats the eligibility criteria themselves as fixed, as the outer boundary of the addressable problem. That assumption is the deeper issue. The hardest part of lung cancer detection is not getting eligible people screened. It is that most people who go on to develop lung cancer were never eligible in the first place. For anyone building or buying health technology in this space, that distinction is not academic. It determines what problem your tools are allowed to solve.
The pack-year filter is a data model, and it is the wrong one.

Current screening eligibility runs through a smoking history filter: roughly twenty or more pack-years, within a defined age band. That criterion exists for a defensible reason. It was derived from the trial populations that established LDCT’s mortality benefit, and within those populations it concentrates screening on the highest-yield candidates.
The trouble is what it excludes, and the size of that exclusion is now well documented. In the Boston Lung Cancer Study, an analysis of more than 7,000 patients who had lung cancer, only 46 percent met current USPSTF screening criteria. The majority would not have qualified. And the excluded group is not random: it skews toward never-smokers, lower-pack-year smokers, women, and people whose risk comes from environmental exposure rather than tobacco. The equity data is just as stark. In the Southern Community Cohort Study, among smokers who developed lung cancer, only 32 percent of African American patients were eligible for screening, compared with 56 percent of white patients (Aldrich et al., JAMA Oncology, 2019). A fixed pack-year cutoff systematically excludes patients who are, in clinical reality, at high risk.
Frame it the way a machine learning team would. Pack-years is a single, self-reported, hard-thresholded feature standing in for a complex, multifactorial risk. It discards exposure history that is not tobacco. It has no awareness of air quality, occupational exposure, family history, or biomarker signals. And it is applied as a binary gate rather than a continuous, calibrated score. If a vendor proposed a risk model with those properties today, it would not survive technical review. We tolerate it in screening eligibility only because it is incumbent.
Case-finding is the architectural shift, and the guidelines just endorsed it.
The meaningful move is from population screening gated by a single criterion to risk-based case-finding: systematically identifying higher-risk individuals using whatever predictive signal is available, then routing them toward appropriate evaluation. Screening starts at the scanner and asks who qualifies. Case-finding starts upstream, in routinely collected data and accessible signals, and asks whom we are missing.
This is no longer a fringe position. The 2026 GOLD report, which sets global strategy for chronic obstructive pulmonary disease, added a formal case-finding algorithm and a framework for understanding underdiagnosis, alongside an entirely new chapter on artificial intelligence and emerging technologies, the first time AI has earned a dedicated chapter in that guideline. Its stance is pragmatic rather than promotional: emerging tools may help identify the undiagnosed, but generalizability and bias must be handled carefully before deployment. The signal for health IT is that the most influential body in respiratory medicine has stopped treating technology-assisted case-finding as speculative and started treating it as a workflow to be governed.
That reframes the build target. The question is no longer only how to get eligible patients scanned. It becomes how to identify elevated-risk patients the criteria miss, and route them responsibly. That is a risk-stratification and care orchestration problem, which is to say a software problem, sitting on top of data most systems already hold.
Three implications follow for the people designing these tools.
First, the input layer has to widen beyond smoking history. A case-finding model worth deploying needs to ingest environmental and occupational exposure, demographic risk factors, relevant biomarkers, and longitudinal signals from routine encounters, not a single pack-year field. The data engineering for this is harder than the model, and underrated as a barrier.
Second, representativeness is a safety property, not a fairness afterthought. The populations most underserved by pack-year criteria- never-smokers, environmentally exposed patients, and groups underrepresented in training data, are precisely the ones a case-finding tool exists to catch. A model trained predominantly on Western smoking cohorts will be miscalibrated for them, which means it will fail exactly where it is most needed. Building representative datasets is not compliance overhead. It is the core engineering challenge of the field.
Third, the workflow has to respect the line between flagging risk and making a diagnosis. A case-finding tool earns its place by helping a clinician decide who deserves a closer look, not by claiming to detect disease. How a risk signal surfaces in the EHR, who acts on it, and how downstream evaluation is orchestrated is where most of the real-world value, and most of the real-world risk, actually live.
The strategic picture is clear enough to act on now. The biggest lever in lung cancer detection is not better adherence to an eligibility rule that misses most cases. It is changing the front door, from a smoking-history checkbox to a risk-based, data-driven case-finding workflow that can see the patients the old model was never built to see. The guidelines are moving. The evidence is moving. The infrastructure is the part that has to catch up.
About Komal Sharma
Komal Sharma is the founder of Mednoia, an early-stage company developing AI-based approaches to respiratory risk stratification, and a director of product at UGenome. She holds an MBA and an MSc in biotechnology, and her work focuses on multimodal AI diagnostic tools for respiratory health. She writes in a personal capacity.
About Pratik Sharma
Pratik Sharma is an AI/ML engineer specializing in multimodal machine learning and data-fusion architectures, and a Cloud Solution Architect at Microsoft. He is pursuing an M.Eng at the University of Illinois Chicago and serves on the advisory board of Mednoia. He writes in a personal capacity.
References
1. American Cancer Society, 2025 Lung Cancer Data: Only 1 in 5 Eligible Adults in U.S. Screened for Lung Cancer; 62,000 Lives Over 5 Years Could be Saved if All Eligible Screened. pressroom.cancer.org/2025- lung-cancer-data
2. Boston Lung Cancer Study, “Assessing Lung Cancer Screening Eligibility of Patients With Lung Cancer” (7,186 cases; 46.1% met USPSTF criteria), 2025. PubMed 39864777
3. Aldrich MC et al., “Evaluation of USPSTF Lung Cancer Screening Guidelines Among African American Adult Smokers,” JAMA Oncology, 2019 (32% vs 56% eligibility; Southern Community Cohort Study). 4. GOLD 2026 Report: Chapter 6 (Artificial Intelligence & Emerging Technologies); new case-finding algorithm and underdiagnosis framework. goldcopd.org
