통계 방법론 리뷰

category
overviews

이 논문의 관계도

List view
Current PaperStatisticsOverviews · Radiology
Neoadjuvant / Perioperative Immunotherapy연결 근거: neoadjuvant, perioperative, resectable, surgical attrition, pathologic response
  1. Methods방법론 색인 — 어떤 분석을 어떤 논문이 썼나연결어: neoadjuvant, perioperative, resectable
  2. ReviewClinical Oncology 종합 리뷰연결어: neoadjuvant, perioperative, resectable
  3. Yang 2024Treatment patterns and clinical outcomes of patients with resectable non-small cell lung cancer…연결어: neoadjuvant, perioperative, resectable
  4. Bang 2025Perioperative Pembrolizumab for Locally Advanced Thymic Epithelial Tumors: A Single-Arm, Phase…연결어: neoadjuvant, perioperative, pathologic response
  5. Forde 2025Overall Survival with Neoadjuvant Nivolumab plus Chemotherapy in Lung Cancer연결어: neoadjuvant, perioperative, resectable
  6. +44 more
Immunotherapy Response and Resistance연결 근거: progressive disease, immune checkpoint inhibitors, pembrolizumab
  1. ReviewClinical Oncology 종합 리뷰연결어: pseudoprogression, dissociated response, acquired resistance
  2. Methods방법론 색인 — 어떤 분석을 어떤 논문이 썼나연결어: pseudoprogression, acquired resistance, progressive disease
  3. Masse 2024[18F]FDG-PET/CT atypical response patterns to immunotherapy in non-small cell lung cancer patie…연결어: pseudoprogression, dissociated response, immune checkpoint inhibitors
  4. Wang 2022Spatial interplay patterns of cancer nuclei and tumor-infiltrating lymphocytes (TILs) predict c…연결어: progressive disease, immune checkpoint inhibitors, pembrolizumab
  5. ReviewBiomedical Imaging 종합 리뷰연결어: progressive disease, immune checkpoint inhibitors, pembrolizumab
  6. +36 more
Extranodal Extension Imaging연결 근거: extranodal extension, ene, nodal
  1. Henson 2024Criteria for the diagnosis of extranodal extension detected on radiological imaging in head and…연결어: extranodal extension, ene, nodal
  2. Jang 2023Radiologic Extranodal Extension of Metastatic Lymph Nodes in Patients With Non-Small Cell Lung…연결어: extranodal extension, ene, nodal
  3. Jang 2023Radiologic Extranodal Extension of Metastatic Lymph Nodes in Patients With Non-Small Cell Lung…연결어: extranodal extension, ene, nodal
  4. Chohan 2025Pathological extranodal extension in head and neck cancer: A prognostic biomarker with therapeu…연결어: extranodal extension, ene, nodal
  5. ReviewBiomedical Imaging 종합 리뷰연결어: extranodal extension, ene, nodal
  6. +8 more
Radiomics and Radiogenomics연결 근거: radiomics, radiomic, radiogenomic, intratumor heterogeneity, subsampling
  1. Methods방법론 색인 — 어떤 분석을 어떤 논문이 썼나연결어: radiomics, radiomic, radiogenomic
  2. ReviewBiomedical Imaging 종합 리뷰연결어: radiomics, radiomic, radiogenomic
  3. Su 2023Radiogenomic-based multiomic analysis reveals imaging intratumor heterogeneity phenotypes and t…연결어: radiomics, radiomic, radiogenomic
  4. Henry 2022Investigation of radiomics based intra-patient inter-tumor heterogeneity and the impact of tumo…연결어: radiomics, radiomic, subsampling
  5. Song 2017Imaging Phenotyping Using Radiomics to Predict Micropapillary Pattern within Lung Adenocarcinoma연결어: radiomics, radiomic
  6. +21 more
Thoracic CT / PET Imaging연결 근거: ct, pet, pulmonary, lung transplantation, churg-strauss
  1. Methods방법론 색인 — 어떤 분석을 어떤 논문이 썼나연결어: ct, pet, pulmonary
  2. ReviewClinical Pulmonology 종합 리뷰연결어: ct, pet, pulmonary
  3. DeFreitas 2021Complications of Lung Transplantation: Update on Imaging Manifestations and Management연결어: ct, pet, pulmonary
  4. Cottin 2016Respiratory manifestations of eosinophilic granulomatosis with polyangiitis (Churg-Strauss)연결어: ct, pulmonary, egpa
  5. ReviewBiomedical Imaging 종합 리뷰연결어: ct, pet, pulmonary
  6. +64 more
Shared Tags이 논문의 tags: adjuvant, durvalumab, neoadjuvant, nivolumab, pembrolizumab
  1. Liu 2024Efficacy of neoadjuvant immunochemotherapy and survival surrogate analysis of neoadjuvant treat…공유 tag: adjuvant, durvalumab, neoadjuvant, nivolumab
  2. Sorin 2024Neoadjuvant Chemoimmunotherapy for NSCLC: A Systematic Review and Meta-Analysis공유 tag: adjuvant, durvalumab, neoadjuvant, nivolumab
  3. Methods방법론 색인 — 어떤 분석을 어떤 논문이 썼나공유 tag: adjuvant, durvalumab, neoadjuvant, nivolumab
  4. Qu 2024Non-invasive prediction for pathologic complete response to neoadjuvant chemoimmunotherapy in l…공유 tag: adjuvant, neoadjuvant, nivolumab, pembrolizumab
  5. Xiong 2025Delta-radiomics features combined with haematological index predict pathological complete respo…공유 tag: adjuvant, durvalumab, neoadjuvant, nivolumab
  6. +66 more

라이브러리 148편에서 실제로 사용된 통계 방법 17가지를 정리했습니다. 각 방법은 개념·적용 상황·이 논문들에서의 사용 방식·설계 시 주의점 순으로 서술하고, 근거 논문을 표로 함께 실었습니다. 맨 아래에 연구 설계용 체크리스트가 있습니다.

생존분석

Kaplan-Meier 생존곡선 — 57편

개념 — Kaplan-Meier 방법은 시간 경과에 따른 사건 발생(예: 사망, 재발)의 누적 확률을 추정하는 비모수적 방법이다. 이 라이브러리의 논문들은 주로 Overall Survival (OS), Progression-free Survival (PFS), Recurrence-free Survival (RFS), Disease-free Survival (DFS) 및 Cancer-specific Survival (CSS)과 같은 time-to-event endpoint를 시각화하고 비교하기 위해 이를 사용한다 [1, 2, 3, 4, 6, 7, 8, 9, 11, 12].

언제 쓰나 — 연구자들은 주로 Retrospective cohort study [1, 3, 5, 6, 7, 8, 12] 또는 Prospective phase 2 trial [11]에서 특정 임상 변수(예: 폐 침윤 유무 [1], TMBRB biomarker [2], Radiomics score [3, 5, 6, 7])나 치료군(예: P-mono vs P-combo [8], SLR vs SBRT [9]) 간의 예후 차이를 평가할 때 선택했다. 특히, High-risk와 Low-risk 군으로 분류한 후 생존율의 유의미한 차이를 확인하거나, 다양한 종양 유형 및 치료 접근법(Neoadjuvant immunotherapy [10, 11])의 장기적 효과를 비교하는 맥락에서 활용되었다.

이 라이브러리에서의 사용 — 대부분의 연구는 진단 시점 [1, 4, 7, 12] 또는 수술/치료 시작 시점 [8, 9, 11]을 baseline으로 설정하여 추적 관찰 기간(follow-up period) 동안의 생존 곡선을 도출했다. 표본 규모는 단일 기관 소규모 코호트(n=42 [1], n=206 [12])부터 다기관 대규모 데이터셋(N=1,474 [6], N=40,967 [4])까지 다양하며, 일부 연구에서는 Propensity Score Matching (PSM) [8, 9]을 통해 군 간 균형을 맞춘 후 생존 분석을 수행했다. 보고 방식은 주로 Kaplan-Meier 곡선을 제시하고, Log-rank test 등을 통해 군 간 차이의 통계적 유의성을 검정하는 형태로 일관되게 나타났다 [1, 2, 3, 6, 7, 8, 9, 12].

설계 시 주의점 — 연구 설계 시 baseline 정의의 명확성이 중요하며, 특히 변이형 종양(T-SCLC)과原発성 종양(P-SCLC)처럼 시작 시점이 다른 경우(병리적 변형 시점 vs 초기 진단 시점 [12])에는 해석에 주의해야 한다. 또한, Retrospective study에서 선택 편향(selection bias)을 최소화하기 위해 PSM [8, 9]이나 다기관 검증(Multi-institutional validation [6, 7])을 병행하는 것이 권장된다. 데이터 전처리 단계에서는 Radiomics features의 경우 Standard-Scaler [7]나 Z-score normalization [3]을 적용하여 변수 간 스케일 차이를 보정해야 하며, 추적 관찰 기간(median follow-up)과 Data cut-off date를 명시적으로 보고해야 한다 [4, 8, 12].

근거 논문

  1. Pulmonary Intravascular Lymphomatosis: Clinical, CT, and PET Findings, Correlation of CT and Pathologic Results, and Survival Outcome — Radiology · 2016
  2. Predicting response to immunotherapy in advanced non-small-cell lung cancer using tumor mutational burden radiomic biomarker — Journal for Immunotherapy of Cancer · 2020
  3. Are radiomics features universally applicable to different organs? — Cancer Imaging · 2021
  4. Deep Learning to Estimate Biological Age From Chest Radiographs — JACC Cardiovascular Imaging · 2021
  5. Imaging Phenotyping Using Radiomics to Predict Micropapillary Pattern within Lung Adenocarcinoma — Journal of Thoracic Oncology · 2017
  6. Radiogenomic-based multiomic analysis reveals imaging intratumor heterogeneity phenotypes and therapeutic targets — Science Advances · 2023
  7. Imaging-Based Biomarkers Predict Programmed Death-Ligand 1 and Survival Outcomes in Advanced NSCLC Treated With Nivolumab and Pembrolizumab: A Multi-Institutional Study — JTO Clinical and Research Reports · 2023
  8. Long-term survival comparison of first-line pembrolizumab versus pembrolizumab plus chemotherapy for patients with advanced non-small cell lung cancer: A multicenter propensity-matched cohort study — Lung Cancer · 2025
  9. Sublobar resection, stereotactic body radiotherapy, and thermal ablation for early-stage non-small cell lung cancer a systematic review and meta-analysis — Lung Cancer · 2026
  10. The rapidly evolving paradigm of neoadjuvant immunotherapy across cancer types — Nature Cancer · 2025
  11. Perioperative Pembrolizumab for Locally Advanced Thymic Epithelial Tumors: A Single-Arm, Phase 2 Trial — Journal of Thoracic Oncology · 2025
  12. Clinical outcomes and neuroendocrine features of transformed versus primary small-cell lung cancer — Lung Cancer · 2025
Cox proportional hazards model — 64편

개념 — 이 라이브러리의 논문들은 Cox proportional hazards model을 사용하여 시간-사건(time-to-event) 데이터에서 특정 변수(예: radiomics score, 치료군, 병리학적 특징)가 생존 결과(Overall Survival, Progression-free Survival, Recurrence-free Survival)에 미치는 영향을 추정하고, Hazard ratio (HR)를 통해 위험도의 크기와 방향성을 정량화한다 [1, 3, 4, 5, 7, 8, 9, 10, 11]. 이 모델은 사건 발생까지의 시간과 함께 covariates의 효과를 동시에 고려하여, 다른 변수들을 통제했을 때 각 요인의 독립적인 예후 예측 능력을 평가하는 데 활용된다 [9, 11].

언제 쓰나 — 연구자들은 주로 retrospective cohort study 또는 observational cohort study 설계에서, 진단 시점부터 사망 또는 재발까지의 추적 관찰 기간(follow-up period)이 존재할 때 이 모델을 선택한다 [1, 2, 4, 7, 8, 9, 10, 11]. 특히 advanced NSCLC 환자에서의 Immunotherapy 반응 예측 [3, 9, 10], Radiomics 기반의 영상 표현형(Imaging Phenotyping)이 예후와 연관되는지 확인하는 연구 [4, 6, 7, 8], 그리고 서로 다른 치료 전략(P-mono vs P-combo) 간의 장기 생존 비교 [11]에서 핵심 분석 도구로 사용되었다. 또한, 다기관 데이터나 외부 검증 cohort(external validation cohort)를 포함하여 모델의 일반화 가능성을 평가할 때 적용된다 [5, 7, 8, 10].

이 라이브러리에서의 사용 — 대부분의 연구는 Primary endpoint으로 Overall Survival (OS), Progression-free Survival (PFS), 또는 Recurrence-free Survival (RFS)을 설정하고, Kaplan-Meier 곡선과 함께 Cox regression을 수행하여 Hazard ratio (HR)와 95% Confidence Interval (CI)를 보고한다 [1, 3, 7, 8, 9, 10, 11]. 분석 과정에서는 단변량 생존 분석(Univariate survival analysis)을 통해 유의한 변수를 선별한 후, 다변량 분석(Multivariate analysis)으로 독립적인 예측 인자를 도출하는 패턴이 보인다 [9, 11]. 일부 연구는 Propensity Score Matching (PSM)을 통해 치료군 간 기저 특성의 불균형을 보정한 후 Cox model을 적용하여 confounding을 통제했다 [11]. 또한, Radiomics features나 Deep learning biomarker(TMBRB, CXR-Age)를 연속형 변수 또는 이진 분류(high-risk/low-risk)로 변환하여 모델에 포함시키는 방식이 두드러진다 [3, 4, 5, 7, 8, 10].

설계 시 주의점 — Cox model의 전제 조건인 비례 위험 가정(proportional hazards assumption)을 검증해야 하며, 특히 Radiomics나 Deep learning 기반 biomarker를 사용할 때는 feature extraction의 재현성(Segmentation 신뢰도)과 데이터 전처리(Standard-Scaler, z-score normalization)가 결과에 미치는 영향을 고려해야 한다 [4, 6, 10]. Retrospective study에서는 temporal confounding이나 operator bias와 같은 선택 편향을 최소화하기 위해 Propensity Score Matching (PSM)이나 엄격한 inclusion/exclusion criteria를 적용하는 것이 중요하다 [2, 11]. 또한, 다변량 분석에 포함할 covariates(age, sex, ECOG-PS 등)는 사전에 정의(prespecified)하거나 임상적 근거를 바탕으로 선정해야 하며, 최종 모델의 성능은 AUC, sensitivity, specificity 또는 Decision curve analysis와 함께 보고하여 임상적 유용성을 입증해야 한다 [3, 6, 8, 10].

근거 논문

  1. Pulmonary Intravascular Lymphomatosis: Clinical, CT, and PET Findings, Correlation of CT and Pathologic Results, and Survival Outcome — Radiology · 2016
  2. Combined fluoroscopy-and CT-guided transthoracic needle biopsy using a C-arm cone-beam CT system: comparison with fluoroscopy-guided biopsy — Korean Journal of Radiology · 2011
  3. Predicting response to immunotherapy in advanced non-small-cell lung cancer using tumor mutational burden radiomic biomarker — Journal for Immunotherapy of Cancer · 2020
  4. Are radiomics features universally applicable to different organs? — Cancer Imaging · 2021
  5. Deep Learning to Estimate Biological Age From Chest Radiographs — JACC Cardiovascular Imaging · 2021
  6. Imaging Phenotyping Using Radiomics to Predict Micropapillary Pattern within Lung Adenocarcinoma — Journal of Thoracic Oncology · 2017
  7. Radiogenomic-based multiomic analysis reveals imaging intratumor heterogeneity phenotypes and therapeutic targets — Science Advances · 2023
  8. Development of an Immune-Pathology Informed Radiomics Model for Non-Small Cell Lung Cancer — Scientific Reports · 2018
  9. Spatial interplay patterns of cancer nuclei and tumor-infiltrating lymphocytes (TILs) predict clinical benefit for immune checkpoint inhibitors — Science Advances · 2022
  10. Imaging-Based Biomarkers Predict Programmed Death-Ligand 1 and Survival Outcomes in Advanced NSCLC Treated With Nivolumab and Pembrolizumab: A Multi-Institutional Study — JTO Clinical and Research Reports · 2023
  11. Long-term survival comparison of first-line pembrolizumab versus pembrolizumab plus chemotherapy for patients with advanced non-small cell lung cancer: A multicenter propensity-matched cohort study — Lung Cancer · 2025
  12. The rapidly evolving paradigm of neoadjuvant immunotherapy across cancer types — Nature Cancer · 2025
Log-rank test — 39편

개념 — 이 방법은 Kaplan-Meier 추정법을 통해 생성된 생존 곡선(Survival curves) 간의 통계적 차이를 검정하는 비모수적 방법이다. 주로 Overall Survival (OS), Progression-Free Survival (PFS), Recurrence-free survival (RFS), Event-free survival (EFS)과 같은 시간-사건(Time-to-event) 데이터에서 두 개 이상의 군(Group) 간 생존율의 유의한 차이를 평가한다 [1, 3, 7, 8, 9, 10, 11, 12].

언제 쓰나 — 이 라이브러리의 논문들은 주로 Retrospective cohort study [1, 3, 5, 6, 7, 9, 10] 또는 Randomized controlled trial (RCT) [11, 12]에서 특정 임상적 특징(예: 폐 침윤 유무 [1], TMBRB 고/저위험군 [2], Radiomics score 기반 위험군 [3, 6]), 치료 방식(예: P-mono vs P-combo [8], Neoadjuvant therapy 조합 [11, 12]), 또는 생물학적 표지자 발현 상태(예: PD-L1 발현 [7, 10])에 따라 분류된 군들 간의 예후 차이를 비교할 때 선택했다. 특히 Radiomics 특징을 기반으로 High-risk와 Low-risk로 이분화하거나 다분화한 후 생존 분석을 수행하는 경우가 빈번하다 [3, 6].

이 라이브러리에서의 사용 — 대부분의 연구에서 Primary endpoint으로 OS, PFS, RFS, EFS를 설정하고, 이를 기준으로 Kaplan-Meier curves를 작성한 후 Log-rank test로 p-value를 산출하여 보고했다 [1, 7, 8, 9, 10, 11, 12]. 표본 규모는 단일 기관 연구에서 수십 명(예: n=42 [1], n=206 [9])부터 다기관 RCT에서 수백 명(예: n=358 [12])까지 다양하며, 추적 관찰 기간(Median follow-up)은 11.9개월 [4]에서 75개월 [10]까지 광범위하게 분포한다. 분석 과정에서는 종종 Cox proportional hazards model과 병행하여 Hazard ratio (HR)와 신뢰구간을 함께 보고하는 경향이 있다 [11, 12].

설계 시 주의점 — 군 간 비교 시 기저 특성(Baseline characteristics)의 불균형을 보정하기 위해 Propensity score matching (PSM) [8]이나 다변량 분석(Multivariate analysis) [10, 11]을 병행해야 하며, 특히 Retrospective study에서는 선택 편향(Selection bias)에 주의해야 한다. 또한, T-SCLC와 P-SCLC 비교처럼 각 군의 Baseline 정의 시점(진단 시점 vs 변형 시점)이 다를 경우 [9], 추적 관찰 기간(Data cut-off date)과 중위 추적 시간(Median follow-up time)을 명확히 명시하여 해석의 오류를 방지해야 한다. Radiomics 기반 연구에서는 Feature extraction 전의 데이터 전처리(Normalization 등) [3, 7]와 모델 검증(External validation) [6, 7] 과정이 생존 분석 결과의 신뢰성에 직접적인 영향을 미치므로 방법론적 투명성이 필수적이다.

근거 논문

  1. Pulmonary Intravascular Lymphomatosis: Clinical, CT, and PET Findings, Correlation of CT and Pathologic Results, and Survival Outcome — Radiology · 2016
  2. Predicting response to immunotherapy in advanced non-small-cell lung cancer using tumor mutational burden radiomic biomarker — Journal for Immunotherapy of Cancer · 2020
  3. Are radiomics features universally applicable to different organs? — Cancer Imaging · 2021
  4. Deep Learning to Estimate Biological Age From Chest Radiographs — JACC Cardiovascular Imaging · 2021
  5. Imaging Phenotyping Using Radiomics to Predict Micropapillary Pattern within Lung Adenocarcinoma — Journal of Thoracic Oncology · 2017
  6. Radiogenomic-based multiomic analysis reveals imaging intratumor heterogeneity phenotypes and therapeutic targets — Science Advances · 2023
  7. Imaging-Based Biomarkers Predict Programmed Death-Ligand 1 and Survival Outcomes in Advanced NSCLC Treated With Nivolumab and Pembrolizumab: A Multi-Institutional Study — JTO Clinical and Research Reports · 2023
  8. Long-term survival comparison of first-line pembrolizumab versus pembrolizumab plus chemotherapy for patients with advanced non-small cell lung cancer: A multicenter propensity-matched cohort study — Lung Cancer · 2025
  9. Clinical outcomes and neuroendocrine features of transformed versus primary small-cell lung cancer — Lung Cancer · 2025
  10. PD-L1 Expression is Highly Associated with Tumor-Associated Macrophage Infiltration in Nasopharyngeal Carcinoma — Cancer Management and Research · 2020
  11. Association between pathologic response and survival after neoadjuvant therapy in lung cancer — Nature Medicine · 2024
  12. Neoadjuvant Nivolumab plus Chemotherapy in Resectable Lung Cancer — New England Journal of Medicine · 2022
Competing risk analysis — 7편

개념 — 이 방법은 특정 사건(event)이 발생할 확률을 추정할 때, 다른 경쟁 사건(competing risk, 예: 비관심사망)으로 인해 관찰이 중단되는 경우를 고려하여 실제 발생 위험을 더 정확하게 반영하는 분석 기법이다. 특히 Kaplan-Meier method는 경쟁 사건의 영향을 과소평가하거나 왜곡할 수 있으므로, Cumulative Incidence Function (CIF)을 사용하여 각 원인에 따른 누적 발생률을 추정한다 [7]. 또한 Cause-specific hazard model과 Subdistribution hazard model(Fine-Gray model) 등 회귀 분석 접근법을 통해 경쟁 사건의 존재 하에서 위험비(hazard ratio)를 해석하는 데 활용된다 [7].

언제 쓰나 — 이 라이브러리의 논문들은 주로 장기 추적 관찰(long-term follow-up)이 이루어진 cohort study에서 사망(mortality)이나 재발(recurrence)과 같은 주요 endpoint를 분석할 때 경쟁 위험을 고려했다 [1, 3, 4, 5, 6]. 특히 NSCLC(비소세포폐암) 환자의 재발률(cumulative incidence of recurrence) 추정 [3]이나, 방사선 폐렴(radiation pneumonitis)과 같은 부작용의 발생률을 분석할 때 [5, 6], 다른 원인 사망이 경쟁 사건으로 작용하는 상황을 다루었다. 또한 meta-analysis를 수행할 때, 개별 연구가 competing risk를 적절히 처리했는지 확인하여 포함 기준으로 삼는 경우도 있었다 [4].

이 라이브러리에서의 사용 — 실제 적용에서는 Kaplan-Meier method로 overall survival(OS)을 추정하는 동시에 [5, 6], 재발이나 특정 부작용의 누적 발생률(cumulative incidence)은 CIF를 사용하여 보고했다 [3, 7]. 표본 규모는 수백 명에서 수만 명(N=1439 [5, 6], N=40,967 [1])에 이르기까지 다양하며, retrospective multicenter cohort study나 nationwide claims database를 기반으로 한 연구가 많았다 [2, 5, 6]. 보고 방식에서는 단순히 생존 곡선을 제시하는 것을 넘어, competing risk의 영향을 명시적으로 언급하거나, meta-analysis 시 해당 분석 방법을 사용한 연구만 선별하여 포함하는 등 방법론적 엄격성을 유지했다 [4, 7].

설계 시 주의점 — Kaplan-Meier method를 사용하여 재발률이나 특정 사건 발생률을 추정할 경우, 경쟁 사건(예: 다른 원인 사망)이 있는 환자들을 censored로 처리하면 발생률이 과대평가될 수 있으므로 CIF 사용이 필수적이다 [7]. 또한 meta-analysis를 설계할 때는 원문 연구가 competing risk를 어떻게 처리했는지(methodology transparency)를 포함/배제 기준으로 삼아야 신뢰성 있는 요약 추정치를 얻을 수 있다 [4]. 회귀 분석 시에는 Cause-specific hazard와 Subdistribution hazard 모델 중 연구 질문에 적합한 모델을 선택하고, 그 해석의 차이를 명확히 보고해야 한다 [7].

근거 논문

  1. Deep Learning to Estimate Biological Age From Chest Radiographs — JACC Cardiovascular Imaging · 2021
  2. Long-term survival comparison of first-line pembrolizumab versus pembrolizumab plus chemotherapy for patients with advanced non-small cell lung cancer: A multicenter propensity-matched cohort study — Lung Cancer · 2025
  3. Recurrence After Complete Resection for Non-Small Cell Lung Cancer in the National Lung Screening Trial — The Annals of Thoracic Surgery · 2023
  4. Recurrence-Free Survival in Patients With Surgically Resected Non-Small Cell Lung Cancer: A Systematic Literature Review and Meta-Analysis — Chest · 2024
  5. Durvalumab After Chemoradiotherapy for Locally Advanced NSCLC A Real-World Analysis Using a Nationwide Claims Database in Japan — 2026
  6. Durvalumab After Chemoradiotherapy for Locally Advanced NSCLC A Real-World Analysis Using a Nationwide Claims Database in Japan — 2026
  7. Introduction to the Analysis of Survival Data in the Presence of Competing Risks — 2016
Landmark analysis — 8편

개념 — Landmark analysis는 치료 시작 후 특정 고정된 시점(landmark time)을 설정하고, 해당 시점까지 생존한 환자군만을 대상으로 분석함으로써 guarantee-time bias(GTB)를 제거하는 방법이다 [8]. 이 방법은 follow-up 중 발생하는 사건(classifying event)이 outcome에 미치는 영향을 평가할 때, 사건의 발생 시점이 outcome 측정 시점보다 앞서는 구조적 편향을 보정한다 [8].

언제 쓰나 — Retrospective observational study 또는 multicenter cohort study에서 treatment-related adverse events(irAEs), early discontinuation, 또는 time-dependent covariates의 영향을 평가할 때 선택된다 [3][6][7]. 특히 durvalumab consolidation therapy 후 irAE 발생과 efficacy(PFS/OS) 간의 연관성을 분석하거나, CRT 후 durvalumab 투여 여부가 생존에 미치는 영향을 비교하는 상황에서 GTB를 통제하기 위해 적용되었다 [3][7].

이 라이브러리에서의 사용 — 이 논문들에서는 주로 landmark time을 설정하여 해당 시점까지 생존한 환자(subset)만을 추출한 후, classifying event(예: irAE 발생 여부, durvalumab 투여 여부)에 따라 군을 분류하고 이후의 outcome(OS, PFS)을 비교했다 [3][7][8]. 표본 규모는 48명에서 194명까지 다양하며, propensity score matching(PSM)과 결합하여 기저 특성의 불균형을 추가로 보정하는 방식도 사용되었다 [1][7]. 보고 방식은 landmark time 이후의 Kaplan-Meier 생존 곡선 및 hazard ratio(HR)를 제시하는 것이 일반적이었다 [3][7].

설계 시 주의점 — Landmark time을 설정할 때는 해당 시점까지 충분한 수의 환자가 생존해야 하며, 너무 늦은 시점을 선택하면 표본 크기가 급격히 감소하여 통계적 검정력이 떨어질 수 있다 [8]. 또한, landmark time 이전에 발생한 사건(예: early discontinuation)은 분석에서 제외되므로, 이로 인한 selection bias 가능성을 명시적으로 논의해야 한다 [6][8]. Classifying event의 정의와 landmark time의 선택 근거를 명확히 보고하며, time-varying covariates 모델과의 결과 비교를 통해 robustness를 확인하는 것이 권장된다 [8].

근거 논문

  1. Long-term survival comparison of first-line pembrolizumab versus pembrolizumab plus chemotherapy for patients with advanced non-small cell lung cancer: A multicenter propensity-matched cohort study — Lung Cancer · 2025
  2. Clinical outcomes and neuroendocrine features of transformed versus primary small-cell lung cancer — Lung Cancer · 2025
  3. Association of immune-related adverse events with durvalumab efficacy after chemoradiotherapy in patients with unresectable Stage III non-small cell lung cancer — 2024
  4. 2405-6308/© 2022 The Authors. Published by Elsevier B.V. on behalf of European Society for Radiotherapy and Oncology. This is an open access article under the — 2022
  5. Efficacy of neoadjuvant immunochemotherapy and survival surrogate analysis of neoadjuvant treatment in IB–IIIB lung squamous cell carcinoma — Scientific Reports · 2024
  6. Exploring causes and consequences of early discontinuation of durvalumab after chemoradiotherapy for non-small cell lung cancer — 2023
  7. Derived neutrophil-to-lymphocyte ratio has the potential to predict safety and outcomes of durvalumab after chemoradiation in non-small cell lung cancer — 2024
  8. Challenges of Guarantee-Time Bias — 2013

진단·예측 성능

ROC curve / AUC — 37편

개념 — ROC curve와 AUC는 이진 분류 문제에서 모델의 diagnostic accuracy나 predictive performance를 종합적으로 평가하는 지표로, sensitivity와 specificity의 trade-off 관계를 시각화하고 그 아래 면적을 수치화한다. 본 라이브러리의 논문들은 이를 통해 radiomics, deep learning, 또는 임상 변수 기반 모델이 특정 endpoint(예: malignancy, pCR, immune cell infiltration)를 얼마나 정확하게 구분하는지를 정량적으로 검증한다.

언제 쓰나 — 이 방법론은 주로 medical imaging (CT, chest radiographs, PET/CT) 기반의 biomarker 개발 및 검증 연구에서 선택되었다 [1, 2, 4, 5, 7, 9, 10, 12]. 특히 high-dimensional data(특성 수 > 표본 수)를 다루는 radiomics 파이프라인 [1]이나, deep learning 모델의 generalization 성능을 평가할 때 [3, 4, 6] 필수적으로 사용된다. 또한, neoadjuvant therapy 후 pathological response (pCR, MPR) 예측 [6, 9, 12]이나, tumor microenvironment 내 immune cell signature 분류 [8, 11]와 같은 복잡한 생물학적 현상을 이진화하여 평가하는 경우에도 적용되었다.

이 라이브러리에서의 사용 — 대부분의 연구는 데이터를 training, validation, test 세트로 무작위 분할(random split)하거나 외부 cohort를 통해 external validation을 수행하며 AUC를 산출했다 [3, 4, 5, 6, 9, 10]. 평가 방식은 단순한 point estimate뿐만 아니라, fivefold stratified cross-validation with repeats [1], bootstrap resampling (1000회) [4], 또는 DeLong’s test를 통한 모델 간 AUC 차이의 통계적 유의성 검정 [6]을 포함한다. Primary endpoint는 주로 binary classification (예: High-TMB vs Low-TMB [3], pCR yes/no [6, 9])이었으며, 일부 연구에서는 multiclass 문제(예: MP 비율 세분화 [7])나 survival outcome (OS, PFS)과의 연관성 평가에도 활용되었다.

설계 시 주의점 — 모델의 overfitting을 방지하기 위해 적절한 데이터 분할 전략과 cross-validation 또는 bootstrap 방법이 필수적이다 [1, 4]. 또한, radiomics feature extraction 전후의 pre-processing (예: feature normalization methods [1], Standard-Scaler [10])가 AUC 결과에 미치는 영향을 반드시 고려해야 한다. 모델 간 성능 비교 시에는 DeLong’s test와 같은 통계적 검정을 수행하여 유의성을 입증해야 하며 [6], segmentation의 재현성(ICC ≥ 0.8)을 확인하는 등 전처리 단계의 신뢰도 검증도 중요하다 [7, 9]. 마지막으로, AUC 외에도 sensitivity, specificity, calibration (Brier score) 등의 secondary endpoints를 함께 보고하여 모델의 임상적 유용성을 다각도로 평가해야 한다 [1, 3, 4].

근거 논문

  1. The effect of feature normalization methods in radiomics — Insights Imaging · 2024
  2. Multicentre performance and consistency of two deep learning models for malignancy probability estimation of incidental pulmonary nodules — 2026
  3. Predicting response to immunotherapy in advanced non-small-cell lung cancer using tumor mutational burden radiomic biomarker — Journal for Immunotherapy of Cancer · 2020
  4. Multimodal deep learning for integrating chest radiographs and clinical parameters: a case for transformers — Radiology · 2023
  5. Development and validation of a deep learning algorithm detecting 10 common abnormalities on chest radiographs — European Respiratory Journal · 2021
  6. Non-invasive prediction for pathologic complete response to neoadjuvant chemoimmunotherapy in lung cancer using CT-based deep learning: a multicenter study — Frontiers in Immunology · 2024
  7. Imaging Phenotyping Using Radiomics to Predict Micropapillary Pattern within Lung Adenocarcinoma — Journal of Thoracic Oncology · 2017
  8. Spatial interplay patterns of cancer nuclei and tumor-infiltrating lymphocytes (TILs) predict clinical benefit for immune checkpoint inhibitors — Science Advances · 2022
  9. Delta-radiomics features combined with haematological index predict pathological complete response after neoadjuvant immunochemotherapy in resectable non-small cell lung cancer — Clinical Radiology · 2025
  10. Imaging-Based Biomarkers Predict Programmed Death-Ligand 1 and Survival Outcomes in Advanced NSCLC Treated With Nivolumab and Pembrolizumab: A Multi-Institutional Study — JTO Clinical and Research Reports · 2023
  11. Deciphering the tumor microenvironment through radiomics in non-small cell lung cancer: Correlation with immune profiles — PloS One · 2020
  12. Dynamic (18) F-FDG PET/CT can predict the major pathological response to neoadjuvant immunotherapy in non-small cell lung cancer — Thorac Cancer · 2022
Sensitivity / Specificity — 55편

개념 — 이 라이브러리의 논문들은 주로 영상 기반 진단 도구(CT, X-ray, PET) 또는 AI 모델이 특정 병리 상태(malignancy, pCR, abnormalities)를 식별하거나 예후(OS, PFS)를 예측하는 성능을 정량화하기 위해 sensitivity와 specificity를 사용한다. 이는 binary classification 문제에서 참 양성(True Positive)과 참 음성(True Negative)의 비율을 통해 진단 정확도(diagnostic accuracy)를 평가하는 핵심 지표로 기능한다. 또한 일부 연구에서는 cutoff value 결정[11]이나 치료 반응 기준 정의[12]와 같은 임계값 설정 과정에서도 이 개념이 적용된다.

언제 쓰나 — 연구자들은 새로운 영상 기법(FC-TNB vs F-TNB)[1]이나 AI 알고리즘(DLAD-10, TMBRB)[4, 5, 8]의 임상적 유효성을 입증할 때 주로 이 지표를 선택한다. 특히 ground truth가 명확한 경우(수술 병리 결과[9], 장기 follow-up 데이터[10])나 기존 표준 진단법과의 비교 평가 시 필수적으로 활용된다. 또한 radiomics 파이프라인의 전처리 단계(normalization methods)[3]가 최종 분류 성능에 미치는 영향을 검증할 때도 sensitivity와 specificity를 secondary endpoint로 포함하여 모델의 robustness를 확인한다.

이 라이브러리에서의 사용 — 대부분의 연구는 ROC curve 분석을 통해 AUC, sensitivity, specificity를 함께 보고하며[5, 6, 9], 일부는 bootstrap resampling[6]이나 cross-validation[3]을 통해 추정치의 신뢰구간을 제공한다. 표본 규모는 소규모 코호트(n=74~248)[1, 9]부터 대규모 공개 데이터셋(N>30,000)[6, 8]까지 다양하지만, 공통적으로 training/validation/test set으로 데이터를 분할하여 overfitting을 방지한다. 보고 방식에서는 단일 cutoff point에서의 성능뿐만 아니라, 다양한 threshold에 따른 trade-off를 시각화하거나[5], 특정 임상 시나리오(simulated reading test)[8]에서의 실제 적용 가능성을 함께 제시하는 경향이 있다.

설계 시 주의점 — retrospective study 설계에서 temporal confounding[1]이나 operator bias[1]를 통제하지 않으면 sensitivity/specificity 추정치가 왜곡될 수 있으므로, inclusion/exclusion criteria와 데이터 수집 기간을 명확히 정의해야 한다. 또한 class imbalance가 심한 경우(예: 희귀 질환 또는 pCR 달성 환자)[9], 단순 accuracy보다 sensitivity와 specificity의 균형 또는 AUC를 주요 지표로 삼아야 하며, cutoff value는 clinical utility에 기반하여 사전에 설정하거나[11] 최적화 과정을 투명하게 보고해야 한다. 마지막으로, external validation cohort[4, 8, 9]에서의 성능 저하 가능성을 고려하여, 모델의 generalizability를 평가할 때 sensitivity/specificity의 변화폭을 반드시 비교 분석해야 한다.

근거 논문

  1. Combined fluoroscopy-and CT-guided transthoracic needle biopsy using a C-arm cone-beam CT system: comparison with fluoroscopy-guided biopsy — Korean Journal of Radiology · 2011
  2. Complications of Lung Transplantation: Update on Imaging Manifestations and Management — Radiology: Cardiothoracic Imaging · 2021
  3. The effect of feature normalization methods in radiomics — Insights Imaging · 2024
  4. Multicentre performance and consistency of two deep learning models for malignancy probability estimation of incidental pulmonary nodules — 2026
  5. Predicting response to immunotherapy in advanced non-small-cell lung cancer using tumor mutational burden radiomic biomarker — Journal for Immunotherapy of Cancer · 2020
  6. Multimodal deep learning for integrating chest radiographs and clinical parameters: a case for transformers — Radiology · 2023
  7. Radiomics Quality Score 2.0: towards radiomics readiness levels and clinical translation for personalized medicine — Nature Reviews Clinical Oncology · 2025
  8. Development and validation of a deep learning algorithm detecting 10 common abnormalities on chest radiographs — European Respiratory Journal · 2021
  9. Non-invasive prediction for pathologic complete response to neoadjuvant chemoimmunotherapy in lung cancer using CT-based deep learning: a multicenter study — Frontiers in Immunology · 2024
  10. Deep Learning to Estimate Biological Age From Chest Radiographs — JACC Cardiovascular Imaging · 2021
  11. Development of an Immune-Pathology Informed Radiomics Model for Non-Small Cell Lung Cancer — Scientific Reports · 2018
  12. From RECIST to PERCIST: Evolving Considerations for PET response criteria in solid tumors — Journal of Nuclear Medicine · 2009
C-index / Concordance — 13편

개념 — C-index (Concordance index)는 생존 분석 모델이 두 임의의 환자 간 위험 순서를 얼마나 정확하게 예측하는지를 평가하는 지표로, 0.5(무작위 추측)에서 1.0(완벽한 예측) 사이의 값을 가진다 [2, 8, 11]. 이 지수는 모델이 더 짧은 생존 시간을 가진 환자를 더 높은 위험도로 분류할 확률을 추정하며, 특히 time-to-event 데이터의 예측 성능을 요약하는 데 핵심적으로 사용된다 [8, 11].

언제 쓰나 — 본 라이브러리의 논문들은 주로 Overall Survival (OS), Progression-Free Survival (PFS), Disease-Free Survival (DFS) 또는 Transplant-Free Survival (TFS)과 같은 시간 의존성 endpoint를 예측하는 prognostic model의 성능을 검증할 때 C-index를 선택했다 [2, 3, 4, 8, 9, 10, 11, 12]. 특히 Deep Learning 기반 모델 [2, 8, 10]이나 Radiomics 특징을 결합한 다변량 모델 [3, 4, 11]에서 기존 임상 변수 대비 추가적인 예측 가치를 입증하기 위해 활용되었다. 또한, 단일 시점의 영상 biomarker가 장기적인 사망 위험(all-cause mortality)을 얼마나 잘 반영하는지 평가하는 맥락에서도 빈번하게 등장한다 [2, 8, 11].

이 라이브러리에서의 사용 — 연구들은 대부분 Training cohort와 Validation cohort(또는 External validation cohort)로 데이터를 분할하여 모델의 일반화 성능을 확인했으며, 각 단계에서 C-index를 보고했다 [2, 3, 4, 8, 10]. 예를 들어, HCC 절제술 후 생존 예측 연구에서는 Derivation cohort와 External validation cohort 모두에서 C-index를 계산하여 모델의 robustness를 입증했다 [8]. NSCLC 관련 연구들에서는 Radiomics signature나 circulating tumor antigen (ctA) panel을 Cox regression 또는 Machine Learning 알고리즘(XGBoost, Elastic Net)에 입력한 후, 결과 모델의 discrimination ability를 C-index로 정량화했다 [3, 4, 10]. 일부 연구는 단순히 최종 모델의 성능뿐만 아니라, 개별 biomarker(예: CXR-Age)와 전통적 변수(chronological age) 간의 예측력 차이를 C-index 비교를 통해 시각화하기도 했다 [2].

설계 시 주의점 — C-index 보고 시에는 반드시 Internal validation (예: cross-validation, bootstrap)과 External validation cohort에서의 성능을 구분하여 명시해야 한다 [2, 3, 4, 8]. 또한, 모델 개발 과정에서 feature selection (예: LASSO, hierarchical clustering)이나 hyperparameter tuning이 수행되었다면, 이를 통해 과적합(overfitting)이 발생하지 않았음을 검증하는 과정과 함께 C-index가 보고되어야 신뢰성이 확보된다 [3, 10]. 특히 Deep Learning 기반 연구에서는 Ablation study를 통해 각 모달리티의 기여도를 분석함과 동시에, 최종 통합 모델의 C-index가 단일 모달리티 모델보다 유의미하게 향상되었는지 비교해야 한다 [8]. 마지막으로, C-index는 discrimination만 평가하므로, calibration (예: calibration plot)과 함께 보고하여 모델의 예측 정확도와 일관성을 모두 입증하는 것이 본 라이브러리 연구들의 공통된 경향이다 [2, 8, 10].

근거 논문

  1. Investigation of radiomics based intra-patient inter-tumor heterogeneity and the impact of tumor subsampling strategies — Scientific Reports · 2022
  2. Deep Learning to Estimate Biological Age From Chest Radiographs — JACC Cardiovascular Imaging · 2021
  3. Development of an Immune-Pathology Informed Radiomics Model for Non-Small Cell Lung Cancer — Scientific Reports · 2018
  4. Imaging-Based Biomarkers Predict Programmed Death-Ligand 1 and Survival Outcomes in Advanced NSCLC Treated With Nivolumab and Pembrolizumab: A Multi-Institutional Study — JTO Clinical and Research Reports · 2023
  5. Hyperprogressive disease in non-small cell lung cancer treated with immune checkpoint inhibitor therapy, fact or myth? — Frontiers in Oncology · 2022
  6. Comparison of the morphologic criteria (RECIST) and metabolic criteria (EORTC and PERCIST) in tumor response assessments: a pooled analysis — The Korean Journal of Internal Medicine · 2019
  7. Comparability of PD‐L1 immunohistochemistry assays for non‐small‐cell lung cancer a systematic review — 2020
  8. Interpretable multimodal deep learning for time-resolved survival prediction after hepatocellular carcinoma resection — npj Digital Medicine · 2026
  9. Efficacy of neoadjuvant immunochemotherapy and survival surrogate analysis of neoadjuvant treatment in IB–IIIB lung squamous cell carcinoma — Scientific Reports · 2024
  10. Combined circulating tumor antigen model demonstrates additional prognostic value in first-line treatment of non-small cell lung cancer — 2026
  11. Risk factors and prognostic indicators for progressive fibrosing interstitial lung disease: a deep learning-based CT quantification approach — European Radiology · 2025
  12. park-2026-quantitative-ct-imaging-in-progressive-pulmonary-fibrosis
Calibration — 11편

개념 — Calibration은 예측 모델이 산출한 확률값(predicted probability)과 실제 관측된 사건 발생률(observed event rate) 간의 일치도를 평가하는 지표이다 [9]. 이는 단순히 분류 정확도(AUC)뿐만 아니라, 예측된 위험도가 임상적으로 얼마나 신뢰할 수 있는지를 검증하는 핵심 개념으로, 과대추정(overestimation) 또는 과소추정(underestimation)을 식별한다 [9].

언제 쓰나 — 이 라이브러리의 논문들은 주로 high-dimensional data(특성 수 > 표본 수)를 다루는 radiomics 분석 [1], deep learning 기반의 의료 영상 예측 모델 [3, 7, 8], 그리고 다중 모달리티(multimodal) 데이터를 융합하는 임상 예측 시스템 [10, 11]에서 calibration을 적용했다. 특히, neoadjuvant chemoimmunotherapy 후 pCR(pCR) 예측 [3, 5], HCC 절제술 후 Overall Survival(OS) 예측 [7], 그리고 ILD 환자의 PF-ILD 발생 및 사망률 예측 [8]과 같이 임상적 의사결정에 직접적인 영향을 미치는 위험도 추정 시 필수적으로 수행되었다.

이 라이브러리에서의 사용 — 실제 적용 방식은 주로 Brier score, expected calibration error(ECE), 그리고 calibration curve를 시각화하는 것이다 [1, 9]. 일부 연구에서는 calibration-in-the-large(평균 예측값 vs 전체 사건률), weak calibration(intercept와 slope 분석)을 통해 모델의 체계적 편향을 정량적으로 평가했다 [9]. 또한, radiomics 파이프라인에서 feature normalization methods(z-Score, Min-Max 등)가 model calibration에 미치는 영향을 fivefold stratified cross-validation with 100 repeats를 통해 검증하기도 했다 [1].

설계 시 주의점 — 모델 개발 시에는 overfitting 방지를 위해 LASSO regression [5] 또는 적절한 feature selection을 수행해야 하며, segmentation의 재현성(ICC ≥ 0.8)을 먼저 확인하는 것이 중요하다 [4, 5]. 외부 검증(external validation) cohort를 반드시 포함하여 모델의 일반화 가능성을 입증해야 하며 [3, 7], 결측치가 존재하는 sparse data 환경에서는 gating mechanism이나 robust한 프레임워크(MoE-Health 등)를 고려해야 한다 [10]. 마지막으로, calibration slope가 1에서 벗어나거나 intercept가 0이 아닌 경우, 이는 예측값이 지나치게 극단적이거나 평범함을 의미하므로 모델 보정(calibration adjustment)이 필요할 수 있음을 인지해야 한다 [9].

근거 논문

  1. The effect of feature normalization methods in radiomics — Insights Imaging · 2024
  2. Mass estimates by computed tomography: physical density from CT numbers — American Journal of Roentgenology · 1984
  3. Non-invasive prediction for pathologic complete response to neoadjuvant chemoimmunotherapy in lung cancer using CT-based deep learning: a multicenter study — Frontiers in Immunology · 2024
  4. Imaging Phenotyping Using Radiomics to Predict Micropapillary Pattern within Lung Adenocarcinoma — Journal of Thoracic Oncology · 2017
  5. Delta-radiomics features combined with haematological index predict pathological complete response after neoadjuvant immunochemotherapy in resectable non-small cell lung cancer — Clinical Radiology · 2025
  6. Association between diabetes mellitus and reduced efficacy of pembrolizumab in non-small cell lung cancer — Cancer · 2023
  7. Interpretable multimodal deep learning for time-resolved survival prediction after hepatocellular carcinoma resection — npj Digital Medicine · 2026
  8. Risk factors and prognostic indicators for progressive fibrosing interstitial lung disease: a deep learning-based CT quantification approach — European Radiology · 2025
  9. Calibration: the Achilles heel of predictive analytics — BMC Medicine · 2019
  10. MoE-Health: A Mixture of Experts Framework for Robust Multimodal Healthcare Prediction — arXiv preprint arXiv:2508.21793 · 2025
  11. Making the Most Out of the Limited Context Length: Predictive Power Varies with Clinical Note Type and Note Section — arXiv preprint arXiv:2307.07051 · 2023

회귀·교란 보정

Multivariable regression — 62편

개념 — 이 라이브러리의 논문들은 Multivariable regression을 사용하여 여러 독립 변수(임상 특성, 영상 특징, 유전자 발현 등)가 종양학 예후(Overall Survival, Progression-free Survival, Pathologic Complete Response)에 미치는 영향을 동시에 추정하거나 교란 요인을 통제하기 위해 사용한다. 특히 고차원 데이터(Radiomics features, Gene expression data)에서 예측력 있는 변수를 선별(LASSO regression 등)한 후, 최종 모델의 성능을 검증하거나 치료군 간 비교 시 confounding을 보정하는 데 핵심적으로 활용되었다.

언제 쓰나 — 연구자들은 주로 Retrospective cohort study 또는 Observational cohort study 설계에서 선택적 편향(selection bias)이나 교란 요인(confounding)을 통제해야 할 때 이 방법을 적용했다 [1, 2, 4, 7, 8, 10, 12]. 또한, 다수의 예측 인자(Radiomic features, Clinical variables)가 존재할 때, 어떤 변수가 결과(endpoint)에 독립적으로 유의미한지 규명하거나, 복잡한 비선형 관계를 가진 고차원 데이터에서 핵심 변수를 추출하여 모델의 일반화 가능성을 높이기 위해 사용했다 [3, 6, 9, 10, 11]. 특히 Propensity score matching 후 잔여 교란을 확인하거나, Deep learning 모델의 임상적 유용성을 통계적으로 검증하는 단계에서도 필수적으로 도입되었다 [4, 5, 12].

이 라이브러리에서의 사용 — 실제 적용 방식은 크게 두 가지로 나뉜다. 첫째, 고차원 영상 데이터(Radiomics) 분석에서는 LASSO regression을 통해 수백 개의 특징 중 예측력 있는 소수 변수(예: 9개의 Delta-RFs)를 선별하고, 이를 기반으로 Rad-score를 산출한 후 다변량 회귀로 최종 모델을 검증했다 [10]. 둘째, 임상 코호트 비교 연구에서는 Propensity score matching (PSM)으로 군 간 균형을 맞춘 후, 매칭된 표본에서 Multivariable Cox regression 등을 통해 치료 효과(Pembrolizumab monotherapy vs. combination)의 독립적 영향을 추정했다 [12]. 보고 방식은 대부분 Hazard ratio (HR), 95% Confidence Interval (CI), 그리고 P-value를 명시하며, 모델 성능 평가에는 Area Under the Curve (AUC)Receiver Operating Characteristic (ROC) curve 분석을 병행하여 진단 정확도를 함께 제시했다 [3, 4, 6, 9].

설계 시 주의점 — 이 논문들에서 드러나는 중요한 전제 조건은 표본의 대표성과 데이터의 재현성이다. Radiomics 기반 연구에서는 Region of Interest (ROI) segmentation의 신뢰도를 Intraclass correlation coefficient (ICC)로 검증하여 안정된 특징만 분석에 포함해야 한다 [6, 10]. 또한, 고차원 데이터를 다룰 때는 Overfitting을 방지하기 위해 Training/Validation/Test cohort로 엄격히 분리하거나, 외부 코호트(External validation cohort)에서의 검증이 필수적이다 [3, 4, 7, 8, 11]. 교란 요인 통제를 위해 PSM을 사용할 경우, 매칭 전후의 covariates 균형(예: Age, Sex, ECOG-PS)을 명확히 보고하고, 잔여 편향이 없는지 확인해야 한다 [12]. 마지막으로, 생존 분석(Survival analysis) 시 추적 기간(Follow-up period)과 검열 데이터(Censored data)의 처리 방식을 명시하여 결과의 해석 가능성을 높여야 한다 [1, 5, 9].

근거 논문

  1. Pulmonary Intravascular Lymphomatosis: Clinical, CT, and PET Findings, Correlation of CT and Pathologic Results, and Survival Outcome — Radiology · 2016
  2. Combined fluoroscopy-and CT-guided transthoracic needle biopsy using a C-arm cone-beam CT system: comparison with fluoroscopy-guided biopsy — Korean Journal of Radiology · 2011
  3. Predicting response to immunotherapy in advanced non-small-cell lung cancer using tumor mutational burden radiomic biomarker — Journal for Immunotherapy of Cancer · 2020
  4. Non-invasive prediction for pathologic complete response to neoadjuvant chemoimmunotherapy in lung cancer using CT-based deep learning: a multicenter study — Frontiers in Immunology · 2024
  5. Deep Learning to Estimate Biological Age From Chest Radiographs — JACC Cardiovascular Imaging · 2021
  6. Imaging Phenotyping Using Radiomics to Predict Micropapillary Pattern within Lung Adenocarcinoma — Journal of Thoracic Oncology · 2017
  7. Radiogenomic-based multiomic analysis reveals imaging intratumor heterogeneity phenotypes and therapeutic targets — Science Advances · 2023
  8. Development of an Immune-Pathology Informed Radiomics Model for Non-Small Cell Lung Cancer — Scientific Reports · 2018
  9. Spatial interplay patterns of cancer nuclei and tumor-infiltrating lymphocytes (TILs) predict clinical benefit for immune checkpoint inhibitors — Science Advances · 2022
  10. Delta-radiomics features combined with haematological index predict pathological complete response after neoadjuvant immunochemotherapy in resectable non-small cell lung cancer — Clinical Radiology · 2025
  11. Deciphering the tumor microenvironment through radiomics in non-small cell lung cancer: Correlation with immune profiles — PloS One · 2020
  12. Long-term survival comparison of first-line pembrolizumab versus pembrolizumab plus chemotherapy for patients with advanced non-small cell lung cancer: A multicenter propensity-matched cohort study — Lung Cancer · 2025
Logistic regression — 31편

개념 — Logistic regression은 이진(binary) 또는 다중 범주형(endpoint) 결과 변수와 하나 이상의 예측 변수(independent variables) 간의 관계를 모델링하여, 특정 사건 발생의 확률을 추정하거나 독립적인 연관성을 검정하는 방법이다. 본 라이브러리에서는 주로 질병 진행 여부, 치료 반응(pCR/MPR), biomarker 발현(PD-L1), 또는 병리학적 발견(pN+)과 같은 이진 분류 문제를 해결하기 위해 사용되었다.

언제 쓰나 — 연구자들은 primary endpoint가 명확한 이진 결과(예: malignancy vs benignity [2], PD-L1 ≥50% vs <50% [8], pCR 달성 여부 [5, 7])를 가진 retrospective cohort study 또는 case-control design에서 이 방법을 선택했다. 특히 고차원 데이터(radiomics features)에서 변수 선택(LASSO regression) 후 남은 소수의 예측 인자들과 임상 변수를 결합하여 최종 예측 모델을 구성하거나 [7], 대규모 데이터베이스(NCDB)에서 다변량 분석을 수행할 때 활용되었다 [12]. 또한, univariate analysis에서 유의한 결과를 보인 변수들을 multivariable model에 포함시켜 confounding 효과를 통제하고자 할 때 적용되었다 [10, 11].

이 라이브러리에서의 사용 — 실제 적용에서는 먼저 Chi-square test, t-test, 또는 Fisher’s exact test를 통해 기저 특성의 차이를 확인한 후, 유의한 변수들을 logistic regression 모델에 투입했다 [5, 10, 11]. 모델의 성능 평가는 주로 ROC curve와 Area Under the Curve (AUC)를 계산하여 수행되었으며, 일부 연구에서는 DeLong’s test를 사용하여 모델 간 AUC 차이의 통계적 유의성을 검증하기도 했다 [5]. 보고 방식은 Odds Ratio (OR)과 그 95% Confidence Interval (CI), 그리고 p-value를 제시하는 것이 표준이었다 [12]. 표본 규모는 소규모 코호트(n=44~250)부터 대규모 데이터베이스(n=109,964)까지 다양했으나, 모두 multivariable logistic regression을 통해 독립적인 예측 인자를 도출했다.

설계 시 주의점 — 고차원 radiomics data를 다룰 경우, overfitting 방지를 위해 LASSO regression 등 변수 선택 기법을 먼저 적용한 후 logistic regression을 수행하는 것이 필수적이다 [7]. 또한, temporal confounding이나 operator bias와 같은 편향을 통제하기 위해 inclusion/exclusion criteria를 엄격히 설정하고, 가능하면 multicenter design이나 external validation cohort를 구성하여 모델의 generalizability를 확인해야 한다 [2, 4, 5]. 변수 간 다중공선성(multicollinearity)을 확인하지 않고 모든 변수를 무작위로 투입할 경우 결과 해석이 왜곡될 수 있으므로, 변수 선택 과정과 모델 적합도 평가(calibration, discrimination)를 투명하게 보고해야 한다 [3, 8].

근거 논문

  1. Autoimmune Features or Connective Tissue Disease Shows a — 2026
  2. Combined fluoroscopy-and CT-guided transthoracic needle biopsy using a C-arm cone-beam CT system: comparison with fluoroscopy-guided biopsy — Korean Journal of Radiology · 2011
  3. The effect of feature normalization methods in radiomics — Insights Imaging · 2024
  4. Multicentre performance and consistency of two deep learning models for malignancy probability estimation of incidental pulmonary nodules — 2026
  5. Non-invasive prediction for pathologic complete response to neoadjuvant chemoimmunotherapy in lung cancer using CT-based deep learning: a multicenter study — Frontiers in Immunology · 2024
  6. Imaging Phenotyping Using Radiomics to Predict Micropapillary Pattern within Lung Adenocarcinoma — Journal of Thoracic Oncology · 2017
  7. Delta-radiomics features combined with haematological index predict pathological complete response after neoadjuvant immunochemotherapy in resectable non-small cell lung cancer — Clinical Radiology · 2025
  8. Imaging-Based Biomarkers Predict Programmed Death-Ligand 1 and Survival Outcomes in Advanced NSCLC Treated With Nivolumab and Pembrolizumab: A Multi-Institutional Study — JTO Clinical and Research Reports · 2023
  9. Deciphering the tumor microenvironment through radiomics in non-small cell lung cancer: Correlation with immune profiles — PloS One · 2020
  10. Dynamic (18) F-FDG PET/CT can predict the major pathological response to neoadjuvant immunotherapy in non-small cell lung cancer — Thorac Cancer · 2022
  11. PD-L1 Expression is Highly Associated with Tumor-Associated Macrophage Infiltration in Nasopharyngeal Carcinoma — Cancer Management and Research · 2020
  12. Quantifying the rate and predictors of occult lymph node involvement in patients with clinically node-negative non-small cell lung cancer — Acta Oncologica · 2022
Propensity score matching/weighting — 12편

개념 — Propensity score matching/weighting은 관찰 연구(observation studies)에서 치료군 간 기저 특성의 불균형을 보정하여 교란 변수(confounding)의 영향을 줄이고, 인과적 효과를 추정하기 위한 방법이다 [11]. 이 방법은 treatment assignment의 확률을 모델링(예: logistic regression)하여 각 환자에게 점수를 부여하고, 이를 통해 유사한 특성을 가진 환자들을 짝지거나 가중치를 적용함으로써 선택 편향(selection bias)을 최소화한다 [2, 9, 11]. 궁극적인 목표는 무작위 대조 시험(RCT)에 가까운 비교 환경을 구축하여 치료 효과나 예후 차이를 더 정확하게 평가하는 것이다.

언제 쓰나 — 이 라이브러리의 논문들은 주로 retrospective cohort study 또는 real-world data를 분석할 때, 두 개 이상의 치료군(예: pembrolizumab monotherapy vs combination therapy [2], neoadjuvant immunochemotherapy vs chemotherapy alone [9])을 비교해야 할 상황에서 이 방법을 선택했다. 특히 기저 특성(baseline characteristics)이 치료군 간에 현저히 다를 것으로 예상되거나, 실제 임상 데이터에서 무작위 배정이 이루어지지 않은 경우 교란 효과를 통제하기 위해 적용되었다 [2, 9]. 또한 meta-analysis에서 직접 비교가 가능한 연구만 선별하거나 [3], 다양한 고형암 환자군의 진행성 질환 패턴을 분석할 때 편향을 줄이기 위한 도구로 활용된 것으로 보인다.

이 라이브러리에서의 사용 — 실제 적용 방식은 주로 1:1 nearest-neighbor propensity score matching (PSM)이 가장 흔하며, caliper size는 연구에 따라 0.2 [2] 또는 0.05 [9]로 설정되었다. 일부 연구에서는 without replacement 방식으로 매칭을 수행하여 최종 분석 표본의 균형을 맞췄다 [9]. Primary endpoint로는 overall survival (OS), progression-free survival (PFS), disease-free survival (DFS) 및 pathologic complete response (pCR) 등이 주로 사용되었으며, 매칭 후 Kaplan-Meier 추정과 log-rank test를 통해 생존 곡선의 차이를 검정하는 것이 표준적인 보고 방식이었다 [2, 9]. 표본 규모는 단일 기관 연구에서 수백 명 [2, 9]부터 다기관 연구에서 천 명 이상 [9]에 이르기까지 다양하지만, 매칭 과정에서 일부 데이터가 제외되어 최종 분석 cohort의 크기가 초기 모집단보다 작아지는 특징이 있다.

설계 시 주의점 — PSM을 적용할 때는 포함할 covariates (예: age, sex, ECOG-PS 등)를 신중하게 선정해야 하며, 매칭 후 잔여 불균형(residual imbalance)이 없는지 반드시 확인해야 한다 [2]. Caliper size의 설정은 중요하며, 너무 넓으면 매칭의 질이 떨어지고 너무 좁으면 표본 손실이 커질 수 있으므로 연구 설계 단계에서 적정 값을 결정해야 한다 [2, 9]. 또한 매칭으로 인해 표본 수가 감소할 수 있으므로, 통계적 검정력(power)을 고려하여 초기 모집단 크기를 충분히 확보하는 것이 필요하다. 마지막으로, propensity score 모델링에 사용된 변수와 최종 분석 모델의 일관성을 유지하며, 매칭 전후의 기저 특성 비교 결과를 투명하게 보고해야 신뢰성이 높아진다 [2, 9].

근거 논문

  1. Development of an Immune-Pathology Informed Radiomics Model for Non-Small Cell Lung Cancer — Scientific Reports · 2018
  2. Long-term survival comparison of first-line pembrolizumab versus pembrolizumab plus chemotherapy for patients with advanced non-small cell lung cancer: A multicenter propensity-matched cohort study — Lung Cancer · 2025
  3. Sublobar resection, stereotactic body radiotherapy, and thermal ablation for early-stage non-small cell lung cancer a systematic review and meta-analysis — Lung Cancer · 2026
  4. Clinical outcomes and neuroendocrine features of transformed versus primary small-cell lung cancer — Lung Cancer · 2025
  5. Lee et al. - Lung Cancer 2026 - Surgical Attrition After Neoadjuvant Chemoimmunotherapy for Non-Small Cell Lung Cancer: Real-World Experience and Predictors — Lung Cancer · 2026
  6. Efficacy of neoadjuvant immunochemotherapy and survival surrogate analysis of neoadjuvant treatment in IB–IIIB lung squamous cell carcinoma — Scientific Reports · 2024
  7. Dissection of Progressive Disease Patterns for a Modified Classification for Immunotherapy — JAMA Oncology · 2025
  8. Deep learning for predicting major pathological response to neoadjuvant chemoimmunotherapy in non-small cell lung cancer: A multicentre study — EBioMedicine · 2022
  9. Treatment patterns and clinical outcomes of patients with resectable non–small cell lung cancer receiving neoadjuvant immunochemotherapy: A large-scale, multicenter, real-world study (NeoR-World) — The Journal of Thoracic and Cardiovascular Surgery · 2024
  10. A radiological predictor for pneumomediastinum/pneumothorax in COVID-19 ARDS patients — Journal of Critical Care · 2020
  11. An introduction to propensity score methods for reducing the effects of confounding in observational studies — Multivariate Behavioral Research · 2011
  12. Health system-scale language models are all-purpose prediction engines — Nature · 2023

연구 설계

Subgroup analysis — 32편

개념 — 이 라이브러리의 논문들은 주로 치료군 간 비교, 바이오마커의 예후적/예측적 가치 검증, 또는 새로운 진단 기준 수립을 위해 Subgroup analysis를 활용한다. 이는 전체 코호트 내 특정 하위 집단(예: PD-L1 발현 수준, 병리학적 반응 정도, 림프절 상태)에서 치료 효과나 생존 이득이 어떻게 달라지는지를 정량화하거나, 특정 임상적 특징을 가진 환자군에서의 진단 정확도와 안전성을 평가하는 것을 목적으로 한다.

언제 쓰나 — 연구자들은 치료 반응에 영향을 미칠 수 있는 생물학적 또는 임상적 변수(예: PD-L1 TPS [4, 10], TILs 공간 패턴 [3], TAMs 침윤 밀도 [6])가 존재할 때 이를 적용한다. 또한, 대규모 데이터베이스(NCDB)를 활용한 관찰 연구에서 특정 임상 상황(cN0 NSCLC [11])의 위험도를 정량화하거나, 기존 기준(RECIST)의 한계를 보완하기 위한 새로운 정량적 지표(PERCIST [2], HistoTIL [3])를 도출하고 검증할 때 사용된다. 특히, Phase 3 RCT의 planned final analysis나 exploratory analysis에서 주요 endpoint(EFS, OS) 외에 병리학적 반응(pCR, %RVT)과의 연관성을 규명하기 위해 활용되었다 [7, 8, 9].

이 라이브러리에서의 사용 — 분석 방법은 연구 설계에 따라 다양하게 적용되었다. RCT 및 코호트 연구에서는 Kaplan-Meier curves와 log-rank test를 통해 생존 차이(EFS, OS)를 시각화하고 Hazard ratio(HR)을 계산했다 [3, 7, 8, 9]. 진단 정확도나 연관성 분석에서는 Chi-square test, Kruskal–Wallis test, Multivariable logistic regression (OR), ROC curves (cutoff 도출) 등을 사용했다 [1, 6, 10, 11]. 표본 규모는 단일 기관의 소규모 코호트(74~212명 [1, 3, 5, 6])부터 국가 데이터베이스 기반의 대규모 연구(109,964명 [11]) 및 메타분석(1,893명 [12])까지 포괄한다. 보고 방식은 주로 Hazard ratio (HR), Odds Ratio (OR), Confidence Interval (CI), P-value를 명시하며, 일부 연구에서는 Propensity Score Matching (PSM)을 통해 교란 변수를 보정 후 분석 결과를 제시했다 [4].

설계 시 주의점 — 관찰성 연구에서는 Temporal confounding [1]이나 선택 편향을 통제하기 위해 PSM [4] 또는 Multivariable regression [6, 11]을 적용해야 한다. 바이오마커 cutoff 값을 도출할 때는 Training group과 Validation group으로 나누어 검증하는 과정이 필수적이며 [10], 생존 분석 시에는 Censoring 정보를 정확히 재구성하거나 Stratified log-rank test를 사용하여 군 간 비교의 타당성을 확보해야 한다 [9, 12]. 또한, Subgroup analysis 결과의 임상적 의미를 해석할 때는 Statistical significance뿐만 아니라 Clinical relevance(예: MPR 비율, pCR 정의)를 명확히 정의하고 보고해야 하며, 특히 exploratory analysis인 경우 과잉 해석을 피하기 위해 Prespecified hypothesis와 결과를 구분하여 서술해야 한다 [7, 8].

근거 논문

  1. Combined fluoroscopy-and CT-guided transthoracic needle biopsy using a C-arm cone-beam CT system: comparison with fluoroscopy-guided biopsy — Korean Journal of Radiology · 2011
  2. From RECIST to PERCIST: Evolving Considerations for PET response criteria in solid tumors — Journal of Nuclear Medicine · 2009
  3. Spatial interplay patterns of cancer nuclei and tumor-infiltrating lymphocytes (TILs) predict clinical benefit for immune checkpoint inhibitors — Science Advances · 2022
  4. Long-term survival comparison of first-line pembrolizumab versus pembrolizumab plus chemotherapy for patients with advanced non-small cell lung cancer: A multicenter propensity-matched cohort study — Lung Cancer · 2025
  5. Perioperative Pembrolizumab for Locally Advanced Thymic Epithelial Tumors: A Single-Arm, Phase 2 Trial — Journal of Thoracic Oncology · 2025
  6. PD-L1 Expression is Highly Associated with Tumor-Associated Macrophage Infiltration in Nasopharyngeal Carcinoma — Cancer Management and Research · 2020
  7. Association between pathologic response and survival after neoadjuvant therapy in lung cancer — Nature Medicine · 2024
  8. Neoadjuvant Nivolumab plus Chemotherapy in Resectable Lung Cancer — New England Journal of Medicine · 2022
  9. Overall Survival with Neoadjuvant Nivolumab plus Chemotherapy in Lung Cancer — New England Journal of Medicine · 2025
  10. Pembrolizumab for the treatment of non-small-cell lung cancer — New England Journal of Medicine · 2015
  11. Quantifying the rate and predictors of occult lymph node involvement in patients with clinically node-negative non-small cell lung cancer — Acta Oncologica · 2022
  12. Systematic review and meta-analysis of the prognostic impact of lymph node micrometastasis and isolated tumour cells in patients with stage I-IIIA non-small cell lung cancer — Histopathology · 2023
Systematic review / Meta-analysis — 28편

개념 — 이 라이브러리의 연구들은 특정 임상 질문(예: 치료 효능, 예후 인자, 진단 정확도)에 대해 기존 문헌을 체계적으로 수집·평가하여 증거를 종합하는 방법론을 사용한다. 이는 단일 코호트의 원시 데이터 분석이 아닌, 다수의 선행 연구를 메타-분석(meta-analysis), 네트워크 메타-분석(network meta-analysis), 또는 서술적 합성(narrative synthesis)을 통해 정량적 또는 질적으로 통합하는 것을 목표로 한다 [1][2][3][4][5][6][7][8][9][10][11][12].

언제 쓰나 — 직접적인 임상 시험 수행이 불가능하거나, 서로 다른 치료법 간의 직접 비교 데이터가 부족할 때 사용된다. 예를 들어, early-stage NSCLC에서 SLR, SBRT, IGTA의 상대적 효능을 비교하기 위해 PSM 연구와 IPD를 결합한 경우 [3], advanced NSCLC에서 다양한 first-line immunotherapy combinations의 순위를 결정하기 위해 Bayesian network meta-analysis를 적용한 경우 [12], 또는 RECIST와 PERCIST/EORTC 기준 간 일치도를 평가하기 위해 pooled analysis를 수행한 경우 [10] 등이 해당한다. 또한, radiomics 연구의 품질을 메타-평가(meta-evaluation)하거나 [2], PD-L1 assay의 비교 가능성을 체계적으로 검토할 때 [11]도 이 방법이 선택되었다.

이 라이브러리에서의 사용 — 실제 적용 방식은 연구 목적에 따라 다양하다. 정량적 분석에서는 Freeman-Tukey double-arcsine transformation model을 이용한 proportional meta-analysis [7], multivariate model for joint analysis of survival proportions를 통한 시간별 생존 데이터 재구성 [8], 그리고 unweighted κ statistics를 활용한 진단 기준 일치도 평가 [10]가 사용되었다. 표본 규모는 소규모(예: 216명 [10])부터 대규모(예: 10,493명 [7], 8,278명 [12])까지 다양하며, primary endpoint로는 overall survival(OS), progression-free survival(PFS), cancer-specific survival(CSS) [3][7][8][12] 및 objective response rate(ORR) [10][12] 등이 주로 설정되었다. 보고 방식은 PRISMA 가이드라인 준수 [11], QUADAS-2 또는 QUIPS 도구를 통한 bias risk 평가 [11], 그리고 heterogeneity 고려를 위한 통계적 모델링 [7][8]이 공통적으로 관찰된다.

설계 시 주의점 — 연구 설계 시 포함/제외 기준의 명확성(예: PSM된 연구만 포함 [3], high risk of bias 연구 제외 [11])과 데이터 추출의 엄격함(예: 두 명의 연구자 독립 screening [11])이 필수적이다. 또한, survival data를 재구성할 때는 censoring 정보와 사건 수(events)의 정확한 추출이 중요하며 [8], indirect comparison을 수행할 때는 Bayesian framework와 같은 적절한 통계적 접근법이 필요하다 [12]. 흔한 함정으로는 HYD 정의의 이질성으로 인한 발생률 편차 [5]나, radiomics 연구에서의 방법론적 엄격성 부족 [2]가 있으며, 이를 피하기 위해 표준화된 평가 도구(RQS 2.0 등)의 적용과 명확한 endpoint 정의(pCR, MPR, EFS 등 [4])가 요구된다. 마지막으로, review 형식의 논문이라도 통계적 가정(예: proportional hazards)이나 결측치 처리에 대한 명시적 보고가 필요하며 [1], software 버전 및 분석 코드 등의 재현성 확보 요소도 고려해야 한다.

근거 논문

  1. Complications of Lung Transplantation: Update on Imaging Manifestations and Management — Radiology: Cardiothoracic Imaging · 2021
  2. Radiomics Quality Score 2.0: towards radiomics readiness levels and clinical translation for personalized medicine — Nature Reviews Clinical Oncology · 2025
  3. Sublobar resection, stereotactic body radiotherapy, and thermal ablation for early-stage non-small cell lung cancer a systematic review and meta-analysis — Lung Cancer · 2026
  4. The rapidly evolving paradigm of neoadjuvant immunotherapy across cancer types — Nature Cancer · 2025
  5. Hyperprogressive disease in non-small cell lung cancer treated with immune checkpoint inhibitor therapy, fact or myth? — Frontiers in Oncology · 2022
  6. Pathological extranodal extension in head and neck cancer: A prognostic biomarker with therapeutic ramifications and diagnostic pitfalls — Pathology - Research and Practice · 2025
  7. Attrition with adjuvant, neoadjuvant, and perioperative immunotherapy-based treatment protocols in patients with resectable non-small-cell lung cancer. A meta-analysis of prospective trials — Lung Cancer · 2025
  8. Systematic review and meta-analysis of the prognostic impact of lymph node micrometastasis and isolated tumour cells in patients with stage I-IIIA non-small cell lung cancer — Histopathology · 2023
  9. Eradicating micrometastases with immune checkpoint blockade: Strike while the iron is hot — Cancer Cell · 2021
  10. Comparison of the morphologic criteria (RECIST) and metabolic criteria (EORTC and PERCIST) in tumor response assessments: a pooled analysis — The Korean Journal of Internal Medicine · 2019
  11. Comparability of PD‐L1 immunohistochemistry assays for non‐small‐cell lung cancer a systematic review — 2020
  12. Efficacy and Safety of First-Line Immunotherapy Combinations for Advanced NSCLC: A Systematic Review and Network Meta-Analysis — Journal of Thoracic Oncology · 2021

모델 검증

Cross-validation — 17편

개념 — 이 라이브러리의 논문들은 주로 고차원 데이터(high-dimensional data)나 복잡한 영상 특징(radiomics features)을 기반으로 한 예측 모델의 일반화 성능(generalizability)과 과적합(overfitting) 방지를 위해 Cross-validation을 활용한다. 이는 제한된 표본 크기(N=51~922 등 [1])에서 모델의 안정성을 평가하고, 최종 선택된 특징(feature selection)이 새로운 데이터에서도 일관된 예측력(predictive performance)을 유지하는지 검증하기 위한 핵심 방법론이다. 특히 Radiomics 파이프라인에서는 feature normalization 및 selection 과정 후, fivefold stratified cross-validation with 100 repeats [1] 또는 내부/외부 검증 cohort를 통해 모델의 robustness를 확인한다.

언제 쓰나 — 이 연구들은 주로 표본 수(N)이 상대적으로 작거나, 특성(feature)의 수가 표본 수보다 훨씬 많은 high-dimensional setting에서 Cross-validation을 선택했다 [1]. 예를 들어, NSCLC 환자의 pCR 예측 [2, 7]이나 MPR 예측 [9]과 같은 binary classification 문제, 그리고 OS/PFS와 같은 time-to-event 분석 [5, 6, 8]에서 모델 개발 단계의 성능 평가에 사용되었다. 또한, 단일 기관 데이터(single-institution cohort) [4, 5, 9, 10]나 소규모 코호트(N=44 [9])에서 외부 검증 cohort가 부족할 때, 내부 데이터를 반복적으로 분할하여 신뢰도 높은 성능 지표를 도출하기 위해 적용되었다.

이 라이브러리에서의 사용 — 실제 적용 방식은 연구마다 차이가 있으나, 공통적으로 training set과 validation set으로 데이터를 분할한 후, training set 내에서 cross-validation을 수행하거나 전체 데이터에 대해 k-fold (주로 5-fold)를 적용했다 [1, 7]. 주요 endpoint로는 AUC(area under the curve) [1, 2, 4, 7], sensitivity/specificity [1], 그리고 survival 분석에서의 hazard ratio(HR) 및 Kaplan-Meier 곡선 비교 [6, 8, 10]가 보고되었다. 표본 규모는 N=51~922 [1]부터 N=323 [8]까지 다양하며, 일부 연구에서는 LASSO regression을 통한 feature selection 후 cross-validation으로 최종 모델의 성능을 확정했다 [7]. 보고 방식은 주로 ROC curve 기반의 AUC 값과 함께 calibration metrics (Brier score, expected calibration error) [1] 또는 decision curve analysis [4]를 병행하여 임상적 유용성을 제시했다.

설계 시 주의점 — 이 논문들에서 드러나는 중요한 전제 조건은 충분한 표본 크기 확보와 적절한 데이터 분할 전략이다. Rule of 10(또는 15)에 따라 최종 모델의 feature당 충분한 sample을 확보해야 하며 [3], 그렇지 않을 경우 overfitting 위험이 크다. 흔한 함정으로는 segmentation이나 feature extraction 과정에서의 재현성(reproducibility) 부족으로 인한 편향(bias)이 있으며, 이를 방지하기 위해 intraclass correlation coefficient (ICC)를 통한 신뢰도 평가가 선행되어야 한다 [4, 7]. 또한, cross-validation 결과만 보고하는 것이 아니라, independent external validation cohort [2, 5, 8] 또는 temporal split을 통해 모델의 실제 임상 적용 가능성을 반드시 검증해야 하며, feature normalization 방법(z-Score, Min-Max 등)이 최종 성능에 미치는 영향을 명시적으로 보고해야 한다 [1].

근거 논문

  1. The effect of feature normalization methods in radiomics — Insights Imaging · 2024
  2. Non-invasive prediction for pathologic complete response to neoadjuvant chemoimmunotherapy in lung cancer using CT-based deep learning: a multicenter study — Frontiers in Immunology · 2024
  3. Radiomics in Oncology: A Practical Guide — RadioGraphics · 2021
  4. Imaging Phenotyping Using Radiomics to Predict Micropapillary Pattern within Lung Adenocarcinoma — Journal of Thoracic Oncology · 2017
  5. Development of an Immune-Pathology Informed Radiomics Model for Non-Small Cell Lung Cancer — Scientific Reports · 2018
  6. Spatial interplay patterns of cancer nuclei and tumor-infiltrating lymphocytes (TILs) predict clinical benefit for immune checkpoint inhibitors — Science Advances · 2022
  7. Delta-radiomics features combined with haematological index predict pathological complete response after neoadjuvant immunochemotherapy in resectable non-small cell lung cancer — Clinical Radiology · 2025
  8. Imaging-Based Biomarkers Predict Programmed Death-Ligand 1 and Survival Outcomes in Advanced NSCLC Treated With Nivolumab and Pembrolizumab: A Multi-Institutional Study — JTO Clinical and Research Reports · 2023
  9. Dynamic (18) F-FDG PET/CT can predict the major pathological response to neoadjuvant immunotherapy in non-small cell lung cancer — Thorac Cancer · 2022
  10. Lee et al. - Lung Cancer 2026 - Surgical Attrition After Neoadjuvant Chemoimmunotherapy for Non-Small Cell Lung Cancer: Real-World Experience and Predictors — Lung Cancer · 2026
  11. Efficacy and Safety of First-Line Immunotherapy Combinations for Advanced NSCLC: A Systematic Review and Network Meta-Analysis — Journal of Thoracic Oncology · 2021
  12. The journey of tumor-infiltrating lymphocytes as a biomarker in breast cancer: clinical utility in an era of checkpoint inhibition — Annals of Oncology · 2021
External validation — 57편

개념 — External validation은 개발된 예측 모델(특히 Radiomics 또는 Deep Learning 기반)의 일반화 성능(generalizability)과 과적합(overfitting) 여부를 평가하기 위해, 모델 학습(training) 및 내부 검증(internal validation)에 사용되지 않은 독립적인 외부 데이터셋을 적용하는 과정이다. 이는 모델이 특정 기관이나 코호트의 편향(bias) 없이 다른 환경에서도 동일한 진단 정확도나 예후 예측 능력을 유지하는지 확인하는 핵심 단계이다.

언제 쓰나 — 이 라이브러리의 논문들은 주로 새로운 알고리즘(Deep learning, Radiomics signature)을 개발한 후, 그 임상적 유용성과 신뢰성을 입증하기 위해 external validation을 수행했다 [4, 7, 8, 9, 11, 12]. 특히 단일 기관 데이터로 모델을 학습시킨 경우 [12], 또는 공개 데이터셋(NIH, PadCHEST 등)으로 pre-training한 후 특정 임상 코호트(PLCO, NLST)에서 fine-tuning 및 검증하는 경우 [9]에 필수적으로 적용되었다. 또한, Radiomics feature의 보편성(universality)을 확인하기 위해 서로 다른 장기(폐, 신장, 뇌)나 다른 기관(FUSCC vs DUKE/TCGA)의 데이터를 테스트 세트로 사용하는 경우도 포함된다 [5, 11].

이 라이브러리에서의 사용 — 실제 적용 방식은 크게 두 가지로 나뉜다. 첫째, 완전히 독립적인 외부 코호트(다른 병원 또는 공개 데이터셋)를 test set으로 사용하여 모델 성능(AUC, sensitivity, specificity)을 평가하는 방식이다 [4, 7, 8, 9, 11]. 예를 들어, MIMIC-IV와 내부 ICU 데이터를 training/validation에 쓰고 PadChest나 NLST를 external validation에 사용했다 [4, 9]. 둘째, subsampling 전략이나 feature robustness를 검증하기 위해 ground truth와의 차이를 정량화하는 방식이다 [2, 5]. Endpoint는 주로 진단 정확도(AUROC/AUC) [4, 7, 8] 또는 생존 분석(Overall Survival, Recurrence-Free Survival) [9, 11, 12]이었다. 표본 규모는 수백 명에서 수만 명(N=36,542~146,717)까지 다양하며, 보고 방식은 DeLong’s test를 통한 AUC 비교 [8]나 bootstrap resampling을 통한 신뢰구간 추정 [4] 등 통계적 유의성을 강조한다.

설계 시 주의점 — External validation을 설계할 때는 데이터의 이질성(heterogeneity)과 전처리(preprocessing) 표준화가 중요하다. 서로 다른 장비(CT vs MRI, HU vs arbitrary unit)나 기관 간 영상 강도(intensity) 분포 차이를 보정하기 위해 z-score normalization 등의 표준화 과정이 선행되어야 한다 [5]. 또한, 모델 개발 시 'Rule of 10' 또는 'Rule of 15'(feature당 충분한 sample 수 확보)를 준수하여 과적합을 방지해야 하며 [10], external validation cohort의 임상 특성(예: disease prevalence)이 training cohort와 유사하거나 실제 임상 상황(simulated reading test)을 반영하도록 설계해야 한다 [7]. 마지막으로, temporal confounding이나 operator bias가 결과에 영향을 미치지 않도록 연구 기간과 시술자 변수를 통제하거나 보고해야 한다 [1].

근거 논문

  1. Combined fluoroscopy-and CT-guided transthoracic needle biopsy using a C-arm cone-beam CT system: comparison with fluoroscopy-guided biopsy — Korean Journal of Radiology · 2011
  2. Investigation of radiomics based intra-patient inter-tumor heterogeneity and the impact of tumor subsampling strategies — Scientific Reports · 2022
  3. Criteria for the diagnosis of extranodal extension detected on radiological imaging in head and neck cancer: Head and Neck Cancer International Group consensus recommendations — The Lancet Oncology · 2024
  4. Multimodal deep learning for integrating chest radiographs and clinical parameters: a case for transformers — Radiology · 2023
  5. Are radiomics features universally applicable to different organs? — Cancer Imaging · 2021
  6. Mass estimates by computed tomography: physical density from CT numbers — American Journal of Roentgenology · 1984
  7. Development and validation of a deep learning algorithm detecting 10 common abnormalities on chest radiographs — European Respiratory Journal · 2021
  8. Non-invasive prediction for pathologic complete response to neoadjuvant chemoimmunotherapy in lung cancer using CT-based deep learning: a multicenter study — Frontiers in Immunology · 2024
  9. Deep Learning to Estimate Biological Age From Chest Radiographs — JACC Cardiovascular Imaging · 2021
  10. Radiomics in Oncology: A Practical Guide — RadioGraphics · 2021
  11. Radiogenomic-based multiomic analysis reveals imaging intratumor heterogeneity phenotypes and therapeutic targets — Science Advances · 2023
  12. Development of an Immune-Pathology Informed Radiomics Model for Non-Small Cell Lung Cancer — Scientific Reports · 2018

일치도

Inter-observer agreement (kappa/ICC) — 12편

개념 — Inter-observer agreement는 두 명 이상의 관찰자(예: Radiologists, Pathologists)가 동일한 영상이나 조직 슬라이드를 평가할 때 얼마나 일치하는지를 정량화하여 측정의 신뢰성(Reliability)과 재현성(Reproducibility)을 추정하는 방법이다. 이 지표는 주관적인 판단이 개입되는 영상 판독(Image interpretation)이나 병리학적 분류(Pathological classification)에서 관찰자 간 편차(Observer bias)를 통제하고, 연구 결과의 타당성을 확보하기 위해 필수적으로 수행된다.

언제 쓰나 — 본 라이브러리의 논문들은 주로 Radiomics feature 추출 전의 ROI segmentation [5], [6]이나, 영상 소견(E.g., ENE, Macklin effect)의 존재 여부 및 등급 분류 [2], [8], [9], [11] 시점에 이 방법을 적용했다. 특히 전문가 패널을 통한 진단 기준 합의 과정 [2]이나, TILs 정량화와 같은 정밀한 병리학적 측정 [10]에서도 관찰자 간 일치도를 확인하는 단계로 사용되었다. 이는 연구의 Primary endpoint가 영상 또는 병리 판독에 의존할 때, 그 판독 결과가 개인차가 아닌 객관적인 사실에 기반함을 입증하기 위해 선택된다.

이 라이브러리에서의 사용 — 대부분의 연구는 두 명의 독립된 전문가(Independent radiologists/pathologists)가 Blinded 상태로 영상을 평가한 후 일치도를 계산하는 방식을 취했다 [8], [9], [11]. 통계적 지표로는 범주형 변수(Categorical variables, e.g., ENE presence/absence)에는 Kappa coefficient를, 연속형 변수나 순위형 변수(Ordinal variables, e.g., segmentation stability, TILs quantification)에는 Intraclass correlation coefficient (ICC)를 주로 사용했다 [5], [6], [10]. 보고 방식은 단순히 일치도 수치(Kappa 또는 ICC 값)를 제시하는 것을 넘어, ICC ≥ 0.8과 같은 임계값을 설정하여 신뢰할 수 있는 Features만 선별하거나 [6], 합의도(Consensus threshold)를 명시하여 진단 기준의 강도를 정의하기도 했다 [2].

설계 시 주의점 — 연구 설계 단계에서 관찰자 간 일치도 분석을 위한 별도의 Sub-sample (예: 전체 코호트의 일부, e.g., n=49 [5], n=30 [6])을 미리 선정하거나, 모든 대상에 대해 두 명의 판독자를 배치해야 한다. 특히 Radiomics 연구에서는 ROI segmentation의 재현성이 모델 성능에 직접적인 영향을 미치므로, ICC를 통해 Stable features만 선별하는 과정이 필수적이다 [6]. 또한, Kappa나 ICC 값뿐만 아니라 그 해석 기준(예: Strong agreement threshold)을 사전에 정의하고, 판독자가 서로 독립적으로(Blinded) 평가했음을 명시하여 편향을 배제해야 한다 [8], [11].

근거 논문

  1. Pulmonary Intravascular Lymphomatosis: Clinical, CT, and PET Findings, Correlation of CT and Pathologic Results, and Survival Outcome — Radiology · 2016
  2. Criteria for the diagnosis of extranodal extension detected on radiological imaging in head and neck cancer: Head and Neck Cancer International Group consensus recommendations — The Lancet Oncology · 2024
  3. Are radiomics features universally applicable to different organs? — Cancer Imaging · 2021
  4. Radiomics in Oncology: A Practical Guide — RadioGraphics · 2021
  5. Imaging Phenotyping Using Radiomics to Predict Micropapillary Pattern within Lung Adenocarcinoma — Journal of Thoracic Oncology · 2017
  6. Delta-radiomics features combined with haematological index predict pathological complete response after neoadjuvant immunochemotherapy in resectable non-small cell lung cancer — Clinical Radiology · 2025
  7. Pathological extranodal extension in head and neck cancer: A prognostic biomarker with therapeutic ramifications and diagnostic pitfalls — Pathology - Research and Practice · 2025
  8. Radiologic Extranodal Extension of Metastatic Lymph Nodes in Patients With Non-Small Cell Lung Cancer: Prognostic Utility and Diagnostic Performance — American Journal of Roentgenology · 2023
  9. Radiologic Extranodal Extension of Metastatic Lymph Nodes in Patients With Non–Small Cell Lung Cancer: Prognostic Utility and Diagnostic Performance — American Journal of Roentgenology · 2023
  10. The journey of tumor-infiltrating lymphocytes as a biomarker in breast cancer: clinical utility in an era of checkpoint inhibition — Annals of Oncology · 2021
  11. A radiological predictor for pneumomediastinum/pneumothorax in COVID-19 ARDS patients — Journal of Critical Care · 2020
  12. Pulmonary involvement in Churg-Strauss syndrome: an analysis of CT, clinical, and pathologic findings — European Radiology · 2007

실전 체크리스트 — 연구 설계 전 확인

1. 연구 질문과 endpoint

2. 표본과 교란 변수

3. 통계 방법 선택

4. 검증 전략

5. 보고 시 빠뜨리기 쉬운 것