Interpretable multimodal deep learning for time-resolved survival prediction after hepatocellular carcinoma resection
이 논문의 관계도
List view
Neoadjuvant / Perioperative Immunotherapy연결 근거: perioperative
- Methods방법론 색인 — 어떤 분석을 어떤 논문이 썼나연결어: neoadjuvant, perioperative, resectable
- ReviewClinical Oncology 종합 리뷰연결어: neoadjuvant, perioperative, resectable
- Statistics통계 방법론 리뷰연결어: neoadjuvant, perioperative, resectable
- Yang 2024Treatment patterns and clinical outcomes of patients with resectable non-small cell lung cancer…연결어: neoadjuvant, perioperative, resectable
- Bang 2025Perioperative Pembrolizumab for Locally Advanced Thymic Epithelial Tumors: A Single-Arm, Phase…연결어: neoadjuvant, perioperative, pathologic response
- +44 more
Radiomics and Radiogenomics연결 근거: radiomics, radiomic
- Methods방법론 색인 — 어떤 분석을 어떤 논문이 썼나연결어: radiomics, radiomic, radiogenomic
- ReviewBiomedical Imaging 종합 리뷰연결어: radiomics, radiomic, radiogenomic
- Statistics통계 방법론 리뷰연결어: radiomics, radiomic, radiogenomic
- Su 2023Radiogenomic-based multiomic analysis reveals imaging intratumor heterogeneity phenotypes and t…연결어: radiomics, radiomic, radiogenomic
- Henry 2022Investigation of radiomics based intra-patient inter-tumor heterogeneity and the impact of tumo…연결어: radiomics, radiomic, subsampling
- +21 more
Shared Tags이 논문의 tags: AI, biomarker, deep learning, perioperative, radiomics
- Raghu 2021Deep Learning to Estimate Biological Age From Chest Radiographs공유 tag: AI, biomarker, deep learning
- Lambin 2025Radiomics Quality Score 2.0: towards radiomics readiness levels and clinical translation for pe…공유 tag: AI, deep learning, radiomics
- Lee 2025Risk factors and prognostic indicators for progressive fibrosing interstitial lung disease: a d…공유 tag: AI, biomarker, deep learning
- Wahl 2009From RECIST to PERCIST: Evolving Considerations for PET response criteria in solid tumors공유 tag: AI, biomarker
- Nam 2021Development and validation of a deep learning algorithm detecting 10 common abnormalities on ch…공유 tag: AI, deep learning
- +57 more
요약
Li et al. (2026)은 Hepatocellular carcinoma (HCC) 절제술 후 환자의 Overall survival (OS) 위험을 정적 분류가 아닌 시간 의존적으로 예측하기 위해 TEMPO-HCC라는 Multimodal deep learning 모델을 개발했다. 이 연구는 6개 기관의 1,475명 환자 코호트를 활용하여 Multiphasic MRI, H&E Whole-slide images (WSI), 그리고 Perioperative clinical predictors를 통합했다. TEMPO-HCC는 Discrete-time survival head를 사용하여 수술 후 1, 2, 3, 5년 시점별 사망 확률을 산출하며, External validation cohort에서 C-index 0.751을 기록했다. 기존 Guideline-based staging systems (BCLC, CNLC, AJCC) 및 Unimodal models보다 우수한 Discrimination 성능을 보였으며, 특히 BCLC와 비교했을 때 C-index가 0.598에서 0.751로 유의하게 향상되었다. Hierarchical interpretability를 통해 Radiologic phenotype과 Histopathologic pattern 간의 연관성을 설명 가능한 형태로 제시함으로써 Clinical trust를 높이는 데 기여했다.
방법
Study design 및 Cohort: 본 연구는 Retrospective cohort study로, 2012년 2월부터 2024년 1월까지 6개 기관에서 Curative-intent resection을 받은 Histopathologically confirmed HCC 환자 2,119명을 선별했다. Prior liver-directed therapy 시행자 (n=426), MRI 품질 불량자 (n=119), 주요 Clinical data 누락자 (n=99)를 제외하여 최종 분석 대상은 1,475명이었다. Institution 1-2의 환자 (n=702)는 Derivation cohort로 사용되었으며, 이를 8:2 비율로 Training cohort (n=562)와 Internal test cohort (n=140)로 분할했다. Institution 3-6의 환자 (n=773)는 External validation cohort로 구성되었다.
Data modalities 및 Model architecture: TEMPO-HCC는 세 가지 Modalities를 통합한다. 첫째, Preoperative multiphasic contrast-enhanced MRI (Pre-T1, Arterial, Portal venous, Delayed phases). 둘째, Postoperative H&E-stained whole-slide images (WSI). 셋째, Multivariable Cox regression을 통해 선별된 Independent clinical risk factors (Cirrhosis, Tumor size, Lymph node metastasis [LNMet], Alpha-fetoprotein [AFP])이다. 모델은 Discrete-time survival head를 사용하여 수술 후 12, 24, 36, 60개월 시점에서의 사망 확률을 생성한다.
Endpoint 및 Statistical method: Primary endpoint는 Time-resolved Overall survival (OS) risk prediction이다. 모델 성능 평가에는 C-index, Integrated Brier Score (IBS), 그리고 Time-dependent Area Under the Curve (AUC)를 사용했다. Calibration curves와 Decision curve analysis (DCA)를 통해 예측 정확도와 Clinical utility를 검증했으며, Kaplan-Meier 분석을 통해 High-risk 및 Low-risk 군의 Survival difference를 평가했다.
주요 결과
Predictive performance: External validation cohort에서 TEMPO-HCC는 C-index 0.751 (95% CI: 0.722–0.777)과 IBS 0.087 (95% CI: 0.076–0.099)을 기록했다. Time-dependent AUCs는 12개월 시점 0.836 (0.728–0.932), 24개월 시점 0.781 (0.726–0.827), 36개월 시점 0.812 (0.770–0.846), 60개월 시점 0.680 (0.619–0.744)이었다. Training cohort에서는 C-index 0.845, Internal test cohort에서는 C-index 0.787을 보였다. Calibration curves는 모든 Cohort에서 이상적인 대각선과 근접하여 Predicted risk와 Observed OS risk 간의 일치를 확인했다.
Comparison with staging systems 및 Ablation study: TEMPO-HCC는 BCLC (C-index 0.598), CNLC (0.633), AJCC (0.561)보다 External validation cohort에서 유의하게 우수한 Discrimination 성능을 보였다. TEMPO-HCC를 기존 Staging system에 추가했을 때 C-index와 Time-dependent AUC가 모두 유의하게 향상되었다 (all P < 0.001). Ablation study 결과, Clinical-only model은 C-index 0.685, WSI-only model은 0.752, MRI-only model은 0.792를 기록했다. Radiology와 WSI를 결합한 Bimodal model은 C-index 0.818을 보였으며, 최종 Tri-modal TEMPO-HCC (C-index 0.845)가 가장 높은 성능을 달성했다.
Subgroup analysis 및 Interpretability: External validation cohort의 Subgroup 분석에서 Tumor size, AFP level, Cirrhosis status에 따라 모델 성능이 일관되게 유지되었으며 통계적으로 유의한 차이는 없었다. Center별 C-index는 Center 3 (0.755), Center 4 (0.712), Center 5 (0.745), Center 6 (0.765)로 나타났으며, Center 4의 성능 저하는 상대적으로 작은 Sample size (n=104) 때문으로 추정되었다. Interpretability 분석에서 SHAP analysis는 AFP, Tumor size, Cirrhosis, LNMet를 주요 Risk driver로 식별했다. Grad-CAM maps은 MRI에서 Irregular tumor margins과 Heterogeneously enhancing regions에, WSI에서는 Pathologic lesions의 실제 위치와 일치하는 Hot spots을 강조하며 Normal tissue의 Attention을 효과적으로 억제했다.
통계 분석
분석 설계 — 이 연구는 HCC 절제술 후 환자의 개별화된 Overall Survival (OS) 위험 궤적을 시간별로 예측하는 것을 목표로 한다. 이를 위해 6개 기관에서 수집된 1,475명의 환자 데이터를 사용했으며, Institution 1-2의 702명을 Derivation cohort로, Institution 3-6의 773명을 External validation cohort로 구분했다. Derivation cohort는 Training (n=562)과 Internal test (n=140)으로 8:2 비율로 무작위 분할되었다. Primary endpoint는 수술 후 1, 2, 3, 5년 시점에서의 사망 확률이며, 이는 Discrete-time survival head를 통해 추정되었다.
무엇을 위해 어떤 분석을 썼는가 — 기저 임상 변수 중 OS의 독립적 예측 인자를 선별하기 위해 Multivariable Cox regression을 수행했으며, 이를 통해 Cirrhosis, Tumor size, LNMet, High AFP를 최종 임상 입력값으로 선정했다. 각 모달리티(Clinical, MRI, WSI)가 전체 모델 성능에 기여하는 정도를 정량화하고 다중 모달리티 융합의 유효성을 입증하기 위해 Ablation study를 실시하여 Unimodal 및 Bimodal baseline 모델과 비교했다. 예측된 위험도를 기반으로 High-risk와 Low-risk 군으로 이분화한 후, Kaplan-Meier 분석과 Log-rank test를 사용하여 두 군 간의 생존 곡선 차이의 통계적 유의성을 검정했다. 모델의 임상적 유용성(Net benefit)을 평가하기 위해 Decision curve analysis (DCA)를 수행했으며, 기존 가이드라인 기반 분류 체계(BCLC, CNLC, AJCC)와의 성능 비교를 통해 TEMPO-HCC의 추가적인 예후 판별력(Incremental prognostic value)을 C-index와 Time-dependent AUC로 입증했다. 통계 software 및 버전은 원문에 명시되지 않음.
방법론 평가 — 잘 된 점은 대규모 다기관 코호트(n=1,475)를 활용하여 External validation cohort를 별도로 구성함으로써 모델의 일반화 가능성(Generalizability)을 엄격하게 검증했다는 것이다. 또한, 단순한 이진 분류가 아닌 시간별(Time-resolved) 위험 예측을 제공하며 C-index와 Integrated Brier Score (IBS), Time-dependent AUC 등 다양한 판별력 및 보정(Calibration) 지표를 보고하여 모델 성능을 다각도로 평가했다. 의심스러운 점은 Deep learning 모델의 핵심 가정인 Proportional hazards assumption 검증 과정이 명시되지 않았다는 것이다. 또한, 결측치 처리(Missing data handling) 방법이 원문에 구체적으로 기술되어 있지 않아 데이터 전처리 과정의 투명성이 부족할 수 있다. 마지막으로, 여러 시점(1, 2, 3, 5년)과 다양한 비교군(Ablation study, Staging systems comparison)에 대한 통계적 검정이 반복되었음에도 Multiple testing 보정(예: Bonferroni correction 등)이 수행되었는지 여부가 명시되지 않아 False positive 위험을 배제하기 어렵다.
설계에 참고할 점 — 유사한 연구를 설계할 때는 이 논문처럼 Derivation cohort와 External validation cohort를 기관별로 명확히 분리하여 모델의 과적합(Overfitting)을 방지하고 일반화 성능을 입증하는 접근법을 채택해야 한다. 반면, Deep learning 모델의 해석 가능성과 통계적 엄밀성을 높이기 위해 결측치 처리 전략을 상세히 기술하고, 다중 시점 예측에 대한 Multiple testing 보정을 적용하며, 모델 가정(Assumptions) 검증을 추가하는 것이 필요하다.
강점
이 연구는 6개 기관의 대규모 다기관 코호트 (n=1,475)를 기반으로 External validation을 수행하여 모델의 Generalizability를 엄격하게 검증했다. 기존 정적 예측 모델을 넘어 Time-resolved prediction을 제공함으로써 Postoperative surveillance 계획 수립에 직접 활용 가능한 Dynamic risk assessment를 가능하게 했다. Hierarchical interpretability를 통해 Clinical features, MRI regions, Histopathologic patterns 간의 인과 관계를 추적 가능한 Evidence chain으로 제시하여 Black-box AI의 한계를 극복하고 Clinical trust를 높이는 데 성공했다.
한계
본 연구는 Retrospective design이며, External validation cohort 중 Center 4에서 Sample size가 작아 (n=104) 모델 성능이 다소 저하된 것으로 나타났다. MRI와 WSI 데이터의 전처리 및 Feature extraction 과정에서 기관별 Imaging protocols와 Pathology processing의 이질성이 존재할 수 있으며, 이는 모델의 Robustness에 영향을 미칠 수 있다. 또한, Genomic assays나 Molecular biomarkers가 통합되지 않아 Tumor biology의 미세한 차이를 완전히 반영하지 못할 가능성이 있다.
해석
TEMPO-HCC는 HCC 절제술 후 환자의 OS 위험을 시간적으로 정량화함으로써 Risk-adapted clinical management와 Personalized surveillance 전략 수립에 중요한 도구가 될 수 있다. 기존 Guideline-based staging systems과의 비교에서 입증된 Incremental prognostic value는 Multimodal deep learning이 Clinical decision-making에 실질적인 기여를 할 수 있음을 시사한다. 이 결과는 Oncology와 Biomedical Imaging 교차 분야에서 Computational pathology와 Radiomics의 융합이 단순한 성능 향상을 넘어 해석 가능한 Clinical insight를 제공할 수 있음을 보여준다. LLM Wiki 내 다른 Survival prediction 연구들과 비교할 때, TEMPO-HCC는 Time-dependent risk estimation과 Multi-level interpretability를 결합한 점에서 차별화된 접근법을 제시한다.