文章摘要
急性胰腺炎患者静脉血栓栓塞发生风险预测模型的构建
Establishment of a risk prediction model for venous thromboembolism in patients with acute pancreatitis
  
DOI:10.12089/jca.2026.04.002
中文关键词: 急性胰腺炎  静脉血栓栓塞  机器学习  预测模型
英文关键词: Acute pancreatitis  Venous thromboembolism  Machine learning  Prediction model
基金项目:新疆维吾尔自治区自然科学基金青年项目(2025D01C426);新疆医科大学校级自然科学青年研究项目(2024XYZR42)
作者单位E-mail
李玉倩 830054乌鲁木齐市,新疆医科大学第一附属医院麻醉科  
王佳玲 830054乌鲁木齐市,新疆医科大学第一附属医院麻醉科  
巴特苏荣·巴一娜 830054乌鲁木齐市,新疆医科大学第一附属医院麻醉科  
李文哲 830054乌鲁木齐市,新疆医科大学第一附属医院重症医学科  
叶建荣 830054乌鲁木齐市,新疆医科大学第一附属医院麻醉科 616227972@qq.com 
摘要点击次数: 777
全文下载次数: 190
中文摘要:
      
目的: 基于机器学习构建急性胰腺炎(AP)患者静脉血栓栓塞(VTE)发生风险的预测模型。
方法: 筛选MIMIC-Ⅳ(3.1版本)数据库中AP患者的临床资料,将符合条件的患者按照7∶3比例随机分为训练集和验证集,根据是否发生VTE将患者分为两组:非VTE组和VTE组。采用多因素Logistic回归筛选特征变量,基于Logistic回归、支持向量机(SVM)、梯度提升机(GBM)、自适应增强算法(AdaBoost)、随机森林(RF)、K最近邻算法(KNN)、神经网络及极限梯度提升(XGBoost)8种算法构建AP患者VTE发生风险的预测模型,以网格法进行参数调优并评估模型的预测效能。选择最优模型进行沙普利加性解释(SHAP)分析变量重要性。
结果: 共纳入患者1 553例,其中有318例(20.5%)发生VTE。多因素Logistic回归分析筛选出9个关键特征变量:心率、收缩压、血红蛋白、血小板计数、白蛋白、血糖、活化部分凝血活酶时间(APTT)、是否接受抗凝治疗及是否使用机械通气。8种机器学习算法构建的预测模型中,XGBoost模型表现最佳,其训练集和验证集的AUC及95%可信区间(CI)分别为0.946(95%CI 0.931~0.961)和0.856(95%CI 0.825~0.887)。决策曲线及校准曲线等提示XGBoost模型的综合预测效能及稳定性优于其他算法。SHAP分析显示,血糖、APTT、血红蛋白及白蛋白为模型中的重要变量,具有良好的临床应用可解释性。
结论: 基于机器学习构建的8种AP患者VTE发生风险的预测模型中,XGBoost模型表现出良好的预测效能和稳定性,有助于早期识别VTE高风险患者。
英文摘要:
      
Objective: To establish a predictive model for the risk of venous thromboembolism (VTE) in patients with acute pancreatitis (AP) based on machine learning.
Methods: Clinical data of AP patients were extracted from the MIMIC-IV (Version 3.1) database. The dataset was randomly split into a training set and a validation set at a ratio of 7∶3. Patients were divided into two groups according to the occurrence of VTE: non-VTE group and VTE group. Multivariate logistic regression was used to screen feature variables. Predictive models for VTE risk in AP patients were constructed using eight algorithms: logistic regression, support vector machine (SVM), gradient boosting machine (GBM), adaptive boosting (AdaBoost), random forest (RF), k-nearest neighbor (KNN), neural network, and extreme gradient boosting (XGBoost). Grid search was applied for parameter optimization, and the predictive performance of each model was evaluated. The optimal model was selected for Shapley additive explanations (SHAP) analysis to interpret the feature importance.
Results: A total of 1 553 patients were included, among whom 318 (20.5%) developed VTE. Multivariate logistic regression identified nine key feature variables: heart rate, systolic blood pressure, hemoglobin, platelet count, albumin, blood glucose, activated partial thromboplastin time (APTT), whether anticoagulant therapy was administered, and whether mechanical ventilation was used. Among the eight machine learning models, the XGBoost model exhibited the best performance, with an area under the curve (AUC) of 0.946 (95% CI 0.931-0.961) in the training set and 0.856 (95% CI 0.825-0.887) in the validation set. Decision curve and calibration curve indicated that the XGBoost model had superior comprehensive predictive performance and stability compared with other algorithms. SHAP analysis revealed that blood glucose, APTT, hemoglobin, and albumin were the most important variables in the model, with good clinical interpretability.
Conclusion: Eight predictive models for VTE risk in AP patients was constructed based on machine learning. The XGBoost model demonstrated excellent predictive performance and stability, which can facilitate early and accurate identification of AP patients at high risk of VTE.
查看全文   查看/发表评论  下载PDF阅读器
关闭