02047nas a2200145 4500000000100000000000100001008004100002100002300043700001600066700001800082245016700100856006100267490000700328520156600335 2022 d1 aWojciech Książek1 aFilip Turza1 aPawel Plawiak00aNCA-GA-SVM: A new two-level feature selection method based on neighborhood component analysis and genetic algorithm in hepatocellular carcinoma fatality prognosis uhttps://onlinelibrary.wiley.com/doi/abs/10.1002/cnm.35990 v383 a

Abstract Hepatocellular carcinoma (HCC) is one of the major challenges facing biomedical research. Despite the high lethality, methods to predict mortality for this type of aggressive malignant tumor are insufficient. Machine learning is recognized by many authors as a valuable, yet poorly studied tool in this field. Undoubtedly, searching for new feature selection methods is significant in building an effective machine-learning model. In this study, we propose the novel hybrid model using neighborhood components analysis, genetic algorithm and support vector machine classifier (NCA-GA-SVM). Because SVM works with default parameters characterized by low classification results, we decided to use GA for the proper optimization and feature selection. As reported in the available literature, NCA and GA obtain high classification results. Here, we decided to combine these approaches, building a two-level algorithm for HCC fatality prognosis. We used a well-known dataset collected from 165 patients at Coimbra's Hospital and University Center, Portugal. Our results revealed 96.36\% classification accuracy and 95.52\% F1-score. Additionally, we compared all data for these metrics published so far. We demonstrated that our algorithm achieved the highest accuracy and can be successfully applied for the assessment of hepatocellular carcinoma mortality in the future. Our findings bring methodological value for future HCC studies and emphasize the possibility of using machine-learning techniques to improve the quality of medical decisions.