
Defense of the dissertation of Maxutova Natalya for the degree of Doctor of Philosophy (PhD) in the specialty «8D06103 - Information systems»

L.N. Gumilyov Eurasian National University, a dissertation defense for the degree of Doctor of Philosophy (PhD) by Maxutova Natalya on the topic «Development of an information and analytical system to support diagnostic decision-making in medicine based on the analysis of clinical data and medical examinations» to the educational program «8D06103 – Information systems».
The dissertation was carried out at the «Information Systems education department» of L.N. Gumilyov Eurasian National University.
The language of defense is russian
Official reviewers:
Anargul Shaushenova – Candidate of Technical Sciences, Associate Professor, Head of the Educational Program "Information Systems", S. Seifullin Kazakh Agrotechnical Research University (Astana, Republic of Kazakhstan).
Aslanbek Murzakhmetov – Doctor of Philosophy (PhD), acting Associate Professor, Head of the Department of Information Systems, M.Kh. Dulaty Taraz University (Taraz, Republic of Kazakhstan).
Temporary members of the Dissertation Council:
Vladimir Barakhnin – Doctor of Technical Sciences, Associate Professor, Leading Researcher at the Federal Research Center for Information and Computational Technologies (Novosibirsk, Russia);
Madina Mansurova – Candidate of Physical and Mathematical Sciences, Professor, Head of the Department of Artificial Intelligence and Big Data, Al-Farabi Kazakh National University (Almaty, Republic of Kazakhstan).
Aigerim Yerimbetova – PhD, Candidate of Technical Sciences, Associate Professor, Leading Researcher at the Institute of Information and Computational Technologies of the Science Committee of the Ministry of Science and Higher Education of the Republic of Kazakhstan (Almaty, Republic of Kazakhstan).
Scientific advisors:
Jamalbek Tussupov – Doctor of Physical and Mathematical Sciences, Professor of the Department of Information Systems, L.N. Gumilyov Eurasian National University (Astana, Republic of Kazakhstan).
Kuvvatali Rahimov – Doctor of Philosophy (PhD), Associate Professor, Head of the Department of Applied Mathematics and Informatics, Fergana State University (Fergana, Republic of Uzbekistan).
The defense will take place on August 26, 2026, at 10:00 AM in the Dissertation Council for the training direction «8D061 – Information and communication technologies» in the specialty «8D06103 – Information systems» of L.N. Gumilyov Eurasian National University. The defense meeting is planned to be held online.
Link: https://teams.microsoft.com/meet/42764093079173?p=DvPHFm2uz1HdzefgS6
Address: Astana city, Pushkin Street 11, Educational Building No. 2, Room 222.
Abstract (English): ANNOTATION of the dissertation of Maxutova Natalya Shakhinovna on the topic “Development of an information and analytical system to support diagnostic decision-making in medicine based on the analysis of clinical data and medical examinations”, submitted in fulfillment of the requirements for the degree of Doctor of Philosophy (PhD) in the educational program “8D06103 – Information Systems” Relevance of the topic. Cardiovascular diseases (CVDs) have remained the leading cause of death and disability worldwide over the past decades, which underscores the high social and economic significance of their early diagnosis and prevention. According to the World Health Organization, cardiovascular diseases claim approximately 17.9 million lives each year, accounting for around 31% of all deaths worldwide, with more than 75% of these deaths occurring in low- and middle-income countries. A substantial share of premature mortality (up to 85%) is associated with such forms of CVD as coronary heart disease and stroke. According to international studies, a significant proportion of CVD cases (up to 80%) can be prevented through the timely identification of risk factors and appropriate medical decision-making. Given the growing volume of clinical data, including laboratory indicators, demographic characteristics, and the results of medical examinations, traditional analytical methods are proving insufficiently effective for uncovering complex interrelationships and hidden patterns. The volume of medical data increases annually by 30–40%, which considerably complicates its processing and interpretation using classical statistical methods. Recent advances in artificial intelligence and machine learning open up new possibilities for medical data analysis, enabling the development of highly accurate prediction and risk-assessment models. The use of methods such as ensemble algorithms, neural networks, and gradient boosting contributes to the identification of significant risk factors and the construction of predictive models capable of accounting for nonlinear dependencies and interactions between features. These approaches take on particular importance in the detection of rare events characterized by pronounced class imbalance, which is a typical problem in medical data. At the same time, despite the considerable potential of machine learning methods, their application in clinical practice is constrained by a number of limitations, including the low interpretability of models, the need to ensure the robustness of results, and the difficulties of integrating them into existing medical information systems. This has increased the need for information-analytical decision support systems capable not only of ensuring high prediction accuracy but also of providing well-founded and interpretable recommendations for physicians. In Kazakhstan, cardiovascular disease also remains one of the most pressing challenges facing the healthcare system, exerting a significant impact on mortality rates and the population's quality of life. In recent years, the country has been implementing healthcare digitalization programs aimed at introducing information systems and data analysis technologies into medical practice. However, existing solutions in most cases fail to provide a comprehensive analysis of clinical and biochemical indicators and do not take advantage of modern machine learning methods for identifying hidden risk factors and predicting disease. The relevance of this study stems from the need to improve the effectiveness of cardiovascular disease diagnosis and prediction through the use of machine learning methods and the development of an information-analytical decision support system. Of particular importance is the identification of informative risk factors based on the analysis of clinical data, as well as the development of models that ensure high prediction accuracy and robustness under conditions of heterogeneous and imbalanced data. The development and implementation of an information-analytical system will make it possible to automate the process of medical data analysis, improve the quality of clinical decision-making, and enable a transition to a more personalized approach to the diagnosis and prevention of cardiovascular disease. This, in turn, will help reduce morbidity and mortality rates and improve the efficiency of the healthcare system. Thus, the development of methods for clinical data analysis and the creation of an information-analytical decision support system for cardiovascular disease diagnosis represents a relevant and in-demand area of research aimed at reducing mortality and increasing the effectiveness of clinical decision-making amid the digitalization of healthcare. The purpose of this dissertation is to develop an information-analytical decision support system for the diagnosis of cardiovascular disease, based on the use of machine learning methods, clinical data analysis, and the identification of informative risk factors. In accordance with the stated purpose, the dissertation addresses the following tasks: Review scientific research on the application of machine learning methods for cardiovascular disease diagnosis and risk factor assessment. Develop a machine learning model for diagnosing cardiovascular disease risk based on clinical and laboratory data. Evaluate the effectiveness of the developed model using quality metrics and an analysis of its generalization ability on test data. Develop an information-analytical decision support system for cardiovascular disease diagnosis, including databases, analytical modules, and a user interface. Accomplishing these tasks will make it possible to develop a comprehensive information-analytical solution for decision support in the field of cardiovascular disease diagnosis, contributing to greater accuracy in risk factor identification, earlier diagnosis, a reduced likelihood of disease development, and increased effectiveness of preventive and therapeutic measures within the healthcare system. Object of research – the processes of analyzing clinical data, laboratory indicators, and the results of medical examinations in the diagnosis and prediction of cardiovascular disease. Subject of research – models, algorithms, and methods for analyzing clinical data, identifying informative risk factors, and developing an information-analytical decision support system for cardiovascular disease diagnosis. Research methods – machine learning methods (XGBoost, convolutional neural networks, ensemble methods), model quality assessment metrics (MSE, R², AIC, BIC), methods for handling imbalanced data and optimizing classification thresholds, as well as model interpretation methods (SHAP) and clinical data analysis methods. as well as methods for the design and development of information systems. Scientific novelty of the dissertation Informative factors have been identified that enable early diagnosis and risk assessment of cardiovascular disease based on the analysis of heterogeneous medical data. An ensemble machine learning model has been developed for detecting rare events and assessing cardiovascular disease risk, providing improved accuracy, robustness, and interpretability of results. The methodological basis of this study rests on an integrated approach to clinical data analysis and cardiovascular disease prediction employing machine learning methods, mathematical modeling, and intelligent data analysis. The study is founded on the use of machine learning algorithms to identify informative risk factors and construct predictive models. In particular, it employs gradient boosting methods (XGBoost), convolutional neural networks, and ensemble approaches that improve diagnostic accuracy and model robustness. Model quality assessment metrics (MSE, R², AIC, BIC), methods for handling imbalanced data and optimizing classification thresholds, and model interpretation methods, including SHAP, are additionally used to identify significant factors and ensure the explainability of results. The application of these approaches makes it possible to provide well-founded decision support in medical practice and to increase the effectiveness of early cardiovascular disease diagnosis. Research methodology The research methodology is based on an integrated approach to clinical data analysis and cardiovascular disease prediction using modern machine learning methods, mathematical modeling, and artificial intelligence technologies. The work comprises several stages, encompassing the collection, preprocessing, and analysis of medical data, the development and training of machine learning models, the identification of informative risk factors, the creation of an information-analytical system, and its experimental validation. An integrated approach has been developed for the diagnosis and prediction of cardiovascular disease, based on the application of machine learning methods, including gradient boosting, neural networks, and ensemble models, as well as model interpretation methods. The proposed methodology makes it possible to improve the accuracy of risk prediction, ensure the identification of significant clinical factors, and reduce uncertainty in medical decision-making, thereby contributing to the development of intelligent decision support systems in healthcare. Statements submitted for defense include the following Informative factors for assessing and diagnosing cardiovascular disease risk, identified using machine learning methods and interpretive data analysis approaches (SHAP, among others). Anensemble model for detecting rare events and assessing cardiovascular disease risk, providing improved accuracy, robustness, and interpretability of results. A developed information-analytical decision support system providing automated analysis of medical data, individual risk assessment, and support for clinical decision-making. Theoretical significance of the dissertation lies in the development and substantiation of methods for clinical data analysis and cardiovascular disease prediction using modern machine learning methods. The dissertation proposes algorithmic approaches based on the use of gradient boosting, neural networks, and ensemble models, which improve the accuracy of risk factor identification and disease diagnosis. A comprehensive study was conducted of the informative factors influencing the development of cardiovascular disease using model interpretation methods, including SHAP analysis, which makes it possible not only to identify significant features but also to ensure the explainability of the results obtained. The developed methodology can be used in further scientific research in the field of intelligent medical data analysis, the development of decision support systems, and personalized medicine. Practical significance of the study lies in the creation of an information-analytical decision support system for the diagnosis and prediction of cardiovascular disease, which can be implemented in medical institutions to improve the effectiveness of clinical practice. The developed system provides automation of the process of analyzing clinical and laboratory data, identification of informative risk factors, prediction of disease development probability, and generation of recommendations for physicians. The use of machine learning and model interpretation methods makes it possible to improve diagnostic accuracy and ensure the validity of the decisions made. The implementation of this system in the healthcare system contributes to the early detection of cardiovascular disease, a reduced risk of complications, and improved quality of medical care for patients. Thus, the results of the dissertation research can be used both in scientific research and in the practical work of medical organizations, supporting the digitalization of diagnostic processes, greater effectiveness of prevention, and the development of personalized medicine. Software. The practical implementation of the dissertation is based on the use of modern software tools and technologies that support the development, implementation, and operation of an information-analytical decision support system for cardiovascular disease diagnosis. The system is implemented in the Python programming language, which is used for processing clinical data, building and training machine learning models, and integrating analytical modules. The server side is implemented using the Flask framework, which provides interaction between the user interface, the database, and the prediction models. The Scikit-learn, XGBoost, and TensorFlow/Keras libraries are used as tools for data analysis and model building, enabling the implementation of classification, regression, and cardiovascular disease risk prediction tasks. The SHAP method is used to interpret the results and identify informative factors. Data storage and management are carried out using modern databases that provide structured storage of clinical information, analysis results, and predictions. The client side of the system is implemented using modern web technologies that provide visualization of analysis results, display of risk factors, and convenient user interaction with the system. Specialized libraries are used for data visualization, allowing results to be presented in graphical and analytical form. Thus, the practical implementation of the dissertation is based on the integration of machine learning methods, modern software tools, and data analysis technologies, which ensures the implementation of an effective information-analytical decision support system in medicine. Implementation of results. The developed information-analytical decision support system for cardiovascular disease diagnosis was tested on clinical and laboratory data and demonstrated high effectiveness in the analysis and prediction of risk factors. The system showed high accuracy in the automated processing of medical data, the identification of informative features, and the assessment of disease development probability. The application of machine learning methods, including ensemble models and neural networks, made it possible to substantially improve the accuracy of diagnostic decisions and ensure the robustness of the models when working with heterogeneous data. The use of model interpretation methods, in particular SHAP analysis, made it possible to explain the results obtained and identify the key risk factors influencing the development of cardiovascular disease. The implemented data analysis algorithms make it possible to comprehensively assess a patient's condition, dynamically monitor indicators, and generate well-founded recommendations for medical professionals, thereby improving the quality of diagnosis and the effectiveness of preventive measures. Validation of the dissertation results. The main results of the dissertation research were presented at the following international conferences: Discussion and implementation of the research results: 2 articles published in journals indexed in Scopus and Web of Science: Maxutova N. et al. Assessing risk factors for heart disease using machine learning methods //International Journal of Electrical and Computer Engineering (IJECE). – 2024. – Vol. 14, No. 6, pp. 6734–6742. Maxutova N. et al. A Hybrid Ensemble Framework for Rare Event Detection in Large-Scale Tabular Data //Computers. – 2026. – Vol. 15, No. 3, p. 151. Structure and scope of the dissertation. The dissertation is written in Russian and consists of an introduction, three chapters, a conclusion, a list of references, and appendices. The first chapter analyzes modern and traditional approaches to the diagnosis and prediction of cardiovascular disease, with an emphasis on their limitations amid the growing volume of medical data and the need for personalized medicine. It examines classical risk-assessment methods based on clinical scales and individual biomarkers, as well as their limited ability to account for complex interrelationships between indicators. Particular attention is given to the application of digital technologies, including machine learning, artificial intelligence, and medical information systems, in improving diagnostic accuracy and the effectiveness of clinical decision-making. The second chapter is devoted to the development of a cardiovascular risk prediction model based on patients' clinical and laboratory data. It examines the stages of data preparation and preprocessing, including cleaning, normalization, and the formation of informative features. Key risk factors are analyzed, such as biochemical indicators, clinical characteristics, and behavioral parameters. Particular attention is given to the construction of a hybrid model combining ensemble methods, probabilistic approaches, and latent feature formation methods. The model's effectiveness is evaluated using machine learning metrics, including ROC-AUC, PR-AUC, and balanced accuracy, confirming its high predictive capability. The third chapter examines the practical implementation of the developed model in the form of an information-analytical decision support system. It describes the architecture of the system, implemented on a client-server basis and comprising modules for data input, preprocessing, prediction, and result interpretation. Approaches to processing medical data, generating predictive assessments, and presenting them in a clinically interpretable form are examined. The chapter presents the system testing results, confirming its effectiveness and applicability in clinical practice, and substantiates the feasibility of integrating the developed solution into existing medical information systems. Acknowledgments. The author expresses special gratitude to her scientific advisor, Professor D.A. Tusupov, Doctor of Physical and Mathematical Sciences, for valuable scientific guidance, continuous attention to the work, and support at all stages of the dissertation research. The author is also sincerely grateful to her international scientific consultant, Associate Professor K.O. Rakhimov, PhD (Uzbekistan), for professional advice, constructive comments, and assistance in developing the scientific results of the research.
