Optimization of the SVM Algorithm with Chi-Square and SMOTE for High-Dimensional Stunting Data in Samarinda

Authors

  • Taghfirul Azhima Yoga Siswa Muhammadiyah University of East Kalimantan
  • Naufal Azmi Verdikha Muhammadiyah University of East Kalimantan

DOI:

https://doi.org/10.33394/j-ps.v14i3.21848

Keywords:

Support vector machine, Chi-square, Synthetic minority over-sampling technique, High-dimensional, Stunting

Abstract

Stunting is still considered a serious problem and has become a focus of government attention in the national research priorities for 2020-2024, with a stunting prevalence of 21.6% in 2022. The city of Samarinda ranked second highest after the Kutai Kartanegara district in East Kalimantan Province in 2022, with a percentage of 25.3%. Based on previous data mining research trends, the use of classification methods such as KNN, Naïve Bayes, Random Forest, Neural Network, CART, and SVM has generally produced quite high accuracy but still focuses on low-dimensional data, which can lead to significant information loss, potential overfitting, and difficulty in interpretation. Whereas in research topics related to high-dimensional stunting data, the majority still yield low accuracy. This is further supported by the still prevalent class imbalance found in other studies, which can affect the accuracy and recall values of the built model's performance. The objective of this research is to apply the Support Vector Machine (SVM) algorithm with Chi-Square feature selection and the Synthetic Minority Over-sampling Technique (SMOTE) to address high-dimensional stunting data and handle class imbalance. The dataset for this research is sourced from the Samarinda City Health Office, consisting of 26 community health centers with 20 attributes and 102,534 records. The division of training and testing data uses the k-fold cross-validation technique with k=10. The research results show the model's performance is very good, with an accuracy of 96.6%, supported by precision, recall, and f1-score values of 97% each.

References

Arisandi, R. R. R., Warsito, B., & Hakim, A. R. (2022). Application of the Naive Bayes Classifier (NBC) to classify stunting nutritional status among children under five using k-fold cross-validation. Jurnal Gaussian, 11(1), 130-139.

Azriani, D., Agustian, D., Zuhairini, Y., Yulita, I. N., & Dhamayanti, M. (2025). Prediction models for stunting at 2-years-old from Indonesian newborn population. BMC Pediatrics, 25, Article 718. https://doi.org/10.1186/s12887-025-06096-4

Ayele, M. K., Baye, G. A., Yesuf, S. H., Engda, A. A., & Mitiku, E. T. (2025). Predicting stunting status among under five children in Ethiopia using ensemble machine learning algorithms. Scientific Reports, 15, Article 27907. https://doi.org/10.1038/s41598-025-03206-1

Bernett, J., Blumenthal, D. B., Grimm, D. G., et al. (2024). Guiding questions to avoid data leakage in biological machine learning applications. Nature Methods, 21, 1444-1453. https://doi.org/10.1038/s41592-024-02362-y

Cameron, L., Chase, C., Haque, S., Joseph, G., Pinto, R., & Wang, Q. (2021). Childhood stunting and cognitive effects of water and sanitation in Indonesia. Economics & Human Biology, 40, Article 100944. https://doi.org/10.1016/j.ehb.2020.100944

Chawla, N. V., Bowyer, K. W., Hall, L. O., & Kegelmeyer, W. P. (2002). SMOTE: Synthetic minority over-sampling technique. Journal of Artificial Intelligence Research, 16, 321-357. https://doi.org/10.1613/jair.953

Collins, G. S., Moons, K. G. M., Dhiman, P., Riley, R. D., Beam, A. L., Van Calster, B., et al. (2024). TRIPOD+AI statement: Updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ, 385, Article e078378. https://doi.org/10.1136/bmj-2023-078378

Health Development Policy Agency. (2022). Pocketbook of the 2022 Indonesian Nutritional Status Survey (SSGI) results. Ministry of Health of the Republic of Indonesia.

Hakimah, M., Prabiantissa, C. N., Rozi, N. F., Yamani, L. N., & Puspitasari, I. (2022). Determination of relevant feature combinations for detection stunting status of toddlers. In 2022 5th International Seminar on Research of Information Technology and Intelligent Systems (ISRITI) (pp. 324-329). IEEE. https://doi.org/10.1109/ISRITI56927.2022.10053069

James, G., Witten, D., Hastie, T., & Tibshirani, R. (2021). An introduction to statistical learning: With applications in R (2nd ed.). Springer. https://doi.org/10.1007/978-1-0716-1418-1

Kapoor, S., & Narayanan, A. (2023). Leakage and the reproducibility crisis in machine-learning-based science. Patterns, 4(9), Article 100804. https://doi.org/10.1016/j.patter.2023.100804

Ke, J. X. C., DhakshinaMurthy, A., George, R. B., & Branco, P. (2024). The effect of resampling techniques on the performances of machine learning clinical risk prediction models in the setting of severe class imbalance: Development and internal validation in a retrospective cohort. Discover Artificial Intelligence, 4, Article 91. https://doi.org/10.1007/s44163-024-00199-0

Khan, J. R., Tomal, J. H., & Raheem, E. (2021). Model and variable selection using machine learning methods with applications to childhood stunting in Bangladesh. Informatics for Health and Social Care, 46(4), 425-442. https://doi.org/10.1080/17538157.2021.1904938

Lonang, S., & Normawati, D. (2022). Classification of stunting status among children under five using K-Nearest Neighbor with backward-elimination feature selection. Jurnal Media Informatika Budidarma, 6(1), 49-56. https://doi.org/10.30865/mib.v6i1.3312

Moons, K. G. M., Damen, J. A. A., Kaul, T., Hooft, L., Andaur, N. C., Dhiman, P., et al. (2025). PROBAST+AI: An updated quality, risk of bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods. BMJ, 388, Article e082505. https://doi.org/10.1136/bmj-2024-082505

Mosquera, C., Ferrer, L., Milone, D. H., Luna, D., & Ferrante, E. (2024). Class imbalance on medical image classification: Towards better evaluation practices for discrimination and calibration performance. European Radiology, 34(12), 7895-7903. https://doi.org/10.1007/s00330-024-10834-0

Ndagijimana, S., Kabano, I. H., Masabo, E., & Ntaganda, J. M. (2023). Prediction of stunting among under-5 children in Rwanda using machine learning techniques. Journal of Preventive Medicine and Public Health, 56(1), 41-49. https://doi.org/10.3961/jpmph.22.388

Perdana, A. Y., Latuconsina, R., & Dinimaharawati, A. (2021). Prediction of stunting among children under five using the Random Forest algorithm. e-Proceeding of Engineering, 8(5), 6650-6656.

Pizzol, D., Tudor, F., Racalbuto, V., Bertoldo, A., Veronese, N., & Smith, L. (2021). Systematic review and meta-analysis found that malnutrition was associated with poor cognitive development. Acta Paediatrica, 110(10), 2704-2710. https://doi.org/10.1111/apa.15964

Purbasari, A., Rinawan, F. R., Zulianto, A., Susanti, A. I., & Komara, H. (2021). CRISP-DM for data quality improvement to support machine learning of stunting prediction in infants and toddlers. In 2021 8th International Conference on Advanced Informatics: Concepts, Theory and Applications (ICAICTA). IEEE. https://doi.org/10.1109/ICAICTA53211.2021.9640294

Rahman, S. M. J., Ahmed, N. A. M. F., Abedin, M. M., Ahammed, B., Ali, M., Rahman, M. J., & Maniruzzaman, M. (2021). Investigate the risk factors of stunting, wasting, and underweight among under-five Bangladeshi children and its prediction based on machine learning approach. PLOS ONE, 16(6), Article e0253172. https://doi.org/10.1371/journal.pone.0253172

Rosenblatt, M., Tejavibulya, L., Jiang, R., Noble, S., et al. (2024). Data leakage inflates prediction performance in connectome-based machine learning models. Nature Communications, 15, Article 1829. https://doi.org/10.1038/s41467-024-46150-w

Sasmita, W. G., Nugroho, W. H., & Astuti, A. B. (2022). CART method approach and high dimension simulation data selection and random under-sampling method in stunting case. Journal of Theoretical and Applied Information Technology, 100(24), 4810-4817.

Velliangiri, S., Alagumuthukrishnan, S., & Joseph, S. I. T. (2019). A review of dimensionality reduction techniques for efficient computation. Procedia Computer Science, 165, 104-111. https://doi.org/10.1016/j.procs.2020.01.079

Victora, C. G., Christian, P., Vidaletti, L. P., Gatica-Dominguez, G., Menon, P., & Black, R. E. (2021). Revisiting maternal and child undernutrition in low-income and middle-income countries: Variable progress towards an unfinished agenda. The Lancet, 397(10282), 1388-1399. https://doi.org/10.1016/S0140-6736(21)00394-9

World Health Organization. (2025). Global nutrition targets 2030: Stunting brief. https://www.who.int/publications/i/item/B09383

Downloads

Published

2026-08-10

How to Cite

Siswa, T. A. Y., & Verdikha, N. A. (2026). Optimization of the SVM Algorithm with Chi-Square and SMOTE for High-Dimensional Stunting Data in Samarinda. Prisma Sains : Jurnal Pengkajian Ilmu Dan Pembelajaran Matematika Dan IPA IKIP Mataram, 14(3), 2016–2032. https://doi.org/10.33394/j-ps.v14i3.21848

Issue

Section

Research Articles