Intelligent Analysis of Users Data for Investment in Digital Financial Platforms Using Machine Learning Algorithms
Abstract
Today, due to the high risk of investing in emerging markets, the application of Machine Learning (ML) to design systems tailored to specific financial environments has rapidly expanded, providing new opportunities for analyzing financial data and improving decision-making. Peer-to-Peer (P2P) lending, a form of crowdfunding, has attracted considerable attention from corporate investors, institutional investors, and credit rating agencies by providing an alternative source of financing for individuals, businesses, and entrepreneurs. In this study, to address the class imbalance problem in P2P lending credit scoring, undersampling, oversampling, and synthetic minority oversampling techniques were employed in conjunction with six widely used ML classifiers including Multilayer Perceptron (MLP), Random Forest (RF), Decision Tree (DT), Support Vector Machine (SVM), Logistic Regression (LR), and Light Gradient Boosting Machine (LGBM) based on a large dataset from the Lending Club website. The performance of the models was evaluated using four metrics: precision, accuracy, recall, and the F1 score. The experimental results show that artificial minority oversampling with MLP method has the best performance due to its higher value in the three criteria compared to other methods, as well as previous studies, and achieved an Area Under the Curve (AUC) value of 0.925 in the Receiver Operating Characteristic (ROC) curve, indicating high class discrimination. The findings of this study provide a comprehensive, data-driven framework that can guide credit risk assessment and support decision-making in real-world financial risk management.
Keywords:
Peer-to-peer lending, Online platform, Financial risk, Machine learning, Classification, Imbalanced dataReferences
- [1] Vallee, B., & Zeng, Y. (2019). Marketplace lending: A new banking paradigm? The review of financial studies, 32(5), 1939–1982. https://doi.org/10.1093/rfs/hhy100
- [2] Huang, C. H., & Nguyen, V. T. (2026). Revisiting the shifting landscape of P2P lending: A systematic review based on the affordance actualization perspective. Electronic commerce research, 26(3), 2973–2999. https://doi.org/10.1007/s10660-026-10110-x
- [3] Cummins, M., Lynn, T., an Bhaird, C., & Rosati, P. (2018). Addressing information asymmetries in online peer-to-peer lending. In Disrupting finance: Fintech and strategy in the 21st century (pp. 15–31). Springer. https://doi.org/10.1007/978-3-030-02330-0_2
- [4] Silva, R. R., Garcia, L. A. A., Buso, A. L. Z., de Paula, F. F. S., Marques, D. S., & da Silva Santos, Á. (2022). Novas listas e novas ferramentas tecnológicas sobre medicamentos potencialmente inapropriados para idosos: Uma revisão integrativa. Revista família, ciclos de vida e saúde no contexto social, 10(2), 340-369. https://doi.org/10.18554/refacs.v10i2.5970
- [5] Wu, M. E., Syu, J. H., Lin, J. C. W., & Ho, J. M. (2021). Portfolio management system in equity market neutral using reinforcement learning. Applied Intelligence, 51(11), 8119-8131. https://doi.org/10.1007/s10489-021-02262-0
- [6] Wang, Y. F., Chen, Y. C., & Chien, S. Y. (2025). Citizens’ intention to follow recommendations from a government-supported AI-enabled system. Public policy and administration, 40(2), 372-400. https://doi.org/10.1177/09520767231176126
- [7] Hafez, I. Y., Hafez, A. Y., Saleh, A., Abd El-Mageed, A. A., & Abohany, A. A. (2025). A systematic review of AI-enhanced techniques in credit card fraud detection. Journal of big data, 12(1), 6. https://doi.org/10.1186/s40537-024-01048-8
- [8] Kim, W., Kanezaki, A., & Tanaka, M. (2020). Unsupervised learning of image segmentation based on differentiable feature clustering. IEEE transactions on image processing, 29, 8055–8068. https://doi.org/10.1109/TIP.2020.3011269
- [9] Zakhidov, G. (2024). Economic indicators: Tools for analyzing market trends and predicting future performance. International multidisciplinary journal of universal scientific prospectives, 2(3), 23–29. https://www.researchgate.net/publication/380519295
- [10] Ban, G. Y., El Karoui, N., & Lim, A. E. B. (2018). Machine learning and portfolio optimization. Management science, 64(3), 1136–1154. https://doi.org/10.1287/mnsc.2016.2644
- [11] Chen, L., Pelger, M., & Zhu, J. (2024). Deep learning in asset pricing. Management Science, 70(2), 714–750. https://doi.org/10.1287/mnsc.2023.4695
- [12] Gu, S., Kelly, B., & Xiu, D. (2020). Empirical asset pricing via machine learning. The review of financial studies, 33(5), 2223–2273. https://doi.org/10.1093/rfs/hhaa009
- [13] Gao, B., & Balyan, V. (2022). Construction of a financial default risk prediction model based on the LightGBM algorithm. Journal of intelligent systems, 31(1), 767–779. https://doi.org/10.1515/jisys-2022-0036
- [14] Hu, X., Chen, Y., Ren, L., & Xu, Z. (2023). Investor preference analysis: An online optimization approach with missing information. Information sciences, 633, 27–40. https://doi.org/10.1016/j.ins.2023.03.066
- [15] Shuryhin, K. A., & Zinovatna, S. L. (2024). Recommendation system for financial decision-making using artificial intelligence. Applied aspects of information technology, 7(4), 348–358. https://doi.org/10.15276/aait.07.2024.24
- [16] Sharma, A. K., & Chiu, K. C. (2026). Data-driven prediction of borrower default in P2P lending using feature-optimized ML models. OPSEARCH, 1–24. https://doi.org/10.1007/s12597-026-01166-2
- [17] Najem, R., Amr, M. F., Bahnasse, A., & Talea, M. (2022). Artificial intelligence for digital finance, axes and techniques. Procedia computer science, 203, 633–638. https://doi.org/10.1016/j.procs.2022.07.092
- [18] Kumar, R., & Verma, R. (2012). Classification algorithms for data mining: A survey. International journal of innovations in engineering and technology (IJIET), 1(2), 7–14. https://d1wqtxts1xzle7.cloudfront.net/32141572/data-libre.pdf
- [19] Nikam, S. S. (2015). A comparative study of classification techniques in data mining algorithms. Oriental journal of computer science and technology, 8(1), 13–19. https://www.computerscijournal.org/pdf/vol8no1/vol_8_no_01_13-19.pdf
- [20] Haleem, A., Javaid, M., Qadri, M. A., Singh, R. P., & Suman, R. (2022). Artificial intelligence (AI) applications for marketing: A literature-based study. International journal of intelligent networks, 3, 119–132. https://doi.org/10.1016/j.ijin.2022.08.005
- [21] Odeh, A. J., Keshta, I., & Abdelfattah, E. (2020). Efficient detection of phishing websites using multilayer perceptron. International journal of interactive mobile technologies (IJIM), 14(11), 22–31. https://doi.org/10.3991/ijim.v14i11.13903
- [22] Madane, N., & Nanda, S. (2019). Loan prediction analysis using decision tree. Journal of the gujarat research society, 21(14), 214–221. https://doi.org/10.1088/1757-899X/1022/1/012042
- [23] Addo, P. M., Guegan, D., & Hassani, B. (2018). Credit risk analysis using machine and deep learning models. Risks, 6(2), 38. https://doi.org/10.3390/risks6020038
- [24] Flach, P. (2012). Machine learning: The art and science of algorithms that make sense of data. Cambridge University Press. https://www.amazon.com/Machine-Learning-Science-Algorithms-Sense/dp/1107422221
- [25] Noble, W. S. (2006). What is a support vector machine? Nature biotechnology, 24(12), 1565–1567. https://doi.org/10.1038/nbt1206-1565
- [26] Serrano-Cinca, C., & Gutiérrez-Nieto, B. (2016). The use of profit scoring as an alternative to credit scoring systems in peer-to-peer (P2P) lending. Decision support systems, 89, 113-122. https://doi.org/10.1016/j.dss.2016.06.014
- [27] Peng, C. Y. J., Lee, K. L., & Ingersoll, G. M. (2002). An introduction to logistic regression analysis and reporting. The journal of educational research, 96(1), 3-14. https://doi.org/10.1080/00220670209598786
- [28] Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., ... & Liu, T. Y. (2017). Lightgbm: A highly efficient gradient boosting decision tree. Advances in neural information processing systems, 30. https://proceedings.neurips.cc/paper_files/paper/2017/file/6449f44a102fde848669bdd9eb6b76fa-Paper.pdf
- [29] Chawla, N. V., Bowyer, K. W., Hall, L. O., & Kegelmeyer, W. P. (2002). SMOTE: Synthetic minority over-sampling technique. Journal of artificial intelligence research, 16, 321–357. https://doi.org/10.1613/jair.953