A Machine-Learning Model for Identifying Non-Credible Scholarly Journals

Authors

  • Fouad Abdulameer Salman University of Baghdad, Iraq
  • Bakhtawar Baluch University of Wollongong in Dubai, United Arab Emirates
  • Zuriana Abu Bakar Universiti Malaysia Terengganu, Malaysia
  • Assane Lo University of Wollongong in Dubai, United Arab Emirates

DOI:

https://doi.org/10.64539/msts.v2i2.2026.570

Keywords:

Predatory journals, Legitimate journals, Journal credibility assessment, Machine learning, XGBoost

Abstract

The rapid growth of scholarly publishing has increased the need of reliably identify credible and non-credible journal websites. Evaluating these entities requires close, cautious, thorough, and at times skeptical scrutiny by researchers, institutions, and funding agencies. However, existing automated detection approaches appear to be customized for specific aspects of the criteria and do not effectively detect non-credible entities. Therefore, in this paper, we come up with a comprehensive dataset comprising 4,174 manually verified journal websites. We classify the dataset into two categories: credible journals and non-credible journals, including predatory, commercial, and hijacked journals. We also represent 36 new features derived from different resources of the journal websites. We used these new features to train and test the proposed model to improve accuracy. XGBoost achieved a precision of 0.9778, recall of 0.9601, F1-score of 0.9689, and accuracy of 97.86%, outperforming both GBC and Random Forest in the experimental results. Feature-group experiments show that Web-Based System and Academic Integrity features substantially improve overall classification performance, while HTML-based characteristics provide strong discriminatory power. The proposed approach shows the potential of integrating technical, structural, and scholarly-verification signals to support automated journal-credibility assessment and give researchers an efficient way to identify potentially non-credible scholarly venues.

References

[1] A. M. S. Hinze, D. Stockemer, and T. Reidy, “The predation index: A tool to discover predatory journals,” PS: Political Science & Politics, vol. 59, no. 1, pp. 158–164, 2026. https://doi.org/10.1017/S1049096525101145.

[2] F. A. do Carmo, S. O. Rezende, C. M. C. Paxiúba, and F. M. F. Lobato, “Predatory publishing in the academic landscape: Researchers’ awareness and vulnerabilities,” Scientometrics, vol. 131, no. 5, pp. 3029–3054, 2026. https://doi.org/10.1007/s11192-026-05614-0.

[3] J. Chigwada and P. Ngulube, “Use of artificial intelligence tools in the publishing process: Expectations from publishers through author guidelines,” Frontiers in Research Metrics and Analytics, vol. 11, Art. no. 1740510, 2026. https://doi.org/10.3389/frma.2026.1740510.

[4] Research and Development Department, Iraqi Ministry of Higher Education and Scientific Research, “Journal of Research,” 2026. [Online]. Available: https://jor.rdd.edu.iq.

[5] J. Beall, “Beall’s list of predatory publishers 2016,” Scholarly Open Access, 2016. [Online]. Available: https://web.archive.org/web/20170113114519/https://scholarlyoa.com/2016/01/05/bealls-list-of-predatory-publishers-2016/. [Accessed: Feb. 7, 2017].

[6] L. Giray, “Identifying predatory journals: A practical guide for researchers,” Annals of Biomedical Engineering, pp. 1–10, 2026. https://doi.org/10.1007/s10439-026-04231-5.

[7] W. M. B. Ateeq and H. S. Al-Khalifa, “Intelligent framework for detecting predatory publishing venues,” IEEE Access, vol. 11, pp. 20582–20618, 2023. https://doi.org/10.1109/ACCESS.2023.3250256.

[8] D. J. Dunleavy, “Progressive and degenerative journals: On the growth and appraisal of knowledge in scholarly publishing,” European Journal for Philosophy of Science, vol. 12, no. 4, Art. no. 61, 2022. https://doi.org/10.1007/s13194-022-00492-8.

[9] J. Wang, W. Halffman, and Y. H. Zhang, “Sorting out journals: The proliferation of journal lists in China,” Journal of the Association for Information Science and Technology, vol. 74, no. 10, pp. 1207–1228, 2023. https://doi.org/10.1002/asi.24816.

[10] K. Swargiary, “The illusory index: Indian academia and the shadow of predatory and clone journals,” ERA, U.S., 2025. https://books.google.co.id/books?id=EStiEQAAQBAJ.

[11] L. X. Chen, S. W. Su, C. H. Liao, K. S. Wong, and S. M. Yuan, “An open automation system for predatory journal detection,” Scientific Reports, vol. 13, no. 1, Art. no. 2976, 2023. https://doi.org/10.1038/s41598-023-30176-z.

[12] C. Laine and M. A. Winker, “Identifying predatory or pseudo-journals,” Biochemia Medica, vol. 27, no. 2, pp. 285–291, 2017. https://doi.org/10.11613/BM.2017.031.

[13] T. V. McCann and M. Polacsek, “False gold: Safely navigating open access publishing to avoid predatory publishers and journals,” Journal of Advanced Nursing, vol. 74, no. 4, pp. 809–817, Apr. 2018. https://doi.org/10.1111/jan.13483.

[14] S. Eriksson and G. Helgesson, “Time to stop talking about ‘predatory journals,’” Learned Publishing, vol. 31, no. 2, pp. 1–3, 2017. https://doi.org/10.1002/leap.1135.

[15] J. Olivarez, S. Bales, L. Sare, and W. vanDuinkerken, “Format aside: Applying Beall’s criteria to assess the predatory nature of both OA and non-OA library and information science journals,” College & Research Libraries, vol. 79, no. 1, p. 52, Jan. 2018. https://doi.org/10.5860/crl.79.1.52.

[16] J. A. Teixeira da Silva and P. Tsigaris, “Issues with criteria to create blacklists: An epidemiological approach,” The Journal of Academic Librarianship, vol. 46, no. 1, Art. no. 102070, 2020. https://doi.org/10.1016/j.acalib.2019.102070.

[17] M. A. Shahri, M. D. Jazi, G. Borchardt, and M. Dadkhah, “Detecting hijacked journals by using classification algorithms,” Science and Engineering Ethics, vol. 25, pp. 655–668, Apr. 2017. https://doi.org/10.1007/s11948-017-9914-2.

[18] M. Dadkhah, T. Maliszewski, and V. V. Lyashenko, “An approach for preventing the indexing of hijacked journal articles in scientific databases,” Behaviour & Information Technology, vol. 35, no. 4, pp. 298–303, Apr. 2016. https://doi.org/10.1080/0144929X.2015.1128975.

[19] S. Adnan, S. Anwar, T. Zia, S. Razzaq, F. Maqbool, and Z. U. Rehman, “Beyond Beall’s blacklist: Automatic detection of open access predatory research journals,” in Proc. IEEE 20th Int. Conf. High Performance Computing and Communications, IEEE 16th Int. Conf. Smart City, and IEEE 4th Int. Conf. Data Science and Systems (HPCC/SmartCity/DSS), Exeter, U.K., Jun. 2018, pp. 1692–1697. https://doi.org/10.1109/HPCC/SmartCity/DSS.2018.00274.

[20] M. Maktabar, A. Zainal, M. A. Maarof, and M. N. Kassim, “Content based fraudulent website detection using supervised machine learning techniques,” in Advances in Intelligent Systems and Computing, vol. 734, pp. 294–304, 2018. https://doi.org/10.1007/978-3-319-76351-4_30.

[21] K. Wu, S. Chou, S. Chen, C. Tsai, and S. Yuan, “Application of machine learning to identify counterfeit website,” pp. 321–324, 2018. https://doi.org/10.1145/3282373.3282407.

[22] Y. Li, Z. Yang, X. Chen, H. Yuan, and W. Liu, “A stacking model using URL and HTML features for phishing webpage detection,” Future Generation Computer Systems, vol. 94, pp. 27–39, 2019. https://doi.org/10.1016/j.future.2018.11.004.

[23] A. Abbasi, Z. Zhang, D. Zimbra, H. Chen, and J. F. Nunamaker, Jr., “Detecting fake websites: The contribution of statistical learning theory,” MIS Quarterly: Management Information Systems, no. 3, pp. 435–461, 2010. https://doi.org/10.2307/25750686.

[24] Y. Ding, N. Luktarhan, K. Li, and W. Slamu, “A keyword-based combination approach for detecting phishing webpages,” Computers & Security, vol. 84, pp. 256–275, 2019. https://doi.org/10.1016/j.cose.2019.03.018.

[25] N. A. Mazov and V. N. Gureev, “The editorial boards of scientific journals as a subject of scientometric research: A literature review,” Scientific and Technical Information Processing, vol. 43, no. 3, pp. 144–153, 2016. https://doi.org/10.3103/S0147688216030035.

[26] M. Nazim and A. Ali, “An investigation of open access availability of library and information science research,” DESIDOC Journal of Library & Information Technology, vol. 43, no. 2, pp. 101–111, 2023. https://doi.org/10.14429/djlit.43.02.18580.

[27] M. Sánchez-Paniagua, E. Fidalgo, E. Alegre, and F. Jáñez-Martino, “Fraud detection in e-commerce: A comparative analysis of features to enhance machine learning models,” Electronic Commerce Research, vol. 26, no. 2, pp. 2467–2502, 2026. https://doi.org/10.1007/s10660-025-10029-9.

[28] Z. Zhao, H. Wen, J. Ma, J. Zhan, T. Xu, Y. Wei, and Q. Zhang, “ResearchEVO: An end-to-end framework for automated scientific discovery and documentation,” arXiv preprint arXiv:2604.05587, 2026. https://doi.org/10.48550/arXiv.2604.05587.

[29] Anti-Phishing Working Group (APWG), “Phishing activity trends report 2Q,” 2021. [Online]. Available: https://docs.apwg.org/reports/apwg_trends_report_q2_2021.pdf. [Accessed: Oct. 14, 2024].

[30] D. Mills, A. Branford, K. Inouye, N. Robinson, and P. Kingori, “‘Fake’ journals and the fragility of authenticity: Citation indexes, ‘predatory’ publishing, and the African research ecosystem,” Journal of African Cultural Studies, vol. 33, no. 3, pp. 276–296, 2021. https://doi.org/10.1080/13696815.2020.1864304.

[31] J. M. Iv, D. Bhansali, M. Gratian, and M. Cukier, “A comprehensive evaluation of HTTP header features for detecting malicious websites,” in 15th European Dependable Computing Conference (EDCC), pp. 75–82, 2019. https://doi.org/10.1109/EDCC.2019.00025.

[32] Y. Shibuya, K. Mwitondi, and S. Zargari, “Experimental analyses in search of effective mitigation for login cross-site request forgery,” in Cyber Defence in the Age of AI, Smart Societies and Augmented Humanity, Cham, Switzerland: Springer International Publishing, 2020, pp. 233–266. https://doi.org/10.1007/978-3-030-35746-7_12.

[33] D. Singh, S. Singh, and N. Kumar, “Benchmarking credit card fraud detection models: A comprehensive evaluation across diverse datasets,” Computational Economics, pp. 1–30, 2025. https://doi.org/10.1007/s10614-025-11234-2.

[34] E. Rudini and F. Ernawan, “Prediction of Alzheimer’s dementia using soft voting ensemble learning with machine learning,” IJACI: International Journal of Advanced Computing and Informatics, vol. 1, no. 1, pp. 48–55, 2025. https://doi.org/10.71129/ijaci.v1.i1.pp48-55.

570 Cover

Downloads

Published

2026-09-12

How to Cite

Salman, F. A., Baluch, B., Bakar, Z. A., & Lo, A. (2026). A Machine-Learning Model for Identifying Non-Credible Scholarly Journals. Methods in Science and Technology Studies, 2(2), 194–207. https://doi.org/10.64539/msts.v2i2.2026.570