University students in Somaliland may know that cybersecurity matters, but most are not acting on that knowledge, according to one of the largest studies of digital safety ever conducted in the Horn of Africa. Researchers surveyed 1,026 students across public and private universities and then turned to machine learning to answer a question that traditional statistics alone could not: which students are most likely to fall victim to cyber threats, and what factors best predict their vulnerability. The results reveal a striking paradox, with more than 85 percent of students acknowledging the importance of cybersecurity while simultaneously reporting some of the riskiest online behaviors imaginable.
The numbers behind that paradox are sobering. Sixty-five percent of students admitted to frequently reusing passwords or relying on weak ones, a habit that security experts consider one of the most dangerous practices in the digital world, since a single breach at one service can cascade across every account that shares the same credential. Nearly half of respondents, 48 percent, reported not using two-factor authentication on any of their online accounts, leaving a single stolen password as the only barrier between an attacker and their digital identity. Thirty percent said they would potentially click on links in unsolicited emails, the classic entry point for phishing attacks, and 40 percent demonstrated limited understanding of phishing and social engineering concepts altogether.
The study, published in Discover Education by Shuaib Jama Hassan of Amoud University and Abdisalan Ahmed Osman of Kabridahar University, went beyond simply cataloguing these risky behaviors. The researchers designed a cross-sectional survey with four sections covering demographics, a 20-question cybersecurity knowledge quiz, self-reported digital habits, and perceptions of institutional support. After cleaning the data, they combined theoretical knowledge and practical behavior into a single Weighted Awareness Score, weighting each component equally, and classified students into low, moderate, and high awareness categories. The results showed 38 percent of participants in the low awareness group, 34 percent in the moderate group, and just 28 percent demonstrating genuinely high awareness.
What makes this study methodologically interesting is its use of five different machine learning algorithms to predict awareness levels: Logistic Regression, Decision Tree, Random Forest, K-Nearest Neighbors, and Naive Bayes. The researchers deliberately chose these classical approaches over deep learning, arguing that with a dataset of just over a thousand tabular survey responses, interpretability and computational efficiency matter more than the raw power of neural networks. Each model was trained on 70 percent of the data and tested on the remaining 30 percent, with performance measured using accuracy, precision, recall, F1-score, and the area under the receiver operating characteristic curve.
The Random Forest model, an ensemble method that combines hundreds of decision trees to resist overfitting, emerged as the clear winner. It achieved an F1-score of 0.83 and an AUC-ROC of 0.92, indicating a strong balance between precision and sensitivity in identifying at-risk students. Logistic Regression followed with an AUC of 0.88, and K-Nearest Neighbors reached 0.86. The confusion matrix for the winning model showed most predictions concentrated along the diagonal, meaning the algorithm was particularly effective at distinguishing students with low and moderate awareness from their better-protected peers.
Perhaps the most actionable finding came from the feature importance analysis. The Random Forest model revealed that a student’s cybersecurity knowledge score carried the greatest predictive weight at 0.24, followed closely by prior ICT training at 0.21 and password reuse frequency at 0.18. Institutional support contributed 0.15, two-factor authentication adoption 0.12, and faculty type 0.10. Notably, demographic factors like age and gender carried far less weight than behavioral and educational variables, suggesting that vulnerability to cyber threats is not a matter of who students are but of what they have been taught and how they actually behave online.
The statistical analysis reinforced these findings. Students with prior ICT training were three times more likely to fall into the high awareness category, a difference so pronounced that the chi-square test yielded a value of 45.2 with a p-value below 0.001. ICT students scored an average of 68.5 on the knowledge quiz compared to 51.3 for their non-ICT peers, and the correlation analysis revealed a significant negative relationship between perceived institutional challenges and knowledge scores, meaning students who felt their universities failed to support digital safety tended to know less about it. In the logistic regression model, prior ICT training reduced the odds of inadequate awareness by 72 percent, while non-ICT enrollment combined with poor password habits tripled the likelihood of landing in the low awareness group.
By synthesizing the model’s most powerful features, the researchers identified three distinct risk profiles that read almost like a diagnostic manual for university administrators. Non-ICT students with no prior training and poor password habits formed the highest risk group, with 82 percent predicted to have low awareness. ICT students who had received training but maintained poor practical habits fell into the moderate risk group, with 45 percent predicted as low awareness. At the other end of the spectrum, ICT students with both training and good digital practices were predicted to have high awareness 92 percent of the time. These profiles offer a roadmap for institutions with limited resources, allowing them to concentrate interventions where they will have the greatest impact.
The researchers interpret the knowledge-behavior gap through the lens of Protection Motivation Theory, which holds that people act on threats only when they perceive them as personally relevant and feel capable of responding effectively. Somaliland’s students may intellectually recognize that cybercrime exists while underestimating their own likelihood of being targeted, or they may feel ill-equipped to implement security measures like two-factor authentication. This motivational gap, the authors argue, explains why awareness campaigns that simply transfer knowledge so often fail to change behavior. The theory suggests interventions must build both perceived threat relevance and practical self-efficacy, particularly among non-ICT students who showed the highest vulnerability in the study.
The study’s recommendations span two tiers of urgency. In the short term, universities can launch phishing simulation workshops with immediate feedback, hands-on password hygiene sessions, simplified two-factor authentication tutorials, peer mentoring programs in which ICT students coach their non-ICT classmates, and mobile-friendly educational resources suited to how students actually consume information. Longer-term structural changes include mandatory cybersecurity modules integrated across all faculties, institutional password policies, required two-factor authentication for sensitive university systems, and dedicated help desks where students can ask security questions without judgment. The authors also call for national coordination, urging Somaliland’s Ministry of Education and Science to establish minimum cybersecurity standards for tertiary institutions and a national digital literacy framework. The researchers acknowledge their limitations, including the cross-sectional design that prevents causal conclusions and the reliance on self-reported behavior, which is subject to social desirability bias. Still, as Somaliland’s universities rapidly digitize, this first-of-its-kind study provides both a warning and a practical playbook: cybersecurity resilience depends not just on what students know, but on whether their institutions make the safe choice the easy choice.
Subject of Research: Predicting cybersecurity awareness and risky online behavior among university students in Somaliland using machine learning models
Article Title: Predicting cybersecurity knowledge and behavior among Somaliland university students using machine learning
Article References: Hassan, S. J., & Osman, A. A. (2026). Predicting cybersecurity knowledge and behavior among Somaliland university students using machine learning. Discover Education, 5(1), Article 1142. https://doi.org/10.1007/s44217-026-02261-8
Image Credits: AI Generated
DOI: 10.1007/s44217-026-02261-8
Keywords: cybersecurity awareness, machine learning, Somaliland, higher education, phishing, password reuse, two-factor authentication, Random Forest, digital behavior, ICT training, risk profiles, Protection Motivation Theory
Cite Scienmag News
Teresa Odom. (October 8, 2026). Machine Learning Reveals Which Students Are Most at Risk of Cyber Attacks in Somaliland. Scienmag. https://scienmag.com/machine-learning-reveals-which-students-are-most-at-risk-of-cyber-attacks-in-somaliland/
Teresa Odom. "Machine Learning Reveals Which Students Are Most at Risk of Cyber Attacks in Somaliland." Scienmag, 8 October 2026, https://scienmag.com/machine-learning-reveals-which-students-are-most-at-risk-of-cyber-attacks-in-somaliland/. Accessed 8 October 2026.
Teresa Odom. "Machine Learning Reveals Which Students Are Most at Risk of Cyber Attacks in Somaliland." Scienmag. October 8, 2026. https://scienmag.com/machine-learning-reveals-which-students-are-most-at-risk-of-cyber-attacks-in-somaliland/

