Automated bullying-content detection has advanced rapidly for English and other high-resource languages, yet comparable evidence for Swahili remains limited, particularly for short message service (SMS) communication in Tanzania. This study developed and evaluated a context-aware deep learning model for binary classification of bullying and non-bullying Swahili SMS messages. A multi-source corpus of 7,228 messages was initially assembled from prior Swahili datasets, voluntary student contributions, and Google Forms; duplicate records were removed during data cleaning before the train-validation partition was created. Four Kiswahili graduates applied a common annotation framework that considered the target, communicative intent, and surrounding linguistic context rather than treating offensive vocabulary as a sufficient label criterion. The messages were normalised, tokenised with a 10,000-token vocabulary, padded to 100 tokens, and represented using trainable 300-dimensional FastText embeddings. A Bidirectional Long Short-Term Memory network used 160 units in each direction, followed by dropout, a 96-unit rectified linear dense layer with L2 regularisation, and a two-class Softmax output. Candidate configurations were assessed through Keras Tuner and validation-based model selection. On the 1,470-message internal validation set created after duplicate removal, the selected model achieved 90.07% accuracy, 90.09% macro precision, 90.05% macro recall, and 90.06% macro F1-score. The confusion matrix contained 641 true negatives, 683 true positives, 80 false positives, and 66 false negatives. Bullying recall reached 91.19%, indicating that the model identified most harmful messages, although the 8.01-percentage-point training-validation gap and divergent loss curves showed moderate overfitting. The study contributes a Tanzania-focused Swahili SMS resource, an empirically evaluated FastText-BiLSTM architecture, and deployment guidance that positions automated detection as a triage mechanism for human review rather than an autonomous enforcement tool.
Andrea Peter, Gustaph Sanga, G. Tesha· East African Journal of Info...· 0 citations
Savings and Credit Cooperative Societies (SACCOS) are central to financial inclusion in Tanzania; however, their digital transformation has advanced faster than their cybersecurity capabilities. National instruments, including the Cybercrimes Act (CAP 443), the Government Cyber Security Strategy 2022–2027, and the TCDC Guidelines on Cybersecurity and Resilience of SACCOS, establish baseline obligations, but none provide a structured, risk-based mechanism for measuring cybersecurity maturity or tracking improvements over time. This study developed and evaluated a context-specific Risk-Based Cybersecurity Maturity Assessment Framework (RBCMAF) for Tanzanian SACCOS by adapting the NIST Cybersecurity Framework (CSF) 2.0 to local governance, regulatory, and resource conditions. A descriptive, analytical, cross-sectional, mixed-methods design guided by Design Science Research principles was used. The quantitative strand is explicitly positioned as an exploratory pilot baseline rather than a nationally representative survey, drawing on respondents from a small number of purposively selected, anonymised digitised SACCOS using a NIST CSF-aligned questionnaire scored across the six CSF functions. Qualitative data were generated through semi-structured interviews with ICT managers, one per SACCOS, and a structured review of regulatory and supervisory documents. The instrument showed very high internal consistency, which should be read with caution because such values may also indicate item redundancy. The baseline placed the sampled SACCOS at the Developing maturity level overall, with Identify and Protect emerging as the strongest functions, and Respond, Recover, and Detect as the weakest. Gap analysis against an optimised target level confirmed that the largest deficits lay in Respond, Recover, and Detect. A risk-weighted assessment similarly prioritised Respond, Recover, Detect, and Govern as the functions most in need of attention. The resulting RBCMAF comprises five integrated layers operationalised through a six-stage assessment process and six design principles. Evaluation through quantitative application, qualitative triangulation, and regulatory benchmarking demonstrates the framework’s internal coherence, contextual fit, and practical utility for institutional self-assessment and risk-based supervision.
Ayoub Jonathan Kitomari, Gustaph Sanga, S. Wambura· East African Journal of Info...· 0 citations
Phishing conducted in Swahili has become a persistent threat to the millions of Tanzanians who depend on mobile-money services, yet the detection tools in common use are built for English and transfer poorly to a language whose morphology, register, and transactional vocabulary differ sharply from it. This study makes three contributions. It establishes that classical machine learning, given features tuned to Tanzanian Swahili, separates phishing from legitimate messages at near-ceiling accuracy, and that a deep-learning comparator adds nothing of operational consequence. It documents the compact lexical signature on which that separation rests, built from direct imperatives, money terms, and mobile-operator names. And it shows that lexical urgency, treated as a hallmark of phishing throughout the English-language literature, carries almost no discriminating signal in this language, a caution against porting feature assumptions across languages unexamined. The evidence comes from a corpus of 2,408 Tanzanian short messages, 1,377 of them real SMS drawn from the BongoSCAM collection, on which three classical classifiers and a convolutional neural network were compared under five-fold stratified cross-validation and four ablation experiments. The linear support vector machine and the random forest each returned a mean F1-score of 0.9983 (± 0.0016) and the convolutional network 0.9989 (± 0.0014), a difference smaller than one standard deviation. Performance held across every ablation, indicating that the signal is linguistic rather than an artefact of data construction. Classical models therefore offer an accurate and computationally frugal basis for protecting Swahili-speaking users, provided the gap between balanced-corpus evaluation and the low phishing prevalence of live traffic is managed by pairing the classifier with human review.
Rehema Abdallah Njame, Gustaph Sanga, I. Tende· East African Journal of Info...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.