Author
Listed:
- K.Vadivelan
(Research Scholar, PG and Research Department of Computer Science, Government Arts College (Autonomous), Nandanam, Chennai-35, Tamil Nadu, India)
- Dr. M.Sundara Rajan
(Associate Professor, PG and Research Department of Computer Science, Government Arts College (Autonomous), Nandanam, Chennai-35, Tamil Nadu, India)
Abstract
Social media platforms and Twitter in particular, has become a fertile channel for the rapid circulation of harmful discourse, including hateful and offensive language aimed at individuals and communities. Automatically flagging such content remains a demanding task within natural language processing, largely because micro blog text is short, informal, and heavily reliant on context for correct interpretation. This paper presents a hate-speech classification framework built around the Extreme Gradient Boosting (XGBoost) algorithm and evaluated on one of the largest publicly available labeled Twitter corpora, comprising roughly 96,973 tweets divided into three categories: hate speech, offensive language, and normal content. Each tweet is represented through a combined set of feature families: Term Frequency-Inverse Document Frequency (TF-IDF) vectors, dense word-embedding features, and tweet-level metadata such as hash tag frequency, mention count, re tweet count, and capitalization ratio. Prior to feature extraction, tweets are cleaned through tokenization, stop-word removal, lemmatization, and URL stripping to reduce noise in the raw corpus. The hyper parameters of the XGBoost classifier were selected through grid search combined with stratified cross-validation. The tuned model reached an overall accuracy of 93.7%, a macro-averaged F1-score of 0.937, and an area under the curve above 0.94 for every class, surpassing the results obtained with Naive Bayes (78.4%), Support Vector Machines (82.1%), Random Forest (85.6%), LSTM networks (88.3%), and a fine-tuned BERT model (90.1%). These outcomes indicate that gradient boosting, when paired with a carefully engineered feature set, provides an accurate and computationally efficient alternative for large-scale hate-speech detection.
Suggested Citation
Download full text from publisher
Corrections
All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:bjf:ijltem:v:15:y:2026:i:6:a:3043. See general information about how to correct material in RePEc.
If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.
We have no bibliographic references for this item. You can help adding them by using this form .
If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.
For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Dr. Pawan Verma (email available below). General contact details of provider: https://www.ijltemas.in/ .
Please note that corrections may take a couple of weeks to filter through
the various RePEc services.