2026-08-10
2026-06-29
2026-04-24
Manuscript received February 8, 2026; revised March 24, 2026; accepted April 4, 2026; published August 24, 2026.
Abstract—Accurate voice-based gender detection is critical for secure biometric authentication and personalized human-computer interaction, yet conventional single-feature systems often suffer from misidentification. To address these limitations, this research introduces a robust dual-input deep learning framework that integrates a Convolutional Neural Network (CNN) with a Long Short-Term Memory (LSTM) network. The proposed architecture independently processes parallel standard Mel-Frequency Cepstral Coefficients (MFCCs) and high-resolution Mel-Spectrograms to effectively capture both the spatiotemporal characteristics of human speech. Rigorous evaluation was conducted on the Texas Instruments/Massachusetts Institute of Technology (TIMIT) speech corpus using strict speaker-disjoint protocols to prevent data leakage, alongside mathematical class-weighting to resolve dataset imbalances. This strategic weighting preserved the physical integrity of the original acoustic matrices by eliminating the need for synthetic data augmentation. Ultimately, the multi-feature Convolutional Neural Network-Long Short-Term Memory Network (CNN-LSTM) model achieved an exceptional classification accuracy of 98.87% and a Receiver Operating Characteristic-Area Under Curve (ROC-AUC) score of 0.998, scientifically demonstrating its superior capacity to handle class imbalances and accurately identify the complex acoustic patterns differentiating male and female voices. Keywords—gender detection, speech signals, deep learning, Convolutional Neural Networks (CNN), Long Short-Term Memory (LSTM), Hybrid Model, Mel-Frequency Cepstral Coefficients (MFCCs), Human-Computer Interaction (HCI)