Home > Published Issues > 2026 > Volume 21, No. 4, 2026 >
JCM 2026 Vol.21(4): 528-536
Doi: 10.12720/jcm.21.4.528-536

Dual-Input Multi-Feature Convolutional Neural Network-Long Short-Term Memory for Voice Gender Recognition in Intelligent Communication Systems

Sanjit Kumar Dash 1, Ayush Nanda1, Rashmita Barik1, Suleman Alnatheer2,*, Mohammed Altaf Ahmed2, and Abdullah Alsir Mohamed2
1Department of Information Technology, Odisha University of Technology and Research, Bhubaneswar, Odisha, India
2Department of Computer Engineering, College of Computer Engineering and Sciences, Prince Sattam bin Abdulaziz University, Al-Kharj, Saudi Arabia
Email: skdash@outr.ac.in (S.K.D.); ayushnanda.3110@gmail.com (A.N.); barikrashmita111@gmail.com (R.B.); s.alnatheer@psau.edu.sa (S.A.);m.altaf@psau.edu.sa(M.A.A.);a.mhamed@psau.edu.sa (A.A.M.)
*Corresponding author

Manuscript received February 8, 2026; revised March 24, 2026; accepted April 4, 2026; published August 24, 2026.

Abstract—Accurate voice-based gender detection is critical for secure biometric authentication and personalized human-computer interaction, yet conventional single-feature systems often suffer from misidentification. To address these limitations, this research introduces a robust dual-input deep learning framework that integrates a Convolutional Neural Network (CNN) with a Long Short-Term Memory (LSTM) network. The proposed architecture independently processes parallel standard Mel-Frequency Cepstral Coefficients (MFCCs) and high-resolution Mel-Spectrograms to effectively capture both the spatiotemporal characteristics of human speech. Rigorous evaluation was conducted on the Texas Instruments/Massachusetts Institute of Technology (TIMIT) speech corpus using strict speaker-disjoint protocols to prevent data leakage, alongside mathematical class-weighting to resolve dataset imbalances. This strategic weighting preserved the physical integrity of the original acoustic matrices by eliminating the need for synthetic data augmentation. Ultimately, the multi-feature Convolutional Neural Network-Long Short-Term Memory Network (CNN-LSTM) model achieved an exceptional classification accuracy of 98.87% and a Receiver Operating Characteristic-Area Under Curve (ROC-AUC) score of 0.998, scientifically demonstrating its superior capacity to handle class imbalances and accurately identify the complex acoustic patterns differentiating male and female voices.
 
Keywords—gender detection, speech signals, deep learning, Convolutional Neural Networks (CNN), Long Short-Term Memory (LSTM), Hybrid Model, Mel-Frequency Cepstral Coefficients (MFCCs), Human-Computer Interaction (HCI)


Cite: Sanjit Kumar Dash, Ayush Nanda, Rashmita Barik, Suleman Alnatheer,  Mohammed Altaf Ahmed, and Abdullah Alsir Mohamed, “Dual-Input Multi-Feature Convolutional Neural Network-Long Short-Term Memory for Voice Gender Recognition in Intelligent Communication Systems," Journal of Communications, vol. 21, no. 4, pp. 528-536, 2026.

Copyright © 2026 by the authors. This is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited (CC BY 4.0).

Article Metrics in Dimensions