This study proposes speech command classification using MFCC features and SVM with GridSearchCV hyperparameter optimization. Evaluating RBF/linear kernels, C (0.1-100), and gamma (0.001-scale) on Google Speech Commands Dataset (8 classes), the optimal configuration (RBF, $\mathbf{C}=\mathbf{1 0}$, gamma=0.01) as the best configuration, achieving a mean cross-validation accuracy of 83.15%. Evaluation on the independent test set yielded a final classification accuracy of 83.05%, with per-class F1-scores ranging from 0.73 to 0.89 (stop/up/yes). While lower than 3D CNN approaches (89.16%), the optimized SVM offers superior computational efficiency and provides a computationally efficient alternative compared to deep learning approaches with rigorous hyperparameter tuning as a practical baseline for lightweight speech command recognition.
Santoso, T. Sardjono, D. Purwanto· International Seminar on Int...· 0 citations
A novel video-based stress detection system with speaking awareness that dynamically processes facial features according to online detection of speech activities to demonstrate the system's applicability to real-world applications in stress tracking across healthcare, educational, and workplace well-being contexts.
Ahmad Rafiqan, Rachmad Setiawan, T. Sardjono· JAREE (Journal on Advanced R...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.