This study proposes speech command classification using MFCC features and SVM with GridSearchCV hyperparameter optimization. Evaluating RBF/linear kernels, C (0.1-100), and gamma (0.001-scale) on Google Speech Commands Dataset (8 classes), the optimal configuration (RBF, $\mathbf{C}=\mathbf{1 0}$, gamma=0.01) as the best configuration, achieving a mean cross-validation accuracy of 83.15%. Evaluation on the independent test set yielded a final classification accuracy of 83.05%, with per-class F1-scores ranging from 0.73 to 0.89 (stop/up/yes). While lower than 3D CNN approaches (89.16%), the optimized SVM offers superior computational efficiency and provides a computationally efficient alternative compared to deep learning approaches with rigorous hyperparameter tuning as a practical baseline for lightweight speech command recognition.
Santoso, T. Sardjono, D. Purwanto· International Seminar on Int...· 0 citations
The availability and reliability of the power transmission system are crucial to ensuring the continuity of the electrical energy supply. Disturbances caused by natural events, vegetation interference, or animal activity can lead to widespread blackouts, necessitating rapid and accurate fault analysis. This study proposes a multi-representation deep learning framework for fault classification using a hybrid Convolutional Neural Network (CNN) and Long Short-Term Memory (LSTM) architecture driven by raw Disturbance Fault Recorder (DFR) data. To ensure rigorous evaluation and strictly prevent data leakage, an initial dataset of 457 raw DFR recordings is partitioned using a stratified split method. Exactly 87 validation and 54 testing samples are strictly isolated to preserve their real-world integrity, while the remaining 316 training samples are synthetically augmented to a perfectly balanced 1,500 samples (500 per class). These are processed seamlessly into 1D time-series signals and 2D stacked spatial images at a 224x224 resolution. Furthermore, an Uncertainty Rejection strategy utilizing a 60% confidence threshold is implemented to dynamically intercept ambiguous transient anomalies and prevent forced misclassifications. Experimental results demonstrate that the proposed hybrid CNN-LSTM model achieves an overall classification accuracy of 96.3%. While this 3.7% accuracy improvement corresponds to exactly a 2-sample difference on the constrained testing set compared to the established standalone baselines (LSTM and CNN, both at 92.6%), it serves as preliminary evidence demonstrating that the hybrid architecture can resolve specific, ambiguous edge-cases that single-view models fail to classify. The proposed framework offers a reliable, visually-explainable foundation that can be implemented in a monitoring system to assist operators in making faster and more accurate decisions in fault handling.
Hafizh Tri Januar, D. Purwanto, D. Kuswidiastuti· International Seminar on Int...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.