All projects

Documented end to end (acquisition, features, evaluation protocols) · public-release decision pending

EMG Gesture Classification

Signal extraction from tightly constrained data

≈59%

four-class accuracy on unseen subjects (LOSO) · chance is 25%

≈29% → ≈59%

accuracy on unseen subjects after raising the sampling rate from 20 Hz to 500 Hz

15

handcrafted features across 5 families

Problem

A research project for a client: classify four hand gestures (rest, flexion, grip, pinch) from one surface-EMG channel on the forearm, with a prosthetic hand as the target. Two constraints shaped everything. One channel means no spatial information, and flexion and pinch use the same muscle group. The subject pool was small: 7 people in the first dataset, 8 in the second, and only 2 in some evaluations. The recordings were also continuous, with unlabeled rest between repetitions inside every gesture file.

Role & Approach

(1) Diagnose before modeling. The first models on 7 subjects stayed near chance, so I wrote a statistical diagnostic instead of trying a bigger model. Its control experiment settled it: a classifier scored 45.0% on the real recordings and 56.5% on Gaussian noise of the same shape. The cause was the Arduino loop, which sampled at 20 Hz. EMG carries its information between roughly 20 and 450 Hz, so everything above 10 Hz was aliased. One firmware change (delay(50) → delay(2)) raised the rate to 500 Hz, and the data was re-recorded: 110 recordings from 8 subjects. (2) Cleaning: edge trimming, a 4th-order zero-phase Butterworth band-pass (20–249 Hz), and rolling-RMS spike clipping with an IQR threshold. Rest/active RMS separation went from 1.8× to 3.7×. (3) Features: 15 handcrafted features in five families. Time domain (waveform length, zero crossings, slope sign changes), spectral (mean frequency via Welch PSD), complexity (Hjorth mobility and complexity, sample entropy), autoregressive (Burg AR coefficients) and wavelet energy (DWT db4, 3 levels), plus a WL-to-MNF ratio I designed to target the flexion/pinch confusion. (4) Activity gating: an RMS threshold that drops the unlabeled rest windows inside active recordings. (5) Model comparison: Random Forest, Optuna-tuned XGBoost, 1D CNN, CNN-LSTM and two hybrid CNN-XGBoost pipelines, to test handcrafted features against learned ones at this data size. (6) Evaluation audit: every result re-tabulated by protocol, from window-level random splits to Leave-One-Subject-Out, so each number is read against what it actually measures.

Tech Stack

PythonNumPy / SciPyscikit-learnXGBoostOptunaTensorFlow / Keras (CNN-LSTM)PyWavelets (DWT)Burg ARArduino + Olimex EMG shield

Result

For a prosthesis, the number that matters is accuracy on a person the model has never seen. Under Leave-One-Subject-Out, XGBoost on 13 handcrafted features reached 59.98% (2 held-out subjects), and a CNN-LSTM reached 58.89% across 8 subjects, against 25% chance. A later scan found identical multi-second blocks repeated across some recordings (5.5% of windows). That biases both figures up by a few points, so I report them as ≈55–60%. Three findings matter more than the headline. The sampling-rate fix moved accuracy on unseen subjects from ≈29% to ≈59%, more than every modeling change in the project combined. Handcrafted features matched or beat every deep model at this data size. And the remaining errors sit where a single channel predicts they would: 34% of flexion windows are classified as pinch.

On the 87.09% figure from the same project: it came from a random split over 85%-overlapping windows, with every subject on both sides and an activity threshold that used the class label on the test set. It measures how well the model recognizes windows close to ones it trained on, not how it handles a new user, so it is not used as a result here. The project's own report names LOSO as the deployment metric. The next step is one re-run: the 15-feature XGBoost pipeline under LOSO, with normalization fitted inside each fold, a label-free activity gate, and the duplicated blocks trimmed before windowing.

Links

Media

Bar chart: a classifier scores 45.0% on the real 20 Hz EMG recordings and 56.5% on synthetic Gaussian noise of the same shape; chance is 25%
The control that stopped the modeling: at 20 Hz the real recordings scored below matched noise. The fix was in the firmware, not the model.
Confusion matrix, XGBoost with activity threshold under Leave-One-Subject-Out: rest 4,105 of 5,106 correct, grip 2,602 of 3,728, pinch 2,083 of 4,263, flexion 1,216 of 3,585, with 1,215 flexion windows predicted as pinch
LOSO on unseen subjects, 59.98% overall. Flexion is predicted as pinch about as often as it is recognized.
To add — Raw vs. cleaned signal, per subject