Vāksetu ISL Feature Dataset (v1.0) : 300-Class Indian Sign Language Sequences with Adapter and Phase-2 Refinement Splits
Abstract
Fixed-length feature sequences for Indian Sign Language (ISL) recognition, collected for the final-year project "Vāksetu: Edge-Optimized ISL Recognition for Assistive Communication" (Goa College of Engineering). The dataset consolidates all training data sources of the project into one HDF5 file with per-sample provenance. Each sample is a float32 array of shape (20, 506): 20 frames with 506 features per frame, derived from MediaPipe landmarks. Per frame there are 253 base features (left and right hand landmarks, wrist-centred and scale-normalised; the same landmarks expressed relative to the face; a hand-to-face proximity scalar) plus 253 frame-to-frame velocity features. No video or images are included. GROUPS IN dataset_consolidated.h5 (118,114 samples in total; each group has a "data" dataset of shape (N, 20, 506) and a "filenames" dataset; gzip level 4, one sample per chunk) - main_positive (89,883 samples, 300 classes): primary training set. Also stores per-sample "labels" (indices into "class_names"), "weights" (all 1.0, placeholder) and "domains". Domains, by filename prefix: webcam 87,684 (recorded live by the author and, in a small share, a relative); MVI 920 (landmark features extracted by the author from videos of the INCLUDE dataset, Sridhar et al., 2020, CC BY 4.0); cvae 325 (synthetic sequences from a conditional VAE); unknown 954. Webcam-prefixed samples include augmented variants. Filenames are synthetic (hdf5_sample_NNNNNN) because the original filenames were not stored when the set was compiled. Only 49,081 of the 89,883 arrays are unique. - refining_adapter (6,690 samples, 64 classes): sequences collected by the author during deployment and used to train a small MLP output adapter on top of the frozen base ensemble. By filename pattern, 892 (13.3%) have user-corrected labels and 5,798 (86.7%) have model-predicted labels (confidence >= 0.85). The two kinds are not flagged inside the file, so label noise is possible in the model-predicted samples. - refining_phase2 (16,134 samples, 300 classes): positive samples removed from the primary pool by a quality and diversity filter and used in Phase-2 fine-tuning. 16,125 are webcam-prefixed; at least 54.2% carry an augmentation marker (_aug_ or _merge_) in the filename. 11,281 arrays are unique. - negative (3,915 samples, 23 categories): reject class covering transitions, incomplete signs, idle, tracking failures, empty scenes, other people, non-signing hand movement and confusable sign categories. Used in Phase-1 training. - refining_negative (1,492 samples, 23 categories): lower-quality reject samples removed from the same negative pool by the same filter, used in Phase-2 fine-tuning. FILES - dataset_consolidated.h5 (2.22 GB): the dataset.- manifest_consolidated.csv: one row per sample (filename, group, class, source directory, domain, augmentation flag, SHA-256 of the array).- dataset_consolidated.h5.sha256: SHA-256 checksum of the HDF5 file.- quality_flags.csv: rows flagged as tracking failures (no hand landmarks in any of the 20 frames) or as static sequences (all 20 base-feature frames identical). Columns: group, row_index (index into that group's "data"), class, SHA-256 of the array, and the two flags. It does not cover sequences with hands missing in only some frames. The root attributes of the HDF5 file ("notes", "overlap_stats", "domain_counts_*", per-group counts and class counts) hold further provenance details. USAGE import h5pywith h5py.File("dataset_consolidated.h5", "r") as f: x = f["refining_phase2/data"][0] # (20, 506) float32 print(f.attrs["notes"]) LICENSE AND ATTRIBUTION Released under CC BY 4.0. Samples with the MVI prefix (1,076 in total: 920 in main_positive, 155 in negative, 1 in refining_negative) are landmark features extracted from videos of the INCLUDE dataset (Sridhar, Ganesan, Kumar & Khapra, ACM Multimedia 2020, https://doi.org/10.1145/3394171.3413528, CC BY 4.0), signed by other people. They are modified from the source (feature extraction and augmentation). All other data were recorded by the author and, in a small share, a relative, both of whom agreed to its release. LIMITATIONS - Groups are not mutually exclusive: 1,176 refining_phase2 arrays also occur in main_positive and 112 refining_negative arrays in negative, with smaller overlaps between other pairs. refining_adapter is disjoint from all other groups.- Exact duplicate arrays are common: 40,802 extra copies in main_positive and 4,853 in refining_phase2, with one array occurring 327 times. The cause of most of this duplication has not been determined. Duplicates within main_positive share a single label. Deduplicate by hash before making train/test splits.- A small number of sequences are tracking failures with no hand detected in any frame yet carry a sign label: 337 rows (11 unique arrays) in main_positive and 33 rows (6 unique arrays) in refining_phase2, 327 of them a single array labelled "z". They are listed in quality_flags.csv. Empty arrays are expected in the negative categories empty_scene and tracking_failure.- Eight identical arrays carry different labels across groups.- Non-INCLUDE data come almost entirely from a single signer, with a small share from a second person, so generalisation to other signers is very limited.- Signer metadata is not included.