Skip to content

A Deterministic Binary Fingerprinting Framework with Zero-Trained Feature Extraction for Sparse Count Matrices

Jul 2026 · arXiv.org · Vol abs/2607.14596 · 0 citations
Computer Science

TL;DR

MMTB is introduced, a deterministic non-learned binary representation framework that requires no label supervision, model fitting, or gradient-based optimization and is suited for coarse-grained separation and resource-constrained deployments rather than fine-grained subtype discovery.

Abstract

Sparse count matrices from single-cell transcriptomes to k-mer profiles and document-term frequencies are conventionally analyzed via PCA-reduced graph clustering or iterative optimization in continuous embedding spaces. We introduce MMTB, a deterministic non-learned binary representation framework that requires no label supervision, model fitting, or gradient-based optimization. Column-wise Min-Max normalization followed by fixed cutoffs maps each sample to a thermometer fingerprint whose Hamming distances show empirical correspondence with normalized L1 distances, with Pearson correlation approximately 0.92 on single-cell RNA-seq pairs. In a favorable three-cell-line mixture, the 3-threshold fingerprint achieves NMI of 0.99 at 188 bytes per cell. Under a fair Hamming nearest-neighbor graph plus Leiden readout, MMTB approaches PCA plus Leiden on this coarse task. On challenging tissue-like annotations, continuous pipelines often lead; PBMC Seurat NMI is 0.39 for MMTB versus 0.49 for Scanpy, underscoring that MMTB is suited for coarse-grained separation and resource-constrained deployments rather than fine-grained subtype discovery or as a general replacement for continuous embeddings. Relative to dense float32 representations, MMTB fingerprints reduce memory by approximately 10-fold while providing fixed-width Hamming-indexable codes. PCA30 embeddings and sparse CSR may be smaller; we do not claim universal compression. A label-free suitability score is provided as a deployment guideline, not a performance predictor.

View source

Similar papers

Open access Aug 2026

Measuring and removing near-duplicate contamination in alignment-free SARS-CoV-2 lineage classification benchmarks

Alignment-free lineage assignment from k-mer frequency profiles is widely used for SARS-CoV-2 surveillance, and the methods that do it are ranked against each other by margins of one or two percentage points. Those rankings rest on an unchecked protocol. Public repositories hold many near-duplicate genomes, and stratif...

Mohammad Jamhuri, A. Irawan · 0 citations
Preprint Aug 2026

Who Built This Model? Tracing LLM Lineage via Spectral Fingerprints in Weight Space

This work investigates the notion of LLM"biometrics" to ask whether LLMs exhibit intrinsic fingerprints in weight space alone, without access to input data, that reveal their origin and lineage, and proposes a unified geometric fingerprinting framework that analyzes weight matrices from two complementary perspectives.

Yiwei Chen, Bing-Qi Shang, Sijia Liu · 0 citations
Preprint Aug 2026

Neural Fingerprints for Malware Analysis: An Image-Based Metric Learning Approach with Application to Cross-Domain Classification

Identifying the family of a newly observed malware sample is a core task in threat intelligence, yet conventional classifiers must be retrained whenever a new family appears. This chapter develops an image-based metric learning approach that instead learns to extract discriminative neural fingerprints--fixed-length emb...

Manasa Deshagouni, Sayma Akther, Martin Jurecek et al. · 0 citations
2026

An Unsupervised Multi-View Graph Pseudo-Labeling Framework for Automatic Modulation Recognition

Automatic modulation recognition (AMR) in non-cooperative systems is limited by scarce labels. This letter proposes a strictly unsupervised, fully label-free multi-view graph pseudo-labeling framework. Ground-truth labels are never loaded by any training-stage routine: representation learning, graph construction, pseud...

Yu Xi, Shi-Lian Zheng, Chao Wang et al. · 0 citations
#machine learning Preprint Aug 2026

RSLM: Training-Free Vector Quantization for Approximate Nearest Neighbor Search

By introducing RSLM (Rotated Scaled Lloyd-Max), a family of training-free vector quantization codecs compressing embeddings to 1--4 bits per dimension, this work reduces memory cost and memory bandwidth of a typical large-scale Approximate Nearest Neighbor (ANN) search system, while reducing its complexity and keeping...

Rastislav Lenhardt, Teodora Dobos, Thomas Vecchiato et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Multi-View Molecular Representation Learning with Hierarchical Graphs and Contextualized Fingerprints

Analysis of HiFi-Mol reveals that fragment-aware masking improves graph representation quality, and classification results demonstrate dataset-dependent strengths of the individual graph and fingerprint variants, confirming that the two views provide complementary predictive signals.

Gwang-Hyeon Yun, Jong-Hoon Park, Bing Hu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.