A conserved molecular signature of electroacupuncture across organs: integrative reanalysis of public transcriptomic datasets
Abstract
========== 1. RESEARCH QUESTION AND HYPOTHESES ========== Novelty context (verified 2026-09-28 by adversarial literature search): no prior work integrates multiple public acupuncture/electroacupuncture transcriptomic datasets at the raw-data level across organs. Nearest precedents to be cited and differentiated in the manuscript: single-disease multi-dataset bioinformatics (PMID 40636547), target-level cross-disease mixed analyses (PMID 39492981, 36776063), and public-data WGCNA on two CNS regions (PMID 32953563). The claim is accordingly worded: the first raw-data-level, batch-harmonized, multi-organ integration. Question: Does electroacupuncture (EA) / manual acupuncture elicit a transcriptional response program that is conserved across organs, disease models, and species? Which cell types carry this program, along which dynamic trajectories does it unfold, and which druggable targets does it enrich? Hypotheses: (H1) A conserved cross-organ module exists (dominated by neuroimmune axes) distinguishable from organ-specific responses -- framed as module-level molecular corroboration of the established neuroimmune model of electroacupuncture (Ma et al. 2021; Torres-Rosas et al. 2014), not as its discovery. (H2) The conserved module is preferentially carried by macrophage, fibroblast, and sensory-neuron subsets. (H3) The module is enriched for targets of approved analgesic/anti-inflammatory drugs. ========== 2. DATA AND ELIGIBILITY (FROZEN 2026-09-22) ========== Public GEO datasets with an acupuncture/EA vs same-batch control contrast, rat or mouse, expression matrix downloadable, total n >= 2. Candidate pool (33 bulk series, all downloaded and QC-verified as of 2026-09-23; final count set by eligibility after QC, target >=20 for Gate 1): - Spinal cord dorsal horn: GSE21258, GSE21733, GSE21758, GSE24830 - Dorsal root ganglion: GSE327205, GSE158560 - Brain regions: GSE58803 (PAG), GSE30365 (arcuate), GSE24829/24838/35138/48293 (MPTP striatum/thalamus/SN), GSE86392 (hippocampus/PFC/pituitary), GSE124387 (PFC), GSE192885 (cortex), GSE211253, GSE262257, GSE295500, GSE164727, GSE222968 - Colon: GSE57771, GSE227407, GSE164909 - Myocardium: GSE54132, GSE100499, GSE243878, GSE61840 - Skeletal muscle: GSE8911, GSE38874, GSE220191, GSE140857 - Reserve: GSE229796 (adipose), GSE131486 (retina) - Single-cell layer: GSE275032 (mouse ST36 acupoint scRNA, spliced/unspliced layers verified), GSE312510 (human BL23 snRNA), GSE272895 (rat cortex scRNA), GSE271811 (rat midline fascia scRNA) Exclusions: datasets without a same-batch non-acupuncture control; non-classical stimulation devices; acupoint-drug-intervention series. Dataset-level inclusion decisions and reasons are recorded in a public log (first code commit of the analysis repository). Inventory reproducibility and freeze certification. The candidate pool was curated by a systematic NCBI E-utilities sweep of GEO (acupuncture/electroacupuncture/moxibustion/acupoint terms; 75 candidate series screened to 33 eligible bulk + 4 single-cell series), with per-dataset exclusion reasons disclosed in the repository's first commit; this is a single-curator inventory with AI-assisted screening (disclosed as a limitation; no double independent screening is claimed). Datasets deposited after the freeze date (2026-09-22) are excluded from all discovery analyses and usable only for out-of-sample validation. ========== 3. PRE-SPECIFIED ANALYSES ========== 1. QC and processing: probe folding to gene symbol; one-to-one rat<->mouse ortholog mapping (Ensembl); per-dataset contrast: limma (microarray) or DESeq2/edgeR (RNA-seq), EA vs control within model/dataset. Known format pitfalls handled (e.g., UTF-16 matrices, underscore filenames). 2. Cross-dataset integration (primary statistics): gene-level random-effects meta-analysis of log2FC (REML, Paule-Mandel) plus Rank Product as an independent second discovery channel (C1, C5). Primary conserved module = genes significant in the RE channel (BH-FDR<0.05); organ-direction concordance is a descriptive annotation only (C4). Platform moderator + Stouffer sensitivity channel (C1); small-k tau^2 policy and minimum-coverage exclusion (C2); I^2/tau^2 descriptive with prediction intervals (C3). 3. Domain-correction baselines and deep-domain sensitivity: primary results rest on the statistical channel (Sec.3.2). Classical domain-correction baselines -- ComBat and RUV (pre-specified) -- are applied to the genexdataset logFC matrix as cross-platform harmonization diagnostics. As a supplementary sensitivity analysis, a domain-adversarial autoencoder (DAAE; adversarial branch removes organ/platform/species identity; per-gene conservation score from latent space) is compared against ComBat/RUV/rank-product with Spearman agreement reported; divergence from the pre-registered statistical channel is resolved in favor of the statistical result. ComBat is applied only within dataset (lane/batch covariates) or at the contrast level -- never across datasets with dataset as batch; ComBat and RUV are never stacked; RUV uses an external housekeeping-gene control set (C13). 4. Single-cell localization: scVI integration of the 4 sc/snRNA datasets (batch key = dataset); module scoring by UCell; module-carrier cell types ranked per dataset; group comparisons (ACU vs CFA; pre vs post). 5. Dynamic layer (mouse ST36): input pre-check -- velocity is computed only where spliced/unspliced layers are verified (GSE275032 confirmed by byte-level inspection, unspliced nnz~=5.2x10^7; other datasets enter only if unspliced matrices are available, otherwise pseudotime fallback). scVelo dynamical model (primary), stochastic and deepVelo (sensitivities); CellRank absorption/terminal states; fate-shift analysis (exploratory): cell-type x terminal-state fate probabilities compared ACU vs CFA by exact enumeration of all sample-label partitions (non-paired design, different animals per arm; exact enumeration over all C(6,3)=20 partitions; minimum attainable p=0.05 one-sided / 0.10 two-sided; C10) -- accordingly this analysis is pre-designated exploratory: effect sizes and per-cell-type shift magnitudes are the reported quantities, BH-FDR values are descriptive, and no confirmatory claim rests on it. Velocity-derived quantities are state-transition propensities, not causal fate determinations; driver genes reported with velocity-confidence filtering and matched negative-control gene sets; three-model agreement summarized in captions. Deep GRN in supplement. Confound decoupling: the 1h-vs-24h time-point structure and cell-subpopulation specificity are used to partially separate acupuncture effects from disease-progression confounds; human-layer and cross-species outputs are direction-only and hypothesis-generating. Human BL23: direction consistency of top driver genes (velocity if feasible, pseudotime otherwise) as an independent main-figure panel. 6. EA-specificity controls (mandatory): (a) overlap and enrichment contrast between the conserved module and pre-defined generic-response signatures -- MSigDB HALLMARK_Inflammatory_Response, LPS-response, and generic-stress gene sets -- the conserved module must not be reducible to these signatures (residual enrichment after partialling out, or significant module fraction outside generic signatures); (b) opportunistic within-pool contrasts: any pooled dataset containing sham/non-acupoint stimulation arms contributes a specificity contrast; (c) stimulus-specificity triangle: the conserved module is additionally compared against published exercise-training response signatures (e.g., GOTO trial / endurance-exercise multi-tissue programs) -- overlap quantified (Jaccard + enrichment) at the qualitative pathway level only (acute EA vs chronic-training time-scale mismatch disclosed; consistency != shared mechanism), with EA-unique axes reported; all three comparisons reported in a dedicated main-figure panel; (d) control-heterogeneity safeguards and falsification triggers: leave-one-control-type-out robustness; datasets whose control arm is non-comparable are annotated and droppable without narrative change; the module must shrink under negative contrasts (disease-only, non-acupoint EA, exercise signatures) -- failure to shrink falsifies H1 and is reported as such. 7. Druggable-target layer: conserved module mapped to ChEMBL (release frozen at analysis date, recorded) and Open Targets; target prioritization by ESM-2 protein embeddings + STRING PPI graph neural network, with five comparator baselines: three ablations (no-ESM, no-PPI, linear score) plus two non-DL methods -- STRING network propagation and Open Targets evidence scoring. The ranking is unsupervised and does not use drug-target annotations (stated in figure captions). Validation per C11: positives curated from DrugBank (ATC N02/M01/M02), Open Targets (>=0.4) and TTD (versions/dates recorded); primary statistic = Mann-Whitney AUC of positives over the full ranked list; recall@k curves {20,50,100,200}; degree-preserving permutations (n=1000); paired comparison vs the strongest baseline; exploratory a priori if |positives|<10; the enrichment endpoint includes length- and expression-matched random gene sets as negative controls -- absent enrichment against these controls, the endpoint is reported negative regardless of nominal p-values. 8. Organ-level predictive validation (pre-specified; methodological extension beyond discovery): for each primary organ category (6 folds; extensible to 8 if the reserve pool opens), module membership is redefined from the remaining k-1 organs, and the held-out organ's per-dataset contrasts are tested for directional concordance of module genes (sign concordance + rank correlation vs. matched random gene sets, n=1000 permutations). Decision rule per C8: the single primary metric is a pooled AUC (fold-standardized module-activity scores pooled across the 6 primary organs; organ-level bootstrap 95% CI); success = CI excludes