Data and code for "Replication-driven mutation spectrum asymmetry causes GC-skew in bacterial genomes"
Abstract
Replication-driven mutation spectrum asymmetry causes GC-skew in bacterial genomesThis repository contains the raw data, analysis scripts, output results and supplementary materials supporting the manuscript "Replication-driven mutation spectrum asymmetry causes GC-skew in bacterial genomes". The codebase analyzes previously published mutation accumulation experiment (MAE) datasets using a replication strand-specific 12-component mutational spectrum framework. It evaluates the evolutionary impact of these spectra on bacterial nucleotide composition, including equilibrium base frequencies. Additionally, the nucleotide composition of reference sequences is investigated in relation to replication strand orientation and gene expression levels. Sources for the original MAE datasets and expression data are referenced in the manuscript.Repository StructureOur workspace is organized into the following directory structure: . ├── Scripts/ ├── Data/ │ ├── MAE_data/ │ ├── Genbank_data/ │ ├── Processed/ │ └── Expression_data/ └── Results/ ├── Figures/ └── Tables/ Directory DescriptionsScripts/ — Contains source code and executable scripts used for data processing and analysis.Data/ — Input datasets and intermediate processed files:MAE_data/ — Mutational spectrum / MAE (Mutation Accumulation Experiment) raw datasets.Genbank_data/ — Reference genomes and annotation files downloaded from GenBank.Processed/ — Cleaned and preprocessed data ready for downstream analysis.Expression_data/ — E.coli gene expression dataset (e.g., RNA-Seq / TPM values).Results/ — Generated output files, tables and figures.Used software:Python 3.10.12+Environment and PrerequisitesPython: >= 3.10.12Environment Manager: venv (standard library) or Conda / MinicondaBuild System: pyproject.toml (all required Python libraries are listed in dependencies and installed automatically)Installation & SetupClone or extract the project repository and navigate into the root directory.Create and activate your preferred virtual environment (e.g., via venv or conda).Install the package and all dependencies in editable mode: pip install -e . Script Execution WorkflowTo reproduce the analysis, run the scripts in the following order:Mutation Spectra AnalysisRaw_MAE_data_standardization.ipynbDescription: Preprocesses raw MAE supplementary datasets for downstream analyses and extracts strand-specific 12-component mutations using oriC and ter sites.Mutation_spectra.ipynbDescription: Calculates 6-component and strand-specific 12-component mutation spectra for bacterial species.Permutation_script.ipynbDescription: Calculates the statistical significance of differences in mutation asymmetry between the leading and lagging strands.Split_mutations_by_strand.ipynbDescription: Partitions full-genome standardized MAE mutational data into two datasets: genes located on the leading strand and genes located on the lagging strand.Gene_strand_mutspec.ipynbDescription: Calculates strand-specific 12-component mutation spectra for leading- and lagging-strand genes across bacterial species.Gene_strand_permutation.ipynbDescription: Calculates the statistical significance of differences in strand-specific mutation spectra between genes located on the leading vs. lagging strands.Mutation_spectra_E_coli_ko.ipynbDescription: Calculates strand-specific 12-component mutation spectra of repair-defective E. coli knockout strains.Mutspec_count_script.ipynbDescription: Contains core counting functions from Mutation_spectra.ipynb refactored as reusable modules for Gene_strand_mutspec.ipynb and Mutation_spectra_E_coli_ko.ipynb.Nucleotide & Codon Frequencies AnalysisNucleotide_frequencies.ipynbDescription: Computes nucleotide counts across different genomic positions (four-fold degenerate sites, 1st, 2nd, and 3rd codon positions, and intergenic regions), both strand-specifically and combined relative to the leading replication strand.Equilibrium_nuc_frequencies.ipynbDescription: Computes expected equilibrium nucleotide frequencies across genomic regions (four-fold degenerate sites, 1st–3rd codon positions, and intergenic regions) derived from mutation spectra.Skew_graphics.ipynbDescription: Calculates various strand asymmetry metrics (GC-skew, AT-skew, and CT/GA-skew) using output data generated by Nucleotide_frequencies.ipynb and Equilibrium_nuc_frequencies.ipynb, and visualizes the results.Codon_frequencies.ipynbDescription: Computes normalized strand-specific codon frequencies across four-fold degenerate positions and calculates the inter-strand frequency ratio for each codon type.Expression Data AnalysisExpression_codon_usage.ipynbDescription: Computes strand-specific four-fold degenerate codon frequencies for highly and lowly expressed genes, calculating their ratio (high-expression vs. low-expression codon frequency).Expression_clusters_GC_skew.ipynbDescription: Performs clustering of expression data and calculates GC-skew for each cluster stratified by replication strand and expression group.Expression_permutation_test.ipynbDescription: Evaluates the statistical significance of strand asymmetry in mutation spectra across high and low gene expression groups.Manuscript StatusThis repository accompanies the manuscript:"Replication-driven mutation spectrum asymmetry causes GC-skew in bacterial genomes" In preparation.